Neural Memory Graph - Knowledge Graph Memory System with MCP SSE endpoint. Provides semantic search, entity extraction, graph reasoning, and sleep-time consolidation for long-term conversational memory.
HippoGraph presents a specialized neurosymbolic memory system with 12 tools covering search, CRUD, and graph maintenance operations. Strengths: most tools have clear, descriptive names starting with action verbs (search_, add_, update_, delete_, get_); input schemas are visible and include type definitions; descriptions exist for all tools and most parameters. Weaknesses: descriptions vary significantly in informativeness (some are excellent, others are formulaic); several tools lack error handling guidance; output schemas are not documented in the source; parameter descriptions sometimes lack constraint clarity (e.g., search_memory's `limit` lacks explanation of why 20 is the max); emotional tone parameters in add_note are domain-specific but lack validation guidance; no confirmation/dry-run pattern for destructive operations (delete_note). Schema completeness is good (types, enums, defaults present) but output structure is not documented, forcing LLMs to infer result shapes.
Add new note with automatic entity extraction, linking, and emotional context. Checks for duplicates.
Delete note by ID
Find notes similar to given content. Useful for checking before adding new notes.
Get graph connections for a specific note
Get version history for a note
Get statistics about stored notes, edges, and entities
Output schemas not documented. Tools return data but LLMs cannot infer response structure. E.g., search_memory returns results but no schema defines fields like 'id', 'content', 'category', 'score'. This forces LLMs to guess at response shape and wastes tokens on clarification attempts.
Destructive operation (delete_note) lacks confirmation/dry-run pattern. Agents can permanently delete notes without preview or confirmation step. No recovery guidance in error handling. Missing pattern:confirmation-request.
Error handling lacks recovery guidance. Tool descriptions do not indicate what the LLM should do if a call fails. E.g., update_note with invalid note_id returns an error but no hint to search_memory() first or try a different ID. Missing pattern:recovery-guide.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 71 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Restore a note to a previous version
Search through notes using spreading activation algorithm
Get search quality monitoring stats: latency percentiles, zero-result queries, phase breakdown. Helps identify retrieval issues.
Set importance level for a note: 'critical' (2x boost), 'normal', or 'low' (0.5x)
Run sleep-time graph maintenance: consolidation (thematic clusters + temporal chains), PageRank recalculation, orphan detection, stale edge decay, duplicate scan. Zero LLM cost. Use dry_run=true to preview without changes.
Update existing note by ID
Emotional context parameters (emotional_tone, emotional_intensity, emotional_reflection) in add_note lack validation constraints. No enum, pattern, or range guidance for what constitutes valid emotional_tone strings. Invites hallucinated values like 'confused_happy' or '11' for intensity.
Parameter descriptions lack constraint clarity. search_memory's 'limit' defaults to 5 with max 20, and 'max_results' defaults to 10 with max 50, why two limits? No explanation. time_after/time_before mention ISO format but no validation error examples. Category parameter is optional but no enum of valid categories documented.
Tool composition inefficiency: search_memory and find_similar both search but with different algorithms (spreading activation vs similarity threshold). No guidance on when to use each. LLMs will waste reasoning cycles deciding between them. Naming distinction is weak.
update_note and restore_note_version both modify note content, but restore is framed as history recovery. No clear separation of concerns. If an LLM wants to correct a typo, should it use update or restore? Parameter semantics are conflated.