Persistent, project-aware, long-term memory for AI coding assistants
Heimdall has well-structured tool definitions with consistent naming patterns (verb_noun: store_memory, recall_memories, session_lessons, memory_status, delete_memory, delete_memories_by_tags). All 6 tools have descriptions and documented input schemas. However, descriptions are functional but lack guidance on when/why to use each tool, parameter descriptions are minimal, output schemas are not documented, and error handling guidance is absent. The server follows good JSON Schema structure with types and constraints (enums, min/max), but descriptions average ~80-120 chars and are more technical than LLM-optimized. No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite two destructive tools. The memory system is domain-specific and well-reasoned, but tool definitions could better support agent decision-making.
Delete all memories that have any of the specified tags. Provides preview for safety.
Delete a single memory by its ID. Provides preview of what will be deleted for safety.
Get cognitive memory system health and statistics
Retrieve memories based on query with rich contextual information
Capture and consolidate key learnings from current session for future reference. This tool encourages metacognitive reflection - thinking about what you've learned that would be valuable after 'critical amnesia' between sessions.
Store new experiences or knowledge in cognitive memory for future recall
Output schemas not documented. Tool descriptions state what each tool does but do not document what fields or structure will be returned. LLMs cannot plan downstream actions or extract relevant data without knowing return types.
No tool annotations (destructiveHint, readOnlyHint, idempotentHint). delete_memory and delete_memories_by_tags are clearly destructive operations but lack structured hints. LLMs need explicit markers to reason about safety and reversibility.
No error recovery guidance. Tool descriptions do not explain what errors might occur, how to interpret them, or what the agent should do next. E.g., recall_memories could fail if Qdrant is down or query is invalid, but descriptions offer no guidance.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 52 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Parameter descriptions are technical but lack LLM-optimized context. E.g., hierarchy_level is described only as '0=concept, 1=context, 2=episode' without explaining when/why an agent would choose each level. The description does not answer: when should I call this with hierarchy_level=1 vs 0?
Destructive tools lack dry-run guidance in descriptions. delete_memory and delete_memories_by_tags both offer dry_run parameters but descriptions do not explain the safety value or recommend always previewing first.
memory_status description is vague. 'Get cognitive memory system health and statistics' does not explain what 'health' means, what statistics are returned, or when an LLM should call this vs recall_memories.
No guidance on tool dependencies or ordering. E.g., should an LLM call memory_status before store_memory to check capacity? Can session_lessons and store_memory coexist or do they compete? Descriptions do not clarify.