Annal demonstrates solid definition quality with clear, descriptive tool names following verb-noun convention and well-documented parameters. All 9 tools have substantive descriptions (avg ~150 chars) and typed input schemas with parameter descriptions. Notable strengths: natural language search semantics, tag-based filtering, cross-project search, and structured output modes. Key weakness: output schemas are not formally documented in the code, search_memories and expand_memories return content but lack explicit schema definitions visible in server.py. Tool compositions are well-designed (search → expand workflow), and error handling instructions are embedded in server instructions rather than tool-level recovery guidance. Security considerations are present (tagging for agent identity) but no explicit permission gates visible.
Permanently delete a single memory by ID.
Fetch full content of specific memories by ID. Use after probe-mode searches.
Get statistics about a project's memory store.
Initialize a project with watch paths for file indexing.
Review and clean up memories that are no longer being accessed.
Fix or refine tags on a memory without changing content.
Search project memories by semantic similarity. Returns results with context.
Output schemas not formally documented. search_memories and expand_memories return content but lack explicit JSON schema definitions describing the structure of results (e.g., field names, types, nesting).
Error handling lacks actionable recovery guidance at tool level. Server instructions mention error recovery, but individual tool descriptions do not explain what errors can occur, how to classify them (retryable vs fatal), or what the LLM should do next.
Destructive operations (delete_memory, prune_stale) lack dry-run or confirmation patterns. prune_stale has a dry_run flag, but delete_memory does not, risking irreversible data loss if an LLM makes a mistake.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 70 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 54 | - | v1 |
Store 2+ memories at once. More efficient than multiple store_memory calls.
Store a piece of knowledge in a project's memory.
Parameter descriptions could better specify constraints. store_memory tags accepts 'array or string' but does not explicitly state length limits, character restrictions, or valid tag naming conventions (though server instructions define conventions).
store_batch lacks explicit validation guidance. The 'memories' array parameter description is minimal ('List of memory objects...') and does not specify the schema of each memory object (required vs optional fields, valid field names).
Permission gates not visible. No explicit tool-level permission checks or scope declarations visible in server.py. Security relies on server-level project isolation rather than per-tool authorization.