Local index and search for AI coding-agent conversation threads. Exposes thread indexing, search, snapshots, knowledge extraction, and decision tracking as MCP tools.
Callimachus MCP provides 19 tools with mostly clear naming and good descriptions. Tool names follow verb_noun convention (search_threads, get_thread, snapshot_session, etc.). Descriptions are generally strong (100-250 chars, well above the 34 char baseline minimum). Parameters have types and descriptions. However, OUTPUT SCHEMAS ARE NOT DOCUMENTED, the tool definitions provided show input schemas but no documented return types or structure. This is a critical gap for pattern:tool and pattern:response-shaper compliance. Additionally, STDIO-only transport caps protocol readiness at 50, which impacts the overall server score. The tools themselves demonstrate good domain design for thread/conversation history management, but lack the polish needed for A-grade (80+) production readiness.
Ask a question and get a synthesized, cited answer drawn from the user's indexed threads. Requires knowledge/LLM enabled. Returns the answer plus [thread N] citations.
Before making a change/decision, check if a related decision has already been made (and retrieve its rationale). Returns settled decisions and their reasoning.
Mark a TODO as complete by its id. Idempotent: marking an already-complete TODO succeeds silently.
Fetch one indexed thread as a packed markdown transcript (budget-limited, ready to drop into context). Pass a threadId from search_threads.
Find the commits that a conversation thread produced, by overlapping the files discussed with git log in the thread's time window.
List saved session snapshots (newest first), optionally scoped to a project-path substring. Returns snapshot metadata with an id to load.
OUTPUT SCHEMAS NOT DOCUMENTED. No tool definition includes a documented return type, structure, or field schema. LLMs cannot infer what fields will be returned, breaking pattern:tool-chain (where tool A's output must include IDs for tool B). This forces agents to make assumptions about response structure and blocks composition verification.
PAGINATION NOT DOCUMENTED IN TOOL DEFINITIONS. search_threads, recent_threads, list_todos, and recall have a 'limit' parameter but no 'offset', 'cursor', or 'next_cursor' fields documented in their schemas or descriptions. Without pagination information, agents cannot iterate through large result sets (pattern:paginated-result violation).
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 65 | 2026-07-28+ | v2 |
List all user-defined tags currently in use (across all threads).
List open TODOs extracted from threads, optionally filtered by text search, project, source.
Load a saved session snapshot by id (from list_snapshots or snapshot_session). Returns the packed transcript and project memory carry-forward.
Retrieve the distilled memory for a project: aggregated decisions, gotchas, and open TODOs across all its threads.
Search within a project's threads only. Faster and more focused than global search when you know which repo you're in.
Semantic recall of remembered facts by query. Returns facts relevant to the query with citations to their source threads.
List recent threads (newest first), optionally filtered by source, project, starred status, or tags.
Find clusters of recurring errors/issues from the thread history, optionally scoped to a project and time window.
Record a fact/decision for later semantic recall. Useful for capturing patterns, lessons, or design decisions that span multiple threads.
Find all threads that discussed or touched a given file path (keyword and semantic search over file mentions).
Search the user's indexed AI coding-agent conversation threads across every tool they use (Claude Code, Codex, Cursor, Gemini, Qwen, Goose, OpenCode, Continue, Cline, Roo Code, Kilo Code, and in-app chats). Keyword full-text by default; set hybrid=true to also use on-device semantic similarity. Returns matching threads with snippets and a threadId to fetch. Use this to recall past decisions, prior solutions, or earlier discussion before redoing work.
Save a resumable SNAPSHOT of a thread: its packed transcript plus a carry-forward block of the project's distilled decisions/gotchas/open TODOs. Use this to checkpoint a session before context is compacted, or to hand work off to another agent or tool. Returns the snapshot id to load later. Pass a threadId from search_threads.
Extract distilled knowledge from a single thread: its decisions, gotchas, and open TODOs (LLM-powered if enabled; heuristic extraction otherwise).
PARAMETER ENUM CONSTRAINTS MISSING. search_threads 'sources' parameter lists valid values in the description ('claude_code, codex, cursor, gemini, qwen, goose, opencode, continue, cline, roo, kilo, in_app') but doesn't declare them as a JSON Schema enum. Same for recent_threads. Free-form string parameters invite LLM hallucination; pattern:constrained-input requires enums for known-set values.
NO ERROR HANDLING GUIDANCE IN DESCRIPTIONS. No tool description tells the LLM what to do when an operation fails (pattern:recovery-guide). E.g. search_threads doesn't explain: What if no threads match? What if the embedding model fails to load? What if the database is unavailable? Error handling must be documented so LLMs can plan recovery.
STATEFUL OPERATIONS NOT MARKED WITH DESTRUCTIVE HINTS. snapshot_session, complete_todo, and remember modify state (write/delete operations), but tool definitions lack 'destructiveHint' or 'idempotentHint' annotations (pattern:tool-annotation). LLMs cannot distinguish safe retries from operations that need confirmation.
PARAMETER RELATIONSHIPS NOT DOCUMENTED. remember accepts both 'project' (optional) and 'rationale' (optional), but no description states whether 'project' defaults to the current repo or whether 'rationale' is required for decisions. undocumented parameter dependencies cause silent misuse (pattern:tool-description violation).
NUMERIC RANGE CONSTRAINTS MISSING. list_todos, search_threads, recent_threads, and recall all have 'limit' parameters but no documented min/max (e.g., 'min: 1, max: 100'). Unbounded limits let LLMs request thousands of results, risking context window exhaustion and timeouts.