Local MCP memory search for Codex CLI and Claude Code conversation history
Server has 6 tools with complete input schemas and descriptions. Naming follows verb_noun pattern (search_*, list_*, get_*, refresh_*, suggest_*). Descriptions are substantive (100-250 chars) and include use-case guidance. However, output schemas are not documented, LLMs cannot predict response structure. Parameters lack some constraints (e.g., max_results unbounded in suggest_skills). Error handling is minimal; tools return plain strings rather than structured errors with recovery guidance. No tool annotations (readOnlyHint/destructiveHint) despite clear risk levels. Composition is sound: tools chain logically (list_sessions → get_session). Parameter descriptions are good but could specify ranges and formats more explicitly.
Retrieve the full conversation for a specific session. session_id: from list_sessions output source: "codex" | "claude" max_messages: limit messages returned (default 30) to control token usage. Prefer search_history for finding context; use this only when you need the full conversation flow.
List available past sessions with their titles, dates, and sources. source: "all" | "codex" | "claude" Returns session IDs needed for get_session.
Manually refresh the persistent graph index and in-memory keyword cache. rebuild=true deletes and recreates the derived graph database. It never modifies Codex or Claude history files.
Search the persistent local knowledge graph built from past conversations. Use this for relationship-heavy questions where exact keywords may differ across sessions, such as "where did I solve a similar auth migration bug?". sources can include "codex" and/or "claude".
Search past AI coding agent conversations for a keyword or topic. Returns matching message excerpts with surrounding context. sources can include "codex" (OpenAI Codex CLI) and/or "claude" (Claude Code). Prefer this over get_session to avoid token bloat. Example: search_history("CUDA illegal address", sources=["codex", "claude"])
Output schemas not documented. LLMs cannot predict response structure (format_hits, format_graph_hits, format_skill_candidates return plain strings). Agents cannot extract IDs or chain downstream calls reliably.
No tool annotations (readOnlyHint/destructiveHint). refresh_history_index is marked WRITE in metadata but tool definition lacks destructiveHint. Agents cannot distinguish safe vs. risky operations.
Error handling returns plain strings ('Invalid sources...', 'No sessions found...') instead of structured errors with recovery guidance. LLMs cannot parse error type or determine if retry is safe.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | B | 72 | 2026-07-28+ | v2 |
Suggest reusable skills that could be created from repeated chat patterns. Returns ranked skill ideas with trigger phrases, supporting sessions, and boundaries. This never writes SKILL.md files and never calls a model.
Parameter constraints incomplete. max_results, max_candidates, days_back lack explicit min/max bounds. suggest_skills.days_back defaults to null (unbounded), should specify range or clarify 'null means all history'.
list_sessions returns session IDs but output format is unspecified. Agents cannot reliably extract session_id for downstream get_session calls if response is plain text.