MCP server that indexes Claude Code conversation history and provides semantic search
Two well-named search tools with excellent descriptions and comprehensive parameter schemas. Tool names (search_project_history, search_global_history) follow verb_noun convention and clearly distinguish scope. Descriptions are 200+ characters, context-rich, and explain WHAT, WHEN, and WHY. All 8 input parameters are typed and documented. However, output schemas are not explicitly documented in the visible code, only mentioned in docstrings as 'returns dict'. The schema completeness is inferred from the Pydantic model_dump() call but the actual structure is not visible. Error handling is minimal, no recovery guidance, no invalid input examples, and no edge case documentation (e.g., what happens if query is empty, or threshold is out of bounds). Both tools are READ_ONLY with appropriate risk marking, but lack the defensive input validation patterns common in production tools.
Search conversation history across ALL PROJECTS. Use this to find cross-project knowledge: user preferences, coding patterns, common solutions, global conventions, and insights from all previous work. This searches ALL previous Claude Code sessions regardless of project.
Search conversation history for the CURRENT PROJECT only. Use this to find project-specific context: past decisions, implementation details, bugs discussed, architecture choices, and previous work done on this codebase. This searches previous Claude Code sessions for the project you're currently working on.
Output schema not explicitly documented. Docstrings mention 'Search results with matched messages, scores, session_id, surrounding context, and pagination info' but the actual response structure (field names, types, nested objects) is not visible in code. LLMs cannot plan follow-up tool calls without knowing the exact schema.
No error handling guidance. Missing recovery patterns for common failures: empty query, out-of-range threshold (must be 0-1), invalid ISO 8601 dates, no matches found, memory/timeout errors. No actionable error messages shown in code.
No input validation visible. Parameters like threshold (0-1 range), max_results (should have max limit), offset (should be >= 0), and ISO 8601 dates (after/before) accept user input without documented constraints or validation. LLM may pass invalid values.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 63 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 53 | - | v1 |
Parameter defaults are not justified. Why default threshold to 0.2? Why context_before_after to 3? Why max_results to 10? Undocumented defaults make it hard for LLMs to predict behavior or request adjustment.
No pagination guidance in descriptions. Both tools mention 'offset' and 'has_more' in return docstrings but do not explain pagination strategy (is max_results + offset the only way? can I use a cursor?) or edge cases (what if offset > total_matches?).