Local-first AI PKM for coding conversations — supports Claude Code, Cursor, Codex CLI, Trae, GitHub Copilot. MCP server for semantic search and task memory management in AI conversation knowledge base.
ChatCrystal presents a well-structured knowledge management MCP server with 7 tools, most of which have clear, task-focused descriptions and documented input schemas. Strengths: all tools have descriptions (100+ chars each), parameter types are explicit, and the tool set follows a logical single-responsibility pattern (search, retrieve, validate, persist). Weaknesses: output schemas are not documented in the visible code, error handling guidance is minimal, no tool annotations (readonly/destructive hints) are present, and the write_task_memory tool lacks confirmation/dry-run patterns typical of destructive operations. The descriptions are LLM-optimized and situational (e.g., recall_for_task explicitly states 'Use this at the beginning...'), which is excellent for agent selection. Schema definitions are complete for inputs but outputs are inferred from descriptions only.
Get the full content of a note including title, summary, key conclusions, code snippets, and tags.
Get related notes for a given note, including relationship type and confidence score.
List note summaries for browsing and narrowing the ChatCrystal knowledge base. Use this when you need paginated notes filtered by tag or title/summary keyword. Use search_knowledge instead for semantic relevance ranking, and get_note when you already have an id and need the full note body. Returns note metadata and summaries, not full note content.
Retrieve reusable task memories before starting substantive coding work. Use this at the beginning of implementation, debugging, migration, configuration, investigation, refactor, or optimization tasks to load project-scoped memories first and optional global lessons second. Use mode="debug" when the user reports an error, failing command, regression, or incident; include error_signatures and related_files when available. Use search_knowledge instead for ad hoc semantic note search that is not tied to the current task. This tool is read-only and returns ranked memories plus optional related-note context without writing anything.
Output schemas not documented. Tool descriptions infer result structure (e.g., 'Returns matching notes ranked by relevance', 'Returns acceptance, rejection reason, warnings...'), but no JSON Schema or TypeScript interfaces are visible in the provided source. LLMs cannot plan downstream tool calls or extract fields reliably without explicit output type documentation.
write_task_memory lacks confirmation/dry-run pattern. This is a destructive write operation (persists to knowledge base) with no dry-run step. The validate_task_memory tool exists and is read-only, but write_task_memory should either (a) require explicit confirmation via MRTR (Multi Round-Trip Requests with result.type='input_required'), or (b) document clear guardrails on what qualifies as 'high-quality' to prevent weak auto-writes. Currently, the description says 'Weak auto writebacks are skipped...' but no error message documents the rejection reason for the agent to learn from.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 63 | 2026-07-28+ | v2 |
Semantic search across your AI conversation knowledge base. Returns matching notes ranked by relevance.
Dry-run validation for a candidate task memory before calling write_task_memory. Use this after meaningful work and before persisting a lesson to check whether the candidate is durable, specific, reusable, and shaped like a high-quality ChatCrystal note. It has no side effects and never writes to the knowledge base. Returns acceptance, rejection reason, warnings, and materialized note fields so agents can revise the candidate or skip weak work logs.
Persist a task memory only when it can become a high-quality ChatCrystal note: specific title, concrete summary, meaningful key conclusions, and a durable reusable lesson such as a pitfall, fix, decision, pattern, or symptom-to-resolution mapping. Do not write one-time environment checks, version/status reports, ordinary progress logs, or vague robustness claims. Weak auto writebacks are skipped by core validation and recorded only as receipts.
No tool annotations (toolAnnotations=false). The server does not emit readOnlyHint, destructiveHint, or idempotentHint in tool definitions. This forces the LLM to infer safety from descriptions alone. Best practice: mark all read-only tools (search_knowledge, get_note, list_notes, get_relations, recall_for_task, validate_task_memory) with readOnlyHint=true, and write_task_memory with destructiveHint=true.
Error handling lacks actionable guidance. Descriptions mention validation ('Returns acceptance, rejection reason...') and note the read-only nature of validate_task_memory, but no explicit error response format or recovery hints are documented. For example, if recall_for_task finds no memories, does it return an empty array or a null result? If write_task_memory rejects a candidate, what should the agent do next, revise and retry, or skip?
Parameter 'limit' in search_knowledge has no range constraints. Default is 10, but no min/max documented. Should this accept 1 - 100, 1 - 1000, or unlimited? Unbounded numeric parameters invite absurd values (limit=999999) that may cause timeouts or memory exhaustion.