Persistent shared memory for Claude Code — MCP server with DuckDB vector search, semantic dedup, and ambient intelligence for team knowledge
Distillery has 19 tools with consistent schemas and descriptions, but quality is uneven. Naming follows verb-noun conventions well (store, get, search, classify, etc.). Most tools have descriptions and input schemas, but many descriptions are generic and lack clarity about WHEN to use each tool versus similar alternatives. Critical gap: no output schemas documented anywhere, LLMs cannot see what fields they'll receive back, forcing them to guess at response structure. Parameter descriptions exist but are often minimal (10-30 chars), below the 72-char baseline. Error handling and security posture are undocumented. Tool composition is reasonable but some overlaps exist (search vs find_similar, watch/poll/rescore are vaguely related). The server correctly avoids exposing secrets as parameters, but permission gates and audit trails are not evident from code.
Classify an entry with custom categories or schemas
Configure server settings and parameters at runtime
Mark an entry as incorrect or flag it for review
Find entries similar to a given entry using vector similarity
Retrieve an entry by ID from the knowledge base
Sync data from GitHub repositories
Ingest a document from a URL or file path into the knowledge base
Output schemas are never documented. LLMs cannot see what fields each tool returns (e.g., does 'search' return score, entry_id, body, tags, all of these?). This forces agents to hallucinate response structure and guess at field names for downstream operations.
Parameter descriptions are too short and generic. Examples: 'Optional source URL' (19 chars), 'Optional tags for categorization' (31 chars), 'Maximum number of entries to return' (35 chars). These fall well below the 72-char baseline and lack context for LLM decision-making. No guidance on format, range, or constraints.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 57 | 2026-07-28+ | v2 |
List entries from the knowledge base with optional filtering
Manually trigger polling of configured feed sources
Discover and manage relationships between entries
Rescore entries based on relevance or other criteria
Resolve a flagged entry review by accepting or rejecting changes
Search the knowledge base using semantic vector search
Get the current status of the Distillery server
Store or update an entry in the knowledge base
Get the status of ongoing sync operations
Get and display the type schemas used by the system
Update an existing entry in the knowledge base
Set up a feed source to continuously watch for new content
Tool descriptions lack context for selection and disambiguation. 'Classify an entry with custom categories or schemas' does not explain WHEN to use classify vs store+tags. 'Rescore entries based on relevance or other criteria' is vague, what criteria? How does it differ from search? LLMs will conflate similar tools.
No error handling guidance. Tools like 'ingest_doc' (fetch from URL), 'gh_sync' (GitHub auth), and 'configure' (set runtime parameters) have obvious failure modes (404, auth errors, invalid config). No indication of what errors are possible, how to recover, or whether retries are safe.
Tool composition has overlaps and unclear distinctions. 'search' (semantic vector search), 'find_similar' (vector similarity), and 'relations' (relationship traversal) all find related entries but with different semantics not explained. 'watch', 'poll', and 'gh_sync' all ingest content but the orchestration is unclear.
No permission gates or scope declarations visible. 'correct', 'update', 'delete', and 'configure' are sensitive mutations, but no indication of required permissions, who can call them, or audit logging. Tool definitions do not state what access level they require.
Pagination support present in 'list' but not documented clearly. No indication of whether results are capped, whether total_count is returned, or how to paginate through large result sets. 'search' and 'find_similar' both have 'limit' but no 'offset' or 'cursor', unclear how to fetch next page.
No idempotency hints or confirmation patterns for destructive/write tools. 'correct', 'update', 'store' all mutate state but offer no dry-run, undo, or confirmation step. Agents retrying on ambiguous failures could create duplicate entries or overwrite data unintentionally.