A FastAPI-based MCP server for storing, searching, and managing project memories using ChromaDB with semantic vector search
The server defines 5 tools with explicit schemas and basic descriptions. Tool naming follows verb_noun convention (store, search, update, delete, all) which is good. However, descriptions are uniformly short (13-31 chars), below the 50-200 char baseline for LLM-optimized tools. Parameter descriptions exist but are minimal (3-18 chars). Input schemas are present and use proper JSON Schema with types, but lack important constraints: enums for risk-based operations, bounds on n_results, pattern validation for IDs. Output schemas are not documented, critical for agent chaining. Error handling is absent, no guidance on recovery, retryability, or what the LLM should do on failure. The server stores no secrets but lacks audit logging and permission gating for destructive operations (delete, update). Tool composition is reasonable (each does one thing), but missing pagination, batch operations, and clear chaining IDs.
List all stored facts
Delete a stored fact by id
Search stored project facts
Store an important fact about the project
Update a stored fact by id
Descriptions are critically short (13-31 chars). Below 50-char baseline, they provide minimal context for LLM tool selection. 'Delete a stored fact by id' doesn't explain consequences, retryability, or when to use it vs alternatives.
Output schemas not documented. Callers cannot infer what fields to expect (does store return id, timestamp, tags?). This breaks agent chaining, downstream tools can't reference returned IDs without guessing.
No error handling or recovery guidance. If memory.delete fails (ID not found), the LLM gets no hint about what to do next (retry, search first, or accept failure?). No distinction between retryable and fatal errors.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 45 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Destructive operation (memory.delete) lacks confirmation or dry-run pattern. Agents make mistakes, no guard against accidental deletion of important facts.
Parameter 'n_results' lacks bounds. Unbounded integers let LLMs pass absurd values (0, -1, 999999) that break the tool or cause timeouts. Should specify 1-100 range with default=5.
memory.search lacks pagination support. Large result sets blow context windows. Schema shows no offset/limit or cursor fields, forcing all results into one response.
memory.all has no limit on returned facts. If a project has 10,000 memories, this returns all of them, exhausting context. Should cap at 20-50 with pagination.
memory.update has optional 'content' and 'tags' but no validation that at least one is provided. LLM could call with only 'id', performing a no-op that wastes a round-trip.
No audit logging or permission gating for destructive operations. Who called delete? When? With what result? No trail for compliance or debugging.
Tool names lack 'codename' context. In multi-tenant scenarios, 'memory.store' is ambiguous, store to which project? Names should reflect namespace or codename selection is implicit/external to tool interface.