An MCP server for semantic memory storage and retrieval using ChromaDB and sentence-transformers embeddings
copilot-memory-mcp demonstrates solid definition quality with consistent naming, complete input schemas, and clear parameter descriptions across all 5 tools. All tools follow verb_noun naming convention (create_, search_, update_, delete_, list_). Schemas are well-formed with proper types and nullable markers. However, tool descriptions lack depth and strategic guidance, they document WHAT each tool does but omit WHEN to use it and dependencies between tools. Error handling guidance is minimal. Output schemas are documented informally in docstrings rather than formally in the schema definitions. The server shows no tool annotations (readOnlyHint/destructiveHint/idempotentHint), which is a gap given the clear risk classifications (WRITE, READ_ONLY, DESTRUCTIVE). Descriptions are between 130-280 chars (within acceptable 10-1024 range per pattern:tool-description baseline), but could be more strategically written to guide LLM composition.
Store a new memory. Args: title: Short human-readable label. content: Full memory body — this text is embedded for semantic search. project_name: Optional project scope (omit for global memories). tags: Optional free-form categorical tags. Returns: ``{ id, title, created_at }``
Permanently delete a memory by ID. Args: id: UUID of the memory to delete. Returns: ``{ id, deleted: true }`` Raises: ValueError: If no memory with *id* exists.
Browse memories without a semantic query (paginated). Returns lightweight records — title only, no content. Args: project_name: Filter by project (omit = global). tags: AND-match filter by tags (omit = no tag filter). page: 1-indexed page number (default 1). page_size: Results per page (default 20, max 100). Returns: ``{ items: [{id, title, project_name, tags, created_at}], page, page_size, total, total_pages }`` Raises: ValueError: If *page* < 1 or *page_size* > 100.
Semantic vector search over stored memories. Args: query: Natural language search query. project_name: Filter to a specific project (omit = global search). tags: AND-match filter by tags (omit = no tag filter). limit: Maximum number of results to return (default 5). Returns: List of ``{ id, title, content, project_name, tags, score, created_at }``.
Tool annotations (readOnlyHint, destructiveHint, idempotentHint) are missing. The server explicitly classifies tools by risk level (READ_ONLY, WRITE, DESTRUCTIVE) in the metadata, but these are not propagated as formal tool annotations in the MCP schema. This prevents clients from making safe assumptions about tool behavior without parsing comments.
Tool descriptions lack strategic guidance on when to use each tool and dependencies. For example, search_memories does not mention that agents should typically search before updating or deleting (to find the ID). Similarly, no description advises that create_memory should be called before searching for newly created memories. Per pattern:tool-description baseline, descriptions should answer 'When should the LLM call it instead of a similar tool?', currently omitted.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | A | 83 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 54 | - | v1 |
Update an existing memory by ID. Only supplied fields are changed; the embedding is recomputed when title or content changes. Args: id: UUID of the memory to update. title: New title (omit to keep existing). content: New content (omit to keep existing). tags: Replacement tag list (omit to keep existing). Returns: ``{ id, title, updated_at }`` Raises: ValueError: If no memory with *id* exists.
Error handling descriptions are minimal or absent. delete_memory and update_memory mention 'ValueError: If no memory with *id* exists', but do not advise the agent what to do, e.g. 'If memory not found, try list_memories() or search_memories() to find the correct ID.' Per pattern:recovery-guide, errors must tell the LLM what to do next.
Output schema for search_memories returns a 'score' field (vector similarity score) that is not documented in the description. LLMs may not understand what 'score' represents or how to interpret it. Per pattern:response-shaper, document all fields in the output.
list_memories description mentions 'Returns lightweight records, title only, no content', but the tool actually returns full metadata including created_at, tags, project_name. The description undersells the output and could confuse agents about what data is available.