RAG knowledge base server — MCP tools + HTTP ingestion endpoints backed by Qdrant.
Archivist has 4 tools with clear names and generally good descriptions. All tools have input schemas with types and descriptions. However, output schemas are not documented, parameter descriptions lack depth in some cases, and error handling is minimal. Tool naming follows verb_noun convention (search, save_memory, recall, recent_memories). Descriptions are concise (10-60 chars) but lack detail on what happens internally, when to use each tool vs others, and what the LLM should expect from responses. Schemas are present with types and defaults, but lack constraints (enums, ranges, patterns). No tool annotations (readOnlyHint/destructiveHint), no documented output structure, and no error recovery guidance.
Semantically search past memories
List the most recent entries in reverse chronological order
persist a memory entry with timestamp and tags
Embed the query and return the closest matching document chunks by cosine similarity.
Output schemas not documented. Tools return results but LLMs cannot plan downstream calls or extract structure without knowing field names and types.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). The MCP spec (2026-07-28) includes these as standard; save_memory and recall should be annotated with readOnlyHint=false and readOnlyHint=true respectively for clarity.
Parameter descriptions lack constraints and examples. 'limit' has no min/max (defaults to 5/10, max 50 for search, but ranges not enforced or documented). 'tags' array in save_memory lacks guidance on valid tag values, length limits, or examples beyond the brief guide.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 62 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 47 | - | v1 |
Descriptions are too brief (50-60 chars). The rubric baseline for A+ tools is 50-200 chars. Current descriptions (e.g., 'Embed the query and return the closest matching document chunks by cosine similarity') do not address WHEN to use each tool, how it differs from similar tools, or what the response structure looks like.
No error handling or recovery guidance. Tools do not document failure modes (e.g., no documents indexed, invalid query, Qdrant service down) or next steps for the LLM.
Pagination not documented for search and recall. No mention of offset, cursor, or total count in responses, yet limit is configurable up to 50. Large result sets could exhaust context without clear pagination guidance.
save_memory description does not explicitly state it is a destructive/mutating operation. LLMs need to know which calls are safe to retry and which have side effects.