A server for memorizing and retrieving texts based on their meaning.
This server has significant gaps across naming, parameter documentation, output schemas, and error handling. All 5 tools are explicitly registered with descriptions, but most lack rigorous parameter documentation and output schema specification. The tool names are verb-based but descriptions are often incomplete. Parameter validation, error recovery guidance, and structured output documentation are largely absent. Most tools return unstructured string responses rather than typed objects, making it difficult for LLMs to parse results and chain operations. The server relies heavily on imperative messaging to the LLM rather than structured, machine-readable responses.
Greet the user with their name and the server's name.
Memorize multiple texts for later retrieval based on relevance in meaning, not just keywords.
Chunk the contents of a PDF file into meaningful segments and store them in memory for later retrieval based on relevance in meaning, not just keywords.
Memorize a text for later retrieval based on relevance in meaning, not just keywords.
Query memory for texts similar in meaning to the query text.
Output schemas are completely undocumented. All tools return unstructured strings instead of typed objects with documented fields. The LLM cannot know what structure to expect or how to extract specific fields for downstream tool calls.
Parameter descriptions are minimal or generic. 'metadata' appears in 4 tools with only a one-line description 'Metadata to associate with the memorized content.' Missing details on what metadata structure is expected, required fields, or how it affects retrieval.
Error responses do not guide recovery. Tools return bare error strings like 'Database is not running or cannot be reached' without actionable next steps. LLM does not know whether to retry, ask user, or escalate. Error strings instruct the user ('Inform the user') rather than guiding the agent.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 47 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 39 | - | v1 |
Tool 'remember_similar_texts' lacks bounds validation. Parameter 'n_results' accepts any integer with no documented min/max. Code validates n_results > 0 but schema has no constraints. For a memory query, unbounded results could bloat context window.
Tool 'memorize_pdf_file' has compound logic and confusing return semantics. It returns a text prompt instructing the LLM to call 'memorize_multiple_texts', effectively deferring work to the agent. This breaks composability, the tool should either memorize internally or return structured metadata for the agent to parse, not embed natural language instructions in responses.
No pagination support documented. Tool 'remember_similar_texts' returns up to n_results matches, but there is no mechanism documented for retrieving additional results (offset, cursor, total_count). If memory grows large, a single call cannot return all matches.
Metadata parameter has unsafe default `{'topic': 'memory'}`. This means all memorized texts share identical metadata unless explicitly overridden. If user forgets to provide metadata, all future retrievals become harder to filter. Default should be empty dict or require explicit metadata.