MCP server for examining and understanding stored memories using RAG (Retrieval Augmented Generation) with vector embeddings and semantic search
The Pensieve has 5 tools with visible schemas and descriptions, but multiple critical quality gaps reduce the overall score. All tools have basic descriptions (50-90 chars) and input schemas with type declarations, but descriptions lack LLM-optimization details (no WHEN to use, no prerequisites, no error guidance). Parameter descriptions are minimal or absent. Output schemas are entirely undocumented, users cannot see what fields to expect from tool responses. Tool naming follows verb_noun convention (process_*, query_*, generate_*) but lacks semantic clarity about differences between similar operations (process_knowledge vs process_markdown). Error handling is not visible in the source code. No per-tool improvement suggestions are evident. The server uses STDIO transport, which is a hard cap at 50 for protocol readiness, making remote testing impossible.
Generate a new markdown document
Process knowledge files from root directory
Process a single markdown file directly
Query the processed knowledge using RAG
Examine your stored memories and knowledge, allowing patterns and connections to emerge
No output schemas documented for any tool. LLMs cannot infer what fields responses contain, forcing them to guess about downstream tool chaining (e.g., does query_knowledge return metadata for follow-up filtering?). This violates pattern:tool and wastes tokens on context-window exploration.
Parameter descriptions are sparse or missing detail. 'query' in query_knowledge lacks guidance on format/length constraints. 'content' in process_markdown does not specify encoding, structure, or size limits. 'topic' in generate_markdown does not hint at complexity expectations. Per-parameter descriptions average ~40 chars (well below the 72-char baseline for production tools).
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 44 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 33 | - | v1 |
Tool descriptions lack WHEN/WHY guidance. 'Process knowledge files from root directory' does not explain when process_knowledge should be called vs process_markdown, what 'force' does, or what side effects occur. Descriptions should be 50-200 chars with explicit prerequisites and state-change warnings (per pattern:tool-description baselines).
Semantic naming ambiguity: process_knowledge, process_markdown, and generate_markdown perform different operations (batch processing, single-file processing, generation) but names do not clarify the distinction. LLMs may conflate them. Suggest renaming to batch_process_knowledge_files, process_single_markdown, and create_markdown_document for clarity.
No error handling guidance visible in tool definitions or descriptions. If process_knowledge fails to read a file, what should the agent do? If query_knowledge finds 0 results, should it retry or fallback? No recovery patterns (pattern:recovery-guide) or error classification (pattern:error-classification) documented.
use_pensieve description is vague: 'Examine your stored memories and knowledge, allowing patterns and connections to emerge' is poetic but not actionable for LLMs. Does this tool summarize knowledge? Return related items? Perform reasoning? The description must be explicit about inputs and outputs.
maxResults parameter in query_knowledge has no minimum/maximum bounds documented. Can an LLM pass maxResults=999999? No constraint prevents runaway result sets that blow context windows. Numeric parameters should declare ranges (e.g., 1 - 100).
No idempotency guarantees. process_knowledge and generate_markdown modify state (write files, embed vectors). If called twice with same input, do they create duplicates or skip? If the agent retries after a network hiccup, will it double-process? Idempotency must be documented (pattern:idempotent-operation).