Self-hosted AI memory server that captures, embeds, and semantically searches unified memory for AI agents. Runs on Raspberry Pi NAS and exposes tools via MCP protocol for Claude Code, Codex CLI, and other AI clients.
STRATA has 10 tools with complete input schemas visible in server.py. All tools have descriptions and most parameters are typed and described. However, several issues prevent a higher score: (1) No output schemas are documented, responses are not formally specified, forcing LLMs to infer structure. (2) Tool descriptions lack 'WHEN to use this' guidance and actionable recovery hints. (3) No error handling guidance is visible, no indication of retryable vs fatal failures, no sample error messages. (4) Destructive tools (delete_thought) lack confirmation/dry-run patterns. (5) Writing tools (capture_thought, update_thought) lack idempotency guarantees. (6) No pagination guidance for list operations despite returning arrays. Average description length ~140 chars (below baseline 194), and parameters average 3.5 per tool (near baseline 4). Tools follow verb_noun naming (capture_, semantic_search, update_, delete_, list_, get_) which is solid, but descriptions are sparse on context.
Store a thought/memory in the unified knowledge base. Automatically generates semantic embeddings and extracts tags.
Permanently delete a thought from the knowledge base.
Retrieve a single thought by ID.
Check server status and connectivity to the database.
Search thoughts by keyword/tag matching. Returns exact matches and tag-based results.
Get all people mentioned in thoughts, with frequency counts.
Retrieve recently captured thoughts, ordered by timestamp (newest first).
No output schemas documented. LLMs cannot predict response structure or plan downstream tool calls. E.g., does capture_thought return thought_id, full object, or just acknowledgment? Does list_recent return array metadata (total_count, next_cursor)?
No error handling guidance or recovery hints. Tools define no sample error messages or categorization (retryable, user-fixable, fatal). E.g., if get_thought returns 404 for missing ID, should LLM call list_recent? If semantic_search times out, is it safe to retry?
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | B | 70 | 2026-07-28+ | v2 |
Get all tags currently in use across the knowledge base, with frequency counts.
Search thoughts by semantic similarity. Returns thoughts ranked by relevance to the query.
Modify an existing thought's content, tags, people, or task status.
Destructive tool (delete_thought) lacks confirmation or dry-run. No indication that deletion is irreversible or that agents should confirm before execution. Pattern: confirmation-request.
Write tools (capture_thought, update_thought) lack idempotency guarantees. If agent retries after a network timeout, will duplicate thoughts be created? Will repeated updates be safe? No mention of idempotent keys or deduplication.
List operations (list_recent, list_tags, list_people) lack pagination metadata. Do they return a total_count? A next_cursor? If semantic_search returns 'limit: 10', are there more results? LLMs cannot detect truncation or plan pagination without this.
Tool descriptions lack 'WHEN to use' guidance. 'Store a thought/memory' doesn't explain when to call capture_thought vs update_thought. 'Search thoughts by semantic similarity' doesn't contrast it with keyword_search. LLMs struggle to choose between similar tools.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). MCP 2026-07-28 spec supports these as optional tool metadata. Missing annotations leave LLMs guessing about side effects (critical for planning and safety).
Parameter descriptions are minimal. 'The thought text to store' is clear but lacks constraints: max length? Format restrictions? 'Optional user-provided tags' doesn't explain: what happens if user tags conflict with auto-extracted tags? What's the max tag count?
Priority parameter (0-5) in capture_thought and update_thought lacks enum constraint. Should be constrained to discrete values or have a description like '0=lowest, 5=highest' for clarity.