A personal AI agent with persistent memory and offline cognition — River Algorithm
This MCP server demonstrates critical deficiencies across naming, descriptions, parameter validation, and error handling. While 8 tools are exposed via HTTP/FastAPI, only 2 have properly structured input schemas. Most parameter descriptions are trivial single-word phrases. No output schemas are documented. Error handling is absent, tools return generic HTTP exceptions without guidance for LLM recovery. The server conflates multiple concerns in single tools (e.g., dispatch_task with 'preview' or other actions is unclear). No tool annotations (readOnlyHint, destructiveHint, idempotentHint) are present despite destructive operations like delete_tool. Security is severely compromised: toggle_tool and delete_tool require admin auth but lack per-tool permission granularity, and no audit trail is evident. The implementation shows Flask/FastAPI confusion in requirements.txt (both listed) and uses stateful in-memory operations (_revert_ops, _pending_restart) that violate stateless MCP protocol design.
Delete a tool from the tool registry
Dispatch an outsource task with preview action
Query financial information
Query health-related information
Analyze and describe images
Toggle tool enabled/disabled status in settings
Transcribe voice/audio input
Search the web for information
Tool descriptions are trivial and non-informative. 6 of 8 tools have descriptions under 35 characters, violating the baseline of 194 chars (median). Many read like auto-generated stubs (e.g., 'Toggle tool enabled/disabled status in settings', 'Dispatch an outsource task with preview action'). These fail to explain WHEN to call the tool or distinguish it from similar tools, making LLM selection unreliable.
Input schemas are incomplete or missing. Only toggle_tool and delete_tool have visible schemas; the others lack type definitions and constraints. Parameter descriptions are single words or circular (e.g., 'Task description' for 'task', 'Search query' for 'query'). No minimum/maximum ranges, enums, regex patterns, or format guidance. LLMs will hallucinate valid inputs.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 37 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Output schemas are entirely absent. No tool documents what fields it returns, their types, or structure. Agents cannot plan downstream tool chains or extract data reliably. E.g., web_search should return structured results (rank, title, URL, snippet) and indicate pagination; currently the agent must parse unstructured text.
Destructive/sensitive tools lack security annotations and proper error handling. delete_tool performs irreversible deletion without confirmation, confirmation request pattern, or dry-run option. toggle_tool modifies settings files without per-tool permission checks. No audit trail of who called what or when. Missing destructiveHint and scope declarations.
Error handling is non-existent. Tools raise generic HTTPException with minimal messages (e.g., 'Failed to update tool', 'Tool not found in registry'). No guidance for LLM recovery: should it retry? Call a different tool? Ask the user? No error classification or actionable remediation.
Tool annotations (readOnlyHint, destructiveHint, idempotentHint) are absent. Without them, agents cannot reason about side effects, retryability, or safety. delete_tool should be marked destructive. web_search, image_describe, finance_query, health_query should be marked read-only. This is a protocol requirement for spec alignment.
Tool naming is inconsistent or vague. 'toggle_tool' does not clarify enable/disable semantics. 'dispatch_task' conflates 'preview' and 'execute'. 'finance_query' and 'health_query' are generic and do not indicate data sources or scope. These fail the 'verb_noun' convention and do not disambiguate when similar tools exist.
Implementation uses stateful patterns (in-memory _revert_ops, _pending_restart flags) that violate MCP's stateless design. Each request should be self-contained. The server maintains state across requests, risking race conditions and making the server non-idempotent. This conflicts with current MCP spec (2026-07-28).
health_query poses a critical safety risk: LLM-generated health advice without expert review or disclaimers could harm users. The tool lacks any indication of evidence quality, medical accuracy, or requirement to consult healthcare providers. This should either be removed or heavily restricted with guardrails and disclaimers.