A productivity-focused MCP server providing semantic search over notes, URL summarization, and task management capabilities
This server has 4 tools with basic schema coverage and descriptions, but significant gaps in parameter documentation, error handling, and output schema clarity. All tools are explicitly registered via mcp.tool() decorators in server.py with visible function definitions. However, parameter descriptions are minimal or absent in several cases, output schemas are not formally documented, and error handling lacks actionable recovery guidance. The descriptions are adequate but terse (averaging ~50-70 chars), falling short of the 194-char production baseline. Input schemas exist for all tools but are incomplete, several parameters lack descriptions, and constraints are not formally expressed. No tool annotations (readOnlyHint/destructiveHint/idempotentHint) are present. Tools follow basic verb_noun naming but lack the operational clarity expected of production tools.
Create a new task and persist it to local JSON store.
List tasks, optionally filtered by status.
Semantic search over local markdown notes.
Fetch a URL and return a concise summary using Claude.
Output schemas not documented. Tools return complex dicts/lists but LLMs cannot see field definitions, types, or structure. search_notes returns {text, source, score} without schema declaration; create_task returns {id, title, description, status, created_at} without field typing; list_tasks returns arrays with no documented structure.
summarize_url lacks input schema documentation, no description for the 'url' parameter explaining format, validation, or constraints. LLM cannot infer whether relative URLs are valid or what URL formats are accepted.
No error handling or recovery guidance. Tools make external calls (HTTP fetch in summarize_url, FAISS index operations in search_notes, file I/O in create_task/list_tasks) but lack try-catch blocks, timeout guards, or actionable error messages. If summarize_url times out or fetch fails, the LLM receives an unhandled exception, not 'URL unreachable. Try a different URL or check network.' Pattern: recovery-guide.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 46 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 37 | - | v1 |
No tool annotations. create_task is destructive (writes to JSON), list_tasks and search_notes are read-only, summarize_url is read-only but calls external API. Absence of readOnlyHint/destructiveHint/idempotentHint prevents LLMs from understanding idempotency and retry safety.
search_notes uses stub embedding (hash-based seeding) instead of real embeddings. Docstring says 'placeholder, swap for a real embed model' but production code is live. This causes semantic search to be non-functional, identical queries produce different embeddings based on hash collisions, not meaning.
search_notes returns results with 'score' (L2 distance) without explaining what the score means or how to interpret it. Lower distance = better match, but LLM cannot infer this. No documentation of distance range or relevance threshold.
create_task description says 'Persist it to local JSON store' but does not mention that it returns the created task with an auto-generated UUID and timestamp. LLM cannot plan downstream calls that reference the task_id.
list_tasks has a 'status' enum (pending|done|all) but the description does not explain what these statuses mean or how they are set. No docstring mentions that only 'pending' and 'done' are actual states; 'all' is a filter mode.
search_notes default top_k=5 but no description of the limit or guidance on why 5 is chosen. Baseline production tools document min/max ranges for numeric parameters. No documentation of what happens if top_k > total chunks indexed.