Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
The server defines 6 tools with reasonable naming (verb-noun format) and valid JSON Schema inputs, but suffers from critical gaps in parameter descriptions, output schema documentation, and error handling guidance. All tools have descriptions but they are sparse (10-50 chars, below the 194-char baseline). Input schemas are present but 4 of 6 tools lack descriptions for most parameters. No output schemas are documented. The server provides basic functionality but falls short of production-grade definition quality.
All tool descriptions are under 40 characters; baseline is 194 chars. Descriptions lack critical context: WHEN to use the tool, WHAT data is returned, and what distinguishes similar tools (e.g., search_code vs list_tasks). LLMs cannot effectively select tools without fuller context.
No output schemas documented for any tool. Callers (including LLMs) cannot know what fields to expect in responses. For example, search_code returns 'results' with what structure? Does list_tasks return a total count for pagination? Does create_task return the new task_id? This forces agents to guess and handle failures blindly.
Expand all tool descriptions to 100-300 characters. Include: (1) What the tool does, (2) When to use it instead of similar tools, (3) What data it returns in summary. Example: 'Search codebase using semantic similarity. Call this to find functions, classes, or code patterns matching a natural language query. Returns matching code snippets with file paths and line numbers. Use search_code for semantic queries; use grep for exact text matches.'
Document output schemas for all tools. For example: search_code should document 'Returns {results: [{file_path: string, snippet: string, line_number: integer, score: float}], total_count: integer}'. This enables LLMs to plan downstream tool chaining.
Add error handling sections to each tool description. Example: 'Errors: If repository_id is invalid, returns 404 with available repo IDs. If query is empty, returns 400. If index is stale, recommend calling index_repository(force_reindex=true).'
Clarify parameter constraints in descriptions. For create_task and update_task, document: 'description (optional, string, 0-5000 characters). Markdown supported. If omitted on update, existing description is retained.' Add similar guidance to notes and planning_references.
Add a 'Pagination' section to list_tasks description: 'Returns {tasks: [...], total_count: integer, limit: integer}. To fetch all tasks, iterate with limit=100 until len(tasks) < limit. Supports filters: status, branch.'
Restore tool chain documentation. After index_repository, document the response includes 'repository_id' (UUID string). Then in search_code, reference: 'If you don't have a repository_id, call index_repository() first.'
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Score history
Overall score trend
↓ 1 points across a rubric change (v1 → v2)
47/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
F
47
<=2025-11-25
v2
2026-03-09
F
48
-
v1
No error handling guidance provided. Tools do not document what errors can occur, how to recover, or whether failures are retryable. For example, index_repository may fail on permission denied, invalid paths, or missing directories, but the definition provides no guidance on recovery actions.
Parameter descriptions are inconsistent and sometimes missing detail. 'description' and 'notes' fields in create_task and update_task lack length constraints or format guidance. 'planning_references' is defined as 'Relative file paths' but should specify format (Unix paths? Glob patterns? Relative to what root?).
Parameter naming inconsistencies invite errors. create_task has 'title' and 'description'; update_task has the same fields. But the relationship is unclear: are both required in update? Can you update only one? This ambiguity forces LLMs to reason about optionality. Explicitly state 'at least one of title|description|notes must be provided' or similar.
No pagination/limit guidance for list_tasks. The tool accepts a 'limit' parameter (default 20, max 100), but the response structure is unknown. Does it return a total_count? A next_cursor? If there are 5000 tasks and limit is 100, how does the agent fetch all? This forces blind iteration.
search_code references 'repository_id' as UUID but does not explain how to discover repository IDs. Agents must call index_repository first to get a repo_id, but index_repository's response is not documented. This breaks the tool chain guidance.
No idempotence guarantees. Can create_task be called twice with the same title and get the same task_id, or does it create a duplicate? Can update_task safely be retried? Agents need to know whether tools are idempotent for safe retry logic.
create_taskupdate_task
Add tool annotations (destructiveHint, idempotentHint, readOnlyHint) via the MCP schema extension. Mark create_task and update_task with destructiveHint=true (they modify state). Mark get_task, list_tasks, search_code with readOnlyHint=true. Mark index_repository with an explicit idempotent=false warning or confirm retry behavior.
For update_task, document mutually exclusive vs coexistent parameters. Example: 'Either title, description, notes, or status can be updated independently. Omitted fields are not modified. At least one field must be provided.' This prevents unnecessary 'update everything' logic in agents.
Add examples to enum parameters. For status in list_tasks and update_task, explicitly list: status: 'need to be done' | 'in-progress' | 'complete'. Then note: 'Use enums exactly as shown; no abbreviations or variations.' LLMs frequently hallucinate enum values.
Document commit format clearly in update_task: 'commit (optional, string). Must be a valid 40-character SHA-1 hash (lowercase hex). Example: "a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b". If provided, task will be associated with this git commit.' This prevents LLMs from passing partial or invalid hashes.