MCP server for ClaudeVN compute instances to call tools for task management, progress reporting, context retrieval, and work coordination with the serving component
Static source inference · medium confidence · evidence: Streamable HTTP
Current-spec patterns detected
Summary
This server implements 14 tools across workflow management, context retrieval, and issue tracking. While all tools have schemas and descriptions, quality is inconsistent. Naming follows verb_noun conventions (get_, report_, submit_, etc.), which is appropriate. However, descriptions are often generic (50-80 chars), lacking WHEN-to-use and dependency guidance. Schemas are present but minimal, most parameters lack detailed validation constraints, range limits, and error guidance. Several tools accept loosely-defined object types (test_results, characterization, context_types arrays) without documented structure. Error handling is not visible in provided code. Tool composition appears reasonable (single-concern tools like get_assignment, report_progress), but output schemas are not documented, making it unclear what fields downstream tools receive. The server runs over HTTP with SSE, which is acceptable, but the STDIO-only variant in stdio_server.py limits accessibility. No tool annotations (readOnlyHint, destructiveHint) are declared despite clear READ vs WRITE classifications visible in the tool metadata.
Missing output schemas. Tools return structured data but no response field definitions are documented. Downstream tools and LLMs cannot plan chained calls without knowing what fields are available (e.g., does get_assignment return branch, commits, estimated_effort?). This forces discovery-by-trial-and-error and wastes context windows.
Loose parameter types in complex objects. Tools like claudevn_report_progress (commits array), claudevn_request_review (test_results object), claudevn_submit_characterization (characterization object) accept object or array parameters with no documented structure. LLMs cannot infer the required keys or nested types.
Document output schemas for all tools. Example: 'Returns: {task_id: string, status: string, assigned_to: string, estimated_hours: number, branch_name: string, created_at: ISO8601}'. This enables LLM planning and prevents broken tool chains.
Expand parameter descriptions for complex objects. For test_results in request_review and report_progress, document the expected structure: '{passed: number, failed: number, skipped: number, duration_seconds: number}' or similar. Use structured examples in descriptions.
Add WHEN-to-use guidance to READ tool descriptions. Example for get_context: 'Call this after get_assignment to retrieve task files, git history, and related issues. Use context_types=["files", "history"] to fetch code; use context_types=["dependencies"] to identify blocking tasks.'
Declare error codes and recovery paths for WRITE tools. Example for report_progress: 'Returns 400 if task_id not found (call get_assignment first); 409 if status transition is invalid (valid transitions: started → in_progress → blocked|review_requested|completed); 500 if database fails (retry with exponential backoff).'
Add tool annotations to schemas. Mark read-only tools with 'readOnlyHint: true'; mark destructive tools like complete_task with 'destructiveHint: true' and 'idempotentHint: false'; mark idempotent tools like submit_decomposition with 'idempotentHint: true'. Example: {"name": "claudevn_complete_task", "description": "...", "inputSchema": {...}, "destructiveHint": true}
Add pagination parameters to discovery tools. get_context should accept limit (1-1000, default 50) and offset/cursor parameters. Document: 'Returns up to limit items; provide offset for next page. Total count returned so LLM knows if more results exist.'
Score history
Overall score trend
↑ 22 points across a rubric change (v1 → v2)
55/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
D
55
2026-07-28+
v2
2026-03-09
F
33
-
v1
claudevn_notify_conflict
writeauth50/100
Notify serving of a merge conflict encountered during work
Generic descriptions lacking WHEN-to-use and dependency guidance. 'Fetch persona definition' does not explain when to call this tool vs get_skill, or what a persona is in this system. Descriptions should answer: what, when, why, and prerequisites. Most descriptions are 50-80 chars; median production quality is 194 chars.
No error handling guidance. Tools declare no error codes, recovery paths, or actionable error messages. If a blocker signal fails, what should the agent retry? Is it a permission error, a network error, or a validation error? Silent failures risk agent loops.
No tool annotations. Tools are classified READ_ONLY vs WRITE in metadata but annotations (readOnlyHint, destructiveHint, idempotentHint) are not declared in MCP schemas. LLMs cannot determine idempotency or side effects from tool definitions alone.
Ambiguous context_types parameter in get_context. Accepts enum ['files', 'history', 'related_tasks', 'dependencies', 'all'] but no documentation of what structure each type returns or whether multiple types can be requested together.
No pagination support declared. Tools like get_context and get_skill may return large result sets but no limit, offset, or cursor parameters are visible. Large unbounded results risk context window exhaustion.
Field naming consistency. Some parameters use task_id, others use goal_id, others use issue_id. If decomposition returns a goal_id, but submit_challenge expects task_id, the LLM must infer the mapping. Field naming should be consistent across related tools.
Standardize ID naming across related tools. If decomposition returns decomposition_id and goal_id, ensure submit_decomposition accepts the same field names. Document relationships in descriptions: 'decomposition_id: string, returned by previous decomposition call; goal_id: string, from goal this decomposition targets'.
Add constraint descriptions to enum fields. For status in report_progress, describe: 'Current task status. Valid transitions: started → in_progress (task work begins) → blocked (external dependency) | review_requested (ready for merge) → completed (merged and live).' This guides LLM state machine reasoning.
Document the persona/skill distinction. get_persona returns 'CLAUDE.md content', what does this mean? Is it a prompt? A behavior profile? A capability list? Add: 'Returns the persona definition (behavior guidelines and capability declaration) from CLAUDE.md. Use this to understand AI agent capabilities before assigning skill-specific work.'
Add batch variants for frequently-looped tools. If agents often add multiple issues in a loop, offer add_issues accepting an array instead of requiring N sequential calls. Current add_issues already supports this, ensure documentation states: 'Accepts 1-100 issues per call; batch for efficiency.'
Remove ambiguous optional parameters or make them conditional. notify_conflict accepts context as optional string, when is it needed? Default? Max length? Add: 'context: optional string (max 2000 chars), additional context about merge conflict resolution. Required if multiple conflicts in same file.'
Add rate-limit and concurrency guidance. Tools that modify state (report_progress, add_issues, submit_decomposition) should document: 'Rate limit: 100 requests/min per task_id. Concurrent calls to the same task may cause version conflicts, call sequentially or expect 409 Conflict responses. If 409 received, fetch current state and retry.'
Document field inclusion in responses. State: 'All responses include task_id, timestamp, and status. Responses for state-changing operations (report_progress, signal_blocker) also return updated_by (agent ID) and next_allowed_transitions (array of valid status strings).' This prevents LLM assumptions about missing fields.