Exposes Corbell's architecture graph, code embeddings, and spec tools as MCP tools for external AI platforms (Cursor, Claude Desktop, Antigravity).
Corbell MCP server has 4 tools with adequate naming and moderate descriptions, but suffers from incomplete parameter documentation, missing output schemas, and insufficient error handling guidance. All tools are read-only, which is positive for safety, but the lack of structured output definitions and parameter constraints significantly reduces usability for LLMs. Schema documentation is minimal, parameters exist but lack type constraints and validation rules. Error handling returns generic error strings rather than actionable recovery guidance. This places the server in the C/D range (fair-to-poor) on the calibration scale.
Semantic search across Corbell's code embedding index. Returns the most relevant code chunks matching the query, ranked by cosine similarity. Useful for finding implementations, patterns, and code examples across the workspace.
Get architecture and code context for a feature without LLM generation.
Query Corbell's architecture graph for service dependencies and details.
List all services in the current Corbell workspace graph. Returns a summary of every service discovered by `corbell graph build`, including language, type, tags, and dependency count.
No output schemas documented for any tool. LLMs cannot plan downstream operations or extract specific fields from responses.
Parameter descriptions lack format constraints, ranges, and validation rules. E.g. 'top_k' in code_search has no minimum/maximum bounds; 'service_id' has no explanation of valid format.
Error handling returns generic error strings ('Error querying graph: {str(e)}') without recovery guidance. LLMs cannot determine if errors are retryable, user-fixable, or fatal.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 56 | 2026-07-28+ | v2 |
No pagination support or result limits documented for list_services and code_search. If these return large datasets, LLMs risk context window exhaustion.
Tool descriptions are present but lack WHEN to use guidance. E.g. graph_query vs get_architecture_context are ambiguous, when should an LLM choose one over the other?