This MCP server has 15 tools with mixed quality. Strengths: tool names are action-verb based and generally clear (parse_*, create_*, analyze_*, search_*, replace_*, query_*). All tools have descriptions present. Schemas are visible and mostly complete with parameter types. Critical weaknesses: (1) Descriptions are frequently generic, short, and lack WHEN/WHY guidance that LLMs need for tool selection. (2) Many descriptions under 50 chars (e.g., 'Check status of the USS Agent services' for uss_agent_status). (3) No output schema documentation, LLMs cannot infer what fields to expect from return values. (4) No error handling guidance, tools fail silently without recovery instructions. (5) Parameter descriptions are minimal; many lack format/constraint details (e.g., 'Source code to parse' but no language enum). (6) No input validation or actionable error messages visible. (7) Critical semantic gap: tools like sync_file_to_graph and ask_uss_agent depend on external services (Neo4j, ChromaDB, LLM) but no dependency documentation or fallback behavior. (8) Parameters with ambiguous meanings: 'graph_path' in query_graph has no format spec; 'mode' has no enum but expects specific strings; 'limit' in semantic_search has no bounds. These gaps force LLMs to guess at tool semantics and output structure.
Create an enhanced Abstract Semantic Graph (ASG) from an AST. This version provides more complete edge detection, including improved scope handling and control flow edges.
parse_code_to_astread onlysource verified75/100
Parse code into an Abstract Syntax Tree (AST) using tree-sitter.
No output schemas documented for any tool. LLMs cannot infer what fields to expect from return values, forcing them to guess at downstream field names and structure. This causes broken tool chains and unnecessary discovery calls.
External service dependencies (Neo4j, ChromaDB, LLM) are not documented in tool descriptions. Tools will fail silently if services are down. No error handling guidance or fallback behavior provided.
Add comprehensive output schemas for all 15 tools. Document return structure (e.g., parse_code_to_ast returns {nodes: [{id, type, text, start_byte, end_byte, children?: []}], tree_height: int, error?: string}). Format as JSON Schema. This is CRITICAL for tool chaining.
Expand tool descriptions to 100-200 characters following pattern:tool-description. For each tool, answer: (1) What does it do? (2) When should I call this instead of similar tools? (3) What are key output fields? Example: 'Parse source code into an Abstract Syntax Tree (AST) using Tree-sitter. Returns tree nodes with type, location, and children. Use this before create_asg_from_ast. Supports Python, JavaScript, Go, Rust, C++, Java via language parameter.'
Add enum constraints for 'language' parameter across all tools. Document supported values: ['python', 'javascript', 'typescript', 'go', 'rust', 'c', 'cpp', 'java']. Remove language detection as fallback, require explicit specification to prevent silent failures.
Document external service dependencies in tool descriptions. Add sections: 'Requires: Neo4j ≥6.0 (via neo4j driver)', 'Requires: ChromaDB service running on CHROMADB_URI', 'Requires: LLM endpoint accessible at OPENAI_API_KEY'. Include error guidance: 'If Neo4j is unavailable, returns error 503 with message: Neo4j not responding on <host>:<port>. Check service status with uss_agent_status.'
Add error handling guidance to all tool descriptions. Distinguish error types: (1) Retryable (timeout, service unavailable) → 'Retry with exponential backoff'. (2) User-fixable (invalid language, bad schema) → 'Valid languages: python, javascript, go, rust, c, cpp, java'. (3) Fatal (permission denied, resource not found) → 'This tool requires Neo4j database; verify database is initialized with code graph data.' Example for parse_code_to_ast: 'On language detection failure: passes filename to Tree-sitter. If still unsupported, returns error: Unsupported language '<ext>'. Valid: .py, .js, .ts, .go, .rs, .c, .cpp, .java.'
Score history
Overall score trend
↑ 39 points across a rubric change (v1 → v2)
53/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
D
53
2026-07-28+
v2
2026-03-09
F
14
-
v1
read onlysource verified75/100
Parse code into an AST incrementally using Tree-sitter. This is an optimized version that can use a previous tree to only parse the changed parts of the code, which is much faster for large files with small changes.
query_graphread onlysource verified57/100
Query nodes from saved graph by type/relationships. Returns matching nodes. Modes: summary, node, traverse, query - summary: Get node/edge counts and type distributions - node: Get full details for node_id (include connected edges) - traverse: Follow edge_types from node_id up to depth - query: Filter nodes by node_type
Many tool descriptions are under 100 characters and lack WHEN/WHY context. This is a simplified version that extracts some basic semantic information.' (132 chars, vague) do not provide LLM-actionable guidance for tool selection.
No enum constraints on parameters with known value sets. Parameter 'language' appears in multiple tools but accepts free-form strings (no enum of supported languages like Python, JavaScript, Go, etc.). Parameter 'mode' in query_graph expects specific strings (summary, node, traverse, query) but is unconstrained. This invites hallucinated values and silent failures.
No error handling guidance in any tool. Tools fail without telling the LLM what to do next. No distinction between retryable errors (service timeout), user-fixable errors (invalid language), and fatal errors (bad schema). Per pattern:error-classification, errors must guide the agent's next step.
Conditional parameter dependencies are documented in descriptions but not structured constraints. E.g., parse_code_to_ast_incremental notes that 'old_code' is 'required if previous_tree is provided', easy for LLMs to miss. query_graph has mode-specific required parameters not enforced.
replace_pattern (WRITE risk) lacks confirmation or dry-run capability. LLMs can invoke irreversible string replacements on code without preview or user confirmation. Per pattern:confirmation-request, destructive operations should support a dry-run step.
Parameter descriptions are minimal and lack format constraints. E.g., 'graph_path' has no format spec (relative? absolute? file extension?), 'limit' has no bounds, 'pattern' in search_pattern includes inline example '$PROP && $PROP()' which LLMs may reuse literally.
No rate limits, timeouts, or result caps documented. query_neo4j_graph can return millions of rows without limit. semantic_search accepts no limit parameter in schema (limit is optional with no bounds). Per pattern:mxe-enforce-result-limits, tools should cap results at 20-50 items and document limits in descriptions.
query_neo4j_graph has no query validation or injection prevention. Cypher queries can be abused to drop databases or exfiltrate data. No access control gate documented. Per pattern:permission-gate and pattern:tool-gateway, destructive database queries must be validated and gated.
query_neo4j_graph
Replace inline examples with enum constraints. Remove pattern '$PROP && $PROP()' from search_pattern description. Instead, add to schema: 'pattern: {type: string, description: "ast-grep pattern (see ast-grep.github.io/guide/pattern-syntax for syntax; common: $PATTERN for wildcard, $EXPR for any expression)"}'.
Add bounds and constraints to numeric/string parameters. semantic_search: 'limit: {type: integer, minimum: 1, maximum: 100, default: 20, description: Maximum results to return}'. query_graph: 'depth: {type: integer, minimum: 1, maximum: 10, description: Max hops in graph traversal}'.
Add dry-run capability to replace_pattern. New parameter: 'dry_run: {type: boolean, default: true, description: If true, return matching replacements without applying changes. Set to false only after confirming results in dry-run mode.}' Align with pattern:confirmation-request.
Document conditional parameter dependencies explicitly in schema, not just descriptions. E.g., parse_code_to_ast_incremental: Note in schema that 'old_code' is required when 'previous_tree' is provided. query_graph: Add oneOf constraint: 'mode=summary requires no additional params, mode=node requires node_id, mode=traverse requires node_id + edge_types, mode=query requires node_type'.
Add result pagination and limits. query_neo4j_graph: 'limit: {type: integer, default: 50, maximum: 1000}', 'offset: {type: integer, default: 0}'. semantic_search already has limit but set maximum: 100. All list-returning tools should return {results: [...], total_count: int, has_more: bool, next_cursor?: string}.
Add query validation and access control documentation to query_neo4j_graph. Description: 'Execute read-only Cypher queries on the code graph. Query whitelist enforced: SELECT, MATCH, RETURN only. DDL (DROP, CREATE INDEX, DELETE) blocked. Queries exceeding 10s timeout return error: Query timeout after 10s.'
Clarify USS Agent architecture and API boundaries. Description for ask_uss_agent, uss_agent_status, semantic_search: 'USS Agent integrates Neo4j and ChromaDB for code search. ask_uss_agent accepts natural language and translates to Cypher; semantic_search translates to ChromaDB vector queries. semantic_search is faster for keyword/topic search; ask_uss_agent for complex graph patterns.'
Add field documentation to sync_file_to_graph output. Current: 'Returns {stored: {ast_id, asg_id, analysis_id}}'. Expand: 'Returns {stored: {ast_id: string (UUID), asg_id: string (UUID), analysis_id: string (UUID)}, file_path: string, timestamp: ISO8601, errors?: []}. Use ast_id as input to create_enhanced_asg_from_ast.'
Document uss_agent_status output structure and success criteria. Description: 'Returns status object: {neo4j: {connected: bool, latency_ms: int}, chromadb: {connected: bool, latency_ms: int}, llm: {available: bool, model: string}. All services must show connected=true for ask_uss_agent to function.'
Add logging and audit trail guidance (per pattern:audit-trail). Document: 'All tools log input parameters, execution duration, and result summary to MCP server stderr. Sensitive parameters (code content >10KB) are hashed, not logged verbatim. Tool calls are traceable by execution timestamp and tool name.'
Validate parameter types at tool entry. parse_code_to_ast 'include_children' is boolean but no validation that non-boolean values are rejected with clear error. All boolean/enum params need validation with actionable error messages, not silent coercion.