Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
Server implements 6 tools for graph database access with acceptable naming and basic documentation. However, parameter descriptions are sparse or missing entirely for many inputs, output schemas are not explicitly documented, and error handling lacks recovery guidance. Tools follow a verb_noun pattern (get_*, execute_*, search_*) which is good, but parameter-level documentation is insufficient. Most tools are read-only and relatively safe, but the semantic search and embedding tools lack clarity on their dependencies and failure modes.
Parameter descriptions are minimal or absent for critical inputs. 'entity_id' and 'query' parameters lack format guidance, constraints, and error context. LLMs cannot infer valid formats without explicit description text.
Output schemas are not documented for any tool. Callers cannot predict field names, types, or structure. For example, get_project_info returns a dict but fields like 'error', 'name', 'path', 'summary' are not formally declared. search_nodes_for_semantic_similarity returns a dict with 'error' or 'results' but the structure of result items is undocumented.
Error handling lacks recovery guidance. Tools return generic error strings (e.g., 'Error: Node not found', 'Error: Could not retrieve source code: {e}') without actionable suggestions. LLMs cannot determine whether to retry, call a different tool, or ask the user.
Recommendations
Add explicit output schema declarations to each tool. Use structured JSON with typed fields (e.g., { 'id': string, 'source_code': string, 'error'?: string } for get_source_code_by_id). Document whether fields are always present or conditional.
Expand parameter descriptions to 50-150 characters, following the format: '[WHAT] The {param_name} [FORMAT] (e.g., a Cypher query matching MATCH (n:Method) [CONSTRAINT] without CREATE/DELETE [WHEN] for querying method signatures).' Example: 'entity_id (string): The unique identifier of the node to retrieve (e.g., com.example.MyClass.method1, acquired from previous search results).'
Wrap error messages with recovery guidance. Instead of 'Error: Node not found', return 'Entity not found. Verify the entity_id is correct by calling search_nodes_for_semantic_similarity() with a keyword related to the entity you are looking for.' This guides LLM recovery.
Document embedding prerequisite in get_graph_schema and add a note to generate_embeddings and search_nodes_for_semantic_similarity: 'Note: This tool requires embeddings to be pre-generated and indexed in the graph (summaryEmbeddings index). If results are empty, the graph may not have embeddings loaded.'
Implement result truncation and pagination for search_nodes_for_semantic_similarity and get_source_code_by_id. For source code, cap output at 5000 characters and return a 'truncated: bool' flag. For semantic search, document that results are capped at num_results (default 5) and include a 'total_available: int' field.
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
get_source_code_by_id and search_nodes_for_semantic_similarity both return 'source_code' and semantic 'results' as free-text strings/dicts. Large source files will be returned verbatim, wasting tokens. No pagination, truncation, or streaming. Results for semantic search are uncapped and could return 5+ unstructured items without a clear schema.
execute_cypher_query implements basic regex-based write-operation filtering (checking for CREATE, SET, DELETE keywords), but this is fragile and incomplete. Injection attacks via comments (e.g., '-- MERGE') or encoded payloads could bypass the check. Sanitization should be more robust.
generate_embeddings and search_nodes_for_semantic_similarity have a hidden dependency: they assume embeddings have already been generated and indexed in the graph. The tool descriptions do not document this prerequisite or what happens if embeddings are missing. An LLM could call these tools unprepared.
Replace regex-based query validation in execute_cypher_query with a proper Cypher parser or whitelist. Use a library like 'neo4j-python-driver' query validation or pre-compile allowed query patterns.
Add enum constraint to execute_cypher_query's 'query' parameter description to hint at safe operations: 'A read-only Cypher query. Permitted keywords: MATCH, OPTIONAL MATCH, WHERE, RETURN, UNWIND, CALL, WITH. Write operations (CREATE, SET, DELETE, MERGE) are blocked.'
Clarify the 'num_results' default in search_nodes_for_semantic_similarity to help LLMs understand pagination: 'num_results (integer, 1-100, default 5): Maximum number of nodes to return. If you need more results, make multiple calls with different query terms or increase num_results.'
For get_source_code_by_id, add a 'strip_comments' or 'summary_only' parameter option so LLMs can request condensed output for large files, reducing token waste.
Document which fields in responses are guaranteed vs. optional. E.g., get_project_info always returns 'name', 'path', 'summary', but may return an 'error' field if the query fails. Make this distinction explicit in the output schema.