MCP server for semantic code search and indexing using vector embeddings
rag-code-mcp has 11 code analysis tools with descriptions and parameters, but exhibits significant quality gaps. All tools are read-only except index_workspace (WRITE risk). Descriptions are present but variable in quality (some are detailed, others are generic). Input schemas are visible with type information, but parameter descriptions lack depth and clarity. Output schemas are undocumented, responses are returned as strings (JSON or markdown) with no formal schema definition. Error handling is minimal; most tools lack actionable error guidance. The server is STDIO-only, which is a hard transport cap. Tools follow a consistent naming pattern (verb_noun), but composition is loose, overlapping search tools (search_code, search_local_index, search_docs, hybrid_search) duplicate functionality without clear guidance on when to use each. No tool annotations (readOnlyHint/destructiveHint), no Multi Round-Trip Requests, no per-request _meta logLevel. The codebase shows testing discipline and some validation (TestGetCodeContextTool_Validation), but tool definitions appear inferred from code structure rather than explicitly registered.
Find the definition of a type (struct, class, interface, etc.)
Find usages of a symbol (function, type, variable) in the codebase
Get call hierarchy for a function showing what it calls and what calls it
Extract code context from a file - returns specific lines with surrounding context
Get detailed information about a function including signature, docstring, and source code
Hybrid search combining semantic and lexical matching
Index/reindex the codebase for search - USUALLY AUTOMATIC on first search. Call manually only if search returns 'workspace not indexed' or after major code changes (git pull, branch switch). Analyzes Go, PHP, Python, HTML files and stores vectors for semantic search.
HARD CAP: STDIO transport prevents remote accessibility and integration with hosted MCP clients
Output schemas are undocumented, tools return string responses (JSON or markdown) with no formal schema definition. LLMs cannot infer downstream tool compatibility or required fields for chaining.
Four overlapping search tools (search_code, search_local_index, search_docs, hybrid_search) have ambiguous selection criteria. Documentation states 'USE THIS FIRST' (search_code) and 'use hybrid_search only when...' but guidance is unclear. LLMs will struggle to choose the right one.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 33 | - | v1 |
List exported symbols (functions, types, classes) from a package
Semantic code search - finds functions, classes, and methods by MEANING, not just keywords. USE THIS FIRST when exploring unfamiliar code. Returns complete source code with file path and line numbers. Better than hybrid_search for general exploration; use hybrid_search only when you need EXACT identifier matches. Supports Go, PHP, Python, HTML.
Search documentation files using vector similarity
Search the local vector index for code snippets
Generic parameter descriptions lack specificity on format, constraints, or examples. 'Search query' and 'Path to a file in the workspace' are vague. No info on file path format (absolute? relative?), query length limits, result limits (default 5 for search_code, but what about hybrid_search?), or encoding.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). The server correctly identifies index_workspace as WRITE and others as READ_ONLY, but this is not exposed via MCP annotations. Modern agents rely on these hints to plan execution order.
Error handling is minimal. Code shows validation (e.g. TestGetCodeContextTool_Validation checks for missing file_path, start_line, end_line), but returned errors to the LLM likely lack recovery guidance. No evidence of actionable error messages like 'User not found. Try search_users() with a partial name.'
index_workspace description mentions 'USUALLY AUTOMATIC on first search' but does not explain idempotency or side effects. If called twice, does it recreate collections? The 'recreate' parameter suggests destructive behavior but lacks clear consequences. Agents need to know whether this is safe to retry.
Parameters like 'language' (Optional: specific language to index), 'limit' (Optional: maximum number of results), and 'output_format' (Optional: 'json' or 'markdown') lack clear defaults. What is the default output_format? Default limit? Default language if not specified?
Tools like search_local_index and search_docs check for configured memory/provider and return fallback error messages (e.g. 'no long-term memories configured'). While this shows defensive programming, the error messages in tests suggest the tool fails gracefully, but error categorization (retryable vs user-fixable vs fatal) is unclear.