MCP server for searching a codebase using semantic similarity with embeddings and vector database
The server defines 3 tools with reasonable naming (all verb_noun style: search_codebase, search_by_file_type, get_codebase_stats). Descriptions are present and substantive (185-240 chars), explaining what each tool does and what it returns. However, critical gaps exist: parameter descriptions lack constraint details (e.g., 'limit' has no min/max enforcement stated in description), output schemas are not formally documented in the tool definitions themselves, and error handling returns generic error dicts rather than actionable recovery guidance. The search tools validate and cap limit to 5-20 at runtime but don't state this constraint in the parameter description text. search_by_file_type filters client-side after fetching, which is inefficient. get_codebase_stats lacks pagination guidance despite potentially returning large language dictionaries. No tool annotations (readOnlyHint, idempotentHint) are present despite all tools being READ_ONLY.
Get statistics about the indexed codebase. Returns: Dictionary with codebase statistics including file counts and index info
Search within files of a specific type/extension. Args: file_extension: File extension to filter by (e.g., "ts", "html", "scss") query: Search query limit: Maximum number of results (default: 5, max: 20) Returns: List of search results filtered by file type
Search the codebase using semantic similarity. Args: query: The search query (e.g., "Angular component", "authentication guard") limit: Maximum number of results to return (default: 5, max: 20) Returns: List of search results with file paths, similarity scores, and code snippets
Parameter descriptions lack constraint details. 'limit' parameter description does not state the enforced range (1-20) or that values >20 are capped. LLMs cannot infer numeric bounds from parameter names alone.
Output schemas are not formally documented in tool definitions. Callers must infer response structure from the code (SearchResult dataclass exists but is not exposed in the schema). search_codebase returns a list of dicts with 'file_path', 'language', 'similarity', 'chunk_index', 'content_preview', 'content_lines'; get_codebase_stats returns a dict with 'total_files', 'total_chunks', 'languages', 'last_indexed', 'average_chunks_per_file', but none of this is declared in the tool schemas.
Error handling returns generic error dicts without recovery guidance. When search fails, the tool returns [{'error': 'Search failed: ...'}]. An LLM receives no hint about whether the error is retryable, or what to do next. Should specify 'Try again', 'Check indexer connectivity', or 'Indexer not initialized, call setup first.'
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 61 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 52 | - | v1 |
No tool annotations. All tools are READ_ONLY and idempotent (safe to retry), but the server does not declare readOnlyHint or idempotentHint in the tool metadata. This forces the LLM to infer safety from descriptions.
search_by_file_type performs filtering client-side (retrieves 3x results, then filters by extension). This is inefficient and wasteful of embeddings computation. Should push the file extension filter into the database query.
No pagination guidance for get_codebase_stats. The 'languages' dict could be large (tool caps at 10 languages per query, but the response structure is not documented). If a codebase has 50+ languages, the agent cannot request a subset or paginate.