Efficient codebase navigation and semantic code search using Tree-sitter parsing and vector embeddings
Sourcerer MCP has 5 tools with clear naming (verb-first: semantic_search, find_similar_chunks, get_chunk_code, index_workspace, get_index_status). All tools have descriptions and input schemas are defined via the mark3labs/mcp-go framework. However, critical issues limit quality: (1) Parameter descriptions are minimal or missing, file_types in semantic_search has a 19-char description ('Filter by file type(s)'), below the 20-char floor for credit. (2) Output schemas are entirely undocumented, no LLM can infer what fields semantic_search or find_similar_chunks return, breaking the tool-chaining pattern. (3) Error handling is generic, CallToolResult returns plain text error strings with no recovery guidance (e.g., 'Search failed: %v'). (4) No input validation descriptions, semantic_search accepts free-form query strings with no constraints stated. (5) The instructions are well-written (detailed guidance on usage), but this is a prompt, not a schema remedy. Per-tool scores average 52 across 5 tools, with semantic_search and find_similar_chunks scoring lowest due to undocumented outputs.
Find code chunks semantically similar to a given chunk
Get the actual code you need to examine
Get the codebase's indexing status
Index all pending files in the workspace
Find relevant code using semantic search
Output schemas completely undocumented. semantic_search, find_similar_chunks, and get_chunk_code return text blobs with no structured schema definition. LLMs cannot infer response field names, types, or structure, breaking downstream tool composition.
Parameter descriptions are minimal or missing detail. 'Filter by file type(s)' (19 chars, below the 20-char floor) does not explain valid values (src, docs, tests are mentioned in instructions but not in schema). Missing description of expected format or enum constraints. No guidance on when to filter vs. search broadly.
Error handling is non-actionable. All errors return plain text like 'Search failed: %v' with no recovery guidance. LLMs cannot determine if the error is transient (retry) or user-fixable (change query).
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 54 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 41 | - | v1 |
No input validation constraints documented. semantic_search accepts any string query with no length, format, or character restrictions stated. index_workspace and get_index_status have no parameters but no documentation of preconditions (e.g., workspace must be initialized).
find_similar_chunks has no guidance on valid chunk ID format. Description says 'The chunk ID to find similar code for' but does not explain how to construct or recognize a valid ID (the instructions mention formats like 'path/to/file.ext::Type' but this is not in the parameter schema).