MCP server for semantic code search with hybrid search (BM25 + vector embeddings)
Code Sage has 4 tools with partial schema definitions visible in the source. All 4 tools have descriptions (good), but schema completeness varies. analyze_code, find_code, and delete_index have input schemas with type info and parameter descriptions. check_status also has a schema. However, NONE of the tools have documented OUTPUT schemas, a critical gap for composition and LLM planning. Tool naming is verb-noun (analyze_, find_, delete_, check_) which is correct. The descriptions are adequate in length (50-200 chars) but lack dependency hints and recovery guidance. Parameter descriptions are present but minimal. The risk annotations (WRITE, READ_ONLY, DESTRUCTIVE) are helpful but not part of the MCP tool definition spec, they appear to be internal metadata. Overall, the definitions show competence in basics but lack the depth and completeness for A-grade production use.
Create a searchable index of your code by analyzing functions, classes, and methods. This enables smart code search with natural language queries.
Check if code analysis is complete, in progress, or failed. Shows percentage done and number of files processed.
Delete the search index for a codebase to free up space or start fresh. Removes all stored code analysis.
Find code using natural language questions. Combines keyword search with AI understanding to locate relevant functions, classes, and code patterns.
No output schemas documented for any tool. LLMs cannot plan downstream calls or extract required fields. Critical for composition and context efficiency.
delete_index (destructive operation) lacks confirmation mechanism or dry-run mode. No recovery guidance in description. Agents could accidentally delete indexes without recovery path.
analyze_code accepts 'force' boolean with no default documented. If LLM omits it, unclear whether re-indexing happens. Defaults must not cause unintended side effects.
find_code limit parameter has no min/max constraint stated. LLM could pass limit=10000 or limit=1, breaking the intent of pagination support.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 55 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 44 | 2024-11-05+ | v1 |
Descriptions lack recovery hints. E.g., analyze_code should say 'Call check_status() to monitor progress' or find_code should say 'If results are incomplete, call analyze_code first.'
'path' parameter in all tools accepts string with minimal guidance. Should specify format: absolute path, permissions required (read/write), and what happens if path doesn't exist.