CLI tool for vectorizing codebases and serving them via MCP
project-vectorizer defines 4 tools with explicit schemas and descriptions. All tools have input schemas visible in the code (JSON Schema with types), descriptions present, and are READ_ONLY. However, the server exhibits several quality gaps: (1) descriptions are minimal, averaging ~25 chars, well below the 194-char baseline and insufficient to guide LLM selection; (2) parameter descriptions are sparse or missing entirely (e.g., 'file_type' in list_files has no description despite being optional); (3) no output schemas are documented, tools return JSON strings but the structure (fields, types, nesting) is not specified; (4) no error handling guidance is provided, tools catch exceptions and return bare error strings ('Error: {e}') without telling the LLM what to do next; (5) no parameter constraints are enforced (e.g., 'limit' and 'threshold' accept any value with no validation); (6) the tool chain is weak, search_code returns results with minimal context for follow-up calls. The naming is clear and verb-based (search_, get_, list_), and all tools are safely READ_ONLY, but overall the definitions lack the depth needed for confident LLM reasoning at scale.
Retrieve file content.
Get project statistics.
List all files in the project.
Search through the vectorized codebase.
Missing output schema documentation. All tools return JSON strings but no structured schema is documented, LLMs cannot plan downstream calls or extract fields reliably.
Descriptions are too brief (avg. 25 chars vs. baseline 194). 'Search through the vectorized codebase' and 'Retrieve file content' lack context on when to use each tool vs. similar ones, and do not explain return structure or prerequisites.
Parameter descriptions missing or insufficient. 'file_type' in list_files has no description. 'limit', 'threshold', 'query', 'file_path' lack constraints (e.g., range, format, enum). LLMs cannot validate inputs before calling.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 48 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 43 | - | v1 |
No error handling or recovery guidance. Tools return 'Error: {e}' with no actionable next steps. When search_code finds 0 results or get_file_content fails, the LLM receives no recovery hint (e.g., 'Try broadening the query' or 'File not found. Call list_files() to discover available paths').
Weak tool composition and chaining. search_code returns minimal metadata (query, total_results, results array) but does not include result structure (e.g., file_path, snippet, score, line_number). LLM must infer what to extract or call get_file_content blindly. Missing chaining IDs.
No input validation or constraint enforcement. 'limit' and 'threshold' accept any integer/float with no bounds check. LLMs could request 1M results or threshold of 999.0, causing server overload or nonsensical behavior.