A Model Context Protocol server for VectorCode, a tool to vectorise repositories for RAG (Retrieval-Augmented Generation)
VectorCode MCP Server provides 5 tools with mostly complete schemas but suffers from weak descriptions, missing parameter-level documentation, and unclear guidance for LLM tool selection. All tools are properly registered with input schemas visible in src/vectorcode/mcp_main.py. Tool names follow verb_noun convention (list_collections, vectorise_files, query_tool, ls_files, rm_files), which is good. However, descriptions are brief (27-156 chars, below the 194-char production baseline) and lack context about when to use each tool relative to others. Most critically, parameter descriptions are present but generic, they state WHAT parameters are but not WHEN or WHY to use them, limiting LLM reasoning. Error handling is minimal, with no recovery guidance or actionable error messages visible. Schemas are well-formed with proper typing, but output schemas are not documented.
Retrieve a list of projects accessible via the VectorCode tools
Retrieve a list of files that have been added to the database for a given project
Query the VectorCode database for relevant files based on keywords
Remove files from the VectorCode database. The files will remain in the file system.
Vectorise files in a project so that they'll be available from the query tool. The paths should be accurate and case sensitive.
Tool descriptions are too brief (27-156 chars vs production baseline of 194 chars) and lack context for LLM selection. Descriptions state WHAT the tool does but not WHEN to use it or how it differs from similar tools (e.g., query_tool vs ls_files vs vectorise_files).
Parameter descriptions exist but are often generic and incomplete. For example, 'project_root' is described as 'Directory to the repository' in 3 different tools but never explains: What happens if I omit it? What is the fallback behavior? Is null accepted?
No output schemas documented. Tools return results (file lists, vector matches, etc.) but the response structure, field types, and data models are not visible in the tool definitions. LLMs cannot plan downstream calls without knowing what fields to expect.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 61 | <=2025-11-25 | v2 |
| 2026-03-09 | C | 62 | - | v1 |
No error handling or recovery guidance visible. Tools do not document what errors are possible, how to interpret them, or what the LLM should do if a call fails (retry, ask user, use fallback tool).
Tool 'query_tool' has ambiguous naming. It does not clearly distinguish whether it searches vectorized file contents, file metadata, or both. Compare to alternatives: 'search_files', 'find_relevant_files', 'query_vectors'. The generic name 'query_tool' invites confusion with other query operations.
Parameter 'n_query' in query_tool lacks constraints. No minimum, maximum, or guidance on valid range. An LLM could pass 0, 1000000, or negative values. Should specify bounds (e.g., 1 - 100) in both schema and description.
Tool descriptions do not clarify dependencies or call order. E.g., should list_collections be called first to discover valid project roots? Should vectorise_files be called before query_tool? Implicit workflows force LLMs to guess.