Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
Files-DB-MCP has 5 tools with schemas and basic descriptions, but quality is inconsistent. Tool names follow verb_noun convention (vector_search, get_file_content, get_model_info, get_collection_stats, reindex_project). Descriptions are present but brief (15-75 chars), below the 50-200 char optimum for LLM selection. Parameters are typed and mostly described, but descriptions lack specificity about constraints, ranges, and when to use each tool. No error handling guidance visible. Output schemas are not documented. The tool definitions appear to be statically registered (from source inspection), but the actual MCP tool registration code is not visible in the provided excerpt, making it unclear if all schema details are properly exposed to clients.
Tool descriptions are too short (15-75 chars vs. 50-200 char optimum). LLMs cannot determine when to select each tool or how it differs from similar tools.
Parameter descriptions lack specificity. 'Filter by file type (e.g., python, javascript)' does not specify allowed values, format, or if any string is accepted. No enum constraint visible in schema.
No output schemas documented. LLMs cannot predict what fields to expect from vector_search results, get_file_content response, or get_collection_stats. This forces trial-and-error.
Expand each tool description to 50-200 characters, explicitly stating WHAT the tool does, WHEN to use it instead of alternatives, and what it returns. Example: 'vector_search: Search the vector database for files semantically matching a natural-language query. Use this to find related code files, documentation, or configs. Returns ranked results with file paths and similarity scores. Call this first to discover relevant files; then use get_file_content to read them in full.'
Add enum constraints to vector_search.file_type parameter. Document allowed file types (python, javascript, java, etc.) as an explicit enum rather than free-form text with examples.
Document output schemas for all tools. For vector_search, specify the result object structure: [{file_path: string, similarity_score: number, file_type: string, preview: string, ...}]. For get_file_content, specify max returned size and truncation behavior.
Add range constraints to numeric parameters. vector_search.limit: 1-100 (default 10), vector_search.threshold: 0.0-1.0 (default 0.6). Include these in the parameter description text so LLMs parse them before calling.
Add pagination support to vector_search: accept limit (1-50, default 20) and offset (default 0), return total_results_count and has_more fields. Cap default limit at 20 to avoid context overflow.
Document error recovery for each tool. Example for vector_search: 'If no results are returned, try lowering the threshold parameter. If threshold is 0.9, reduce it to 0.6 for broader matching. If results are still empty, the query may not match any indexed files.'
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
No error handling guidance. Tools have no documented recovery actions if file_path is invalid, query is empty, reindex fails, or similarity threshold yields no results. LLMs receive raw errors with no 'what to do next' hint.
vector_search parameters lack range constraints. limit has default=10 but no max documented. threshold default=0.6 with no min/max bounds. LLMs could pass threshold=999 or limit=1000000 and silently break the tool.
No pagination documented for vector_search. If a query matches hundreds of files, returning all results wastes context. Should support limit + offset/cursor and return total_count.
reindex_project is a destructive operation (rewrites the vector database) with no dry-run or confirmation step. An LLM could trigger full reindex accidentally. No documented rollback path.
Tool descriptions do not distinguish use cases. When should an agent call vector_search vs get_file_content? When should reindex_project be called, and what triggers it? No guidance for multi-step planning.
vector_searchget_file_contentreindex_project
Add a dry-run parameter to reindex_project: 'If dry_run=true, simulate reindexing and return what would be updated without modifying the database. Use this to preview changes before committing.' This prevents accidental data loss.
Distinguish tools in descriptions. Example: 'get_file_content is for reading entire files once you know the path (from vector_search results or user input). vector_search is for discovering files matching a query. Do not call get_file_content without a confirmed path.'
Add parameter dependency documentation. Example: 'file_extensions filters results only if you also provide file_type, or stands alone. If both are provided, results must match both constraints.'
Return chaining IDs in responses. If vector_search returns file_path results, ensure each result includes the full path and any metadata (file_type, size, modified_date) that downstream tools need, so get_file_content can be called immediately without discovery.