Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
This server has significant quality gaps across definition quality dimensions. Parameter schemas are present but descriptions are largely absent or trivial. No output schemas are documented. Error handling is basic with no recovery guidance. The tool names follow verb-noun patterns reasonably well, but the underlying definitions lack the rigor required for production agent use. Most critically, parameter descriptions are missing or single-word, violating the core requirement that every parameter needs a description explaining what it controls.
Descriptions critically short and unhelpful. 'Upload documents with comprehensive error handling' (7 words) is generic and does not explain WHEN to use this tool vs related tools or what specific error handling is provided.
Parameter descriptions missing or trivial. 'file_paths' has description 'List of file paths to process' which is redundant with the type hint and does not explain format constraints, validation rules, or error behavior. 'limit' and 'top_k' parameters appear to have no descriptions at all in visible code.
No output schemas documented. Tools return complex nested objects (e.g. upload_documents_with_rag returns {'processed_documents': [...], 'failed_documents': [...], 'rag_summary': {...}}) but the MCP definitions do not specify the structure, field names, or types that downstream tools or agents should expect. This violates pattern:tool and pattern:response-shaper.
Recommendations
Rewrite tool descriptions to 50 - 200 characters, explaining WHAT the tool does, WHEN to use it, and any prerequisites. Example: 'Upload documents (PDF, Word, Excel, PPT) up to 100MB each and extract text into a vector database. Use this before search_documents or analyze_documents. Files are split into 512-token chunks with 128-token overlap.' This tells the agent when to call the tool and what to expect.
Add description to every parameter. For 'file_paths', write: 'List of file paths (absolute or relative) to process. Supported formats: PDF, DOCX, XLSX, PPTX. Max file size 100MB each. Returns error with list of unsupported formats if any fail.' For 'limit' and 'top_k', write: 'Number of results to return (1-100, default 10). Higher values return more context but increase token cost.'
Document all output schemas as JSON Schema in the MCP tool definition. Example for upload_documents_with_rag: {'type': 'object', 'properties': {'processed_documents': {'type': 'array', 'items': {'type': 'object', 'properties': {'doc_id': {'type': 'string'}, 'filename': {'type': 'string'}, ...}}}, 'failed_documents': {...}, 'rag_summary': {...}}}. This lets agents and downstream tools understand what to expect.
Add specific error recovery guidance. Instead of logging 'File too large', return {'error': 'File too large', 'max_size_mb': 100, 'suggestion': 'Split file into smaller documents or use a different file format. Call split_document_by_pages(file_path, max_pages=50) to create smaller chunks.'}. This guides the agent's next action.
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
No error recovery guidance. Error handling code in upload_documents_with_rag logs failures but does not return actionable guidance for the agent. Per pattern:recovery-guide, errors should tell the LLM 'what to do next', e.g., if a file is too large, suggest splitting it; if content extraction fails, suggest the file type is unsupported.
clear_document_store is destructive with no confirmation step. Per pattern:confirmation-request, irreversible operations should support a dry-run or confirmation. An agent might call this by mistake and irreversibly delete the entire RAG store. No user confirmation is solicited.
Ambiguous tool separation. upload_documents_with_rag and analyze_document_collection both operate on the document collection. The naming does not clearly distinguish when to call one vs the other. Per pattern:tool, each tool should have one clear responsibility, and similar tools should have unambiguous names.
No parameter constraints or range documentation. 'top_k' and 'limit' parameters accept integers but have no documented bounds (min 1, max value). Unbounded values let LLMs pass absurd requests (limit=999999) that could cause timeouts or resource exhaustion.
Config.MAX_FILE_SIZE_MB is not documented in tool description. Agents do not know the file size limit in advance, so they may repeatedly try to upload oversized files without knowing why they fail. The error message in code says 'File too large. Maximum size: {Config.MAX_FILE_SIZE_MB}MB' but agents cannot know this constraint until after a failure.
upload_documents_with_rag
Add a dry-run flag to clear_document_store: clear_document_store(dry_run=true) returns what would be deleted without deleting it. Then add a confirmation parameter that requires the string 'DELETE_ALL_DOCUMENTS' to actually execute. This prevents accidental data loss.
Rename or add description to disambiguate analyze_document_collection from upload_documents_with_rag. Consider: analyze_document_collection → get_document_collection_statistics or describe_uploaded_documents. Then add description: 'Returns statistics about the currently stored documents: count, total tokens, file types, storage size, average chunk length. Call this to audit what is in the RAG store before retrieval or cleanup.'
Add explicit bounds and enums to parameters. For limit and top_k: add minValue: 1, maxValue: 100 to the schema. For file_paths: add a pattern regex that validates file extensions or a constraint in the description. This prevents the LLM from passing invalid values.
Document file size and format constraints directly in the tool description, not just in error messages. Example: 'Supports PDF, DOCX, XLSX, PPTX up to 100MB each. Larger files will return an error with the max size limit.' This sets expectations upfront.
Add a discovery tool or enrich list responses. E.g., add a list_supported_file_formats() tool that returns {'formats': ['PDF', 'DOCX', 'XLSX', 'PPTX'], 'max_file_size_mb': 100, 'max_total_size_mb': 1000}. This lets agents check constraints before attempting uploads and reduces error-driven retries.
Return IDs and chaining references in all responses. E.g., retrieve_context_for_query should return not just matched text snippets but also doc_id, chunk_id, and metadata that downstream tools (e.g., cite_source, get_document_metadata) can use to follow up. Currently, the response structure is not documented so agents cannot know what fields exist to chain calls.