Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
This RAG server defines 2 tools with explicit schemas and descriptions, but exhibits significant gaps in description depth, parameter documentation, and output schema specification. Tool names are action-based and clear (search, get_document_count), but descriptions are very brief (Japanese text under 30 chars translated) and lack guidance on when/why to use each tool. Parameters have types and defaults, but most lack descriptions explaining their purpose or constraints. Output schemas are not documented, handlers return generic text content structures without formal schema definitions. Error handling is present but generic. No tool annotations (readOnlyHint) are declared despite both tools being read-only. The codebase is well-structured with proper separation of concerns, but the tool interface design prioritizes implementation simplicity over agent usability.
Tool descriptions are extremely brief (< 30 chars) and lack context for LLM selection. 'ベクトル検索を行います' ("Perform vector search") does not explain WHEN to use search vs alternatives, what it returns, or prerequisites (e.g., 'requires indexed documents').
Parameter descriptions are missing or minimal. 'with_context' has description 'bool 前後のチャンクも取得するかどうか' but does not explain what 'context' means (neighboring chunks? document structure?), or why the LLM should enable/disable it. 'context_size' ("Number of chunks to fetch before/after") lacks guidance on valid ranges or performance implications.
Output schemas are not documented. Handlers return {'content': [{'type': 'text', 'text': '...'}], 'isError': true/false}, but this structure is not formally specified anywhere. LLMs cannot plan downstream tool calls or extract structured data without knowing what fields/types to expect.
Recommendations
Expand tool descriptions to 50 - 150 chars. For 'search': 'Search indexed documents by semantic similarity to a query. Returns matching chunks with similarity scores, optional context from adjacent chunks, and optionally the full parent document. Use this to find relevant information before retrieving full documents.' For 'get_document_count': 'Retrieve the total number of documents currently indexed in the RAG database. Call this to check if documents have been indexed before performing searches.'
Add descriptions to ALL parameters. Examples: query: 'Natural language search query (required). Will be embedded and compared against indexed document chunks.'; limit: 'Maximum number of matching chunks to return (1 - 100, default 5). Higher limits return more context but increase response size.'; with_context: 'If true, include adjacent chunks (one before, one after) from the same document for context. Helpful for understanding surrounding content.'; context_size: 'Number of adjacent chunks to fetch on each side when with_context is true (default 1, max 5)'; full_document: 'If true, return the entire parent document containing the best match, in addition to individual matching chunks. Useful for getting complete context.'
Document the output schema formally. Define a SearchResult type with fields: {'results': [{'chunk_id': string, 'file_path': string, 'chunk_index': int, 'content': string, 'similarity': float (0 - 1), 'is_context': bool, 'is_full_document': bool}], 'total_results': int, 'query': string} and a DocumentCountResult type with {'document_count': int}. Include these in a schema/types section of the server config or README.
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) are declared, despite both tools being read-only. Per current MCP spec (2026-07-28), tools should declare their safety profile via annotations so agents can make safe decisions (e.g., retry a read-only call without hesitation).
search tool returns results as unstructured text blocks rather than a paginated, typed list. No limit enforcement, if 1000 matching chunks exist, all could be returned in a single response, bloating context and risking token exhaustion.
Error messages are generic and do not guide recovery. 'search中にエラーが発生しました: {str(e)}' ("Error occurred during search: {e}") dumps exception text; LLM cannot determine if error is retryable, user-fixable, or fatal. Call CLI: python -m src.cli index').
The 'limit' parameter defaults to 5 and has no documented maximum or performance guidance. LLMs may not know that passing limit=1000 is unsafe or will timeout.
search
Add readOnlyHint: true to both tool definitions in server.register_tool() calls, signaling that these tools are safe to retry and produce no side effects.
Implement pagination and result limiting. Modify search_handler to respect the limit parameter strictly (cap at 100) and document this in the description. Add an optional offset/page parameter if the search service supports it. Include 'total_results' in response so LLMs know if more results are available.
Improve error messages to categorize and guide recovery. Examples: 'No documents indexed. Run: python -m src.cli index <path>' (user-fixable). 'Search timed out after 30s. Try a shorter query or smaller limit.' (retryable). 'Invalid limit: 101 > max 100' (user-fixable). Each should specify the category and suggest the next tool/step.
Add range constraints to numeric parameters in descriptions. limit: '1 - 100, default 5'; context_size: '0 - 5, default 1'. Consider adding min/max to the schema if the MCP library supports it.
Translate descriptions to English or provide English alongside Japanese, ensuring clarity for non-Japanese-speaking agents and developers.
Add a 'search_tips' or 'rationale' section to the server's startup/capabilities response, explaining: when to call search vs get_document_count, what query formats work best, and when to enable full_document retrieval for complete context.