MCP server providing vector search tools for document retrieval and RAG capabilities, with support for semantic, keyword, and hybrid search across document collections
The server defines 10 tools (with 3 duplicates: search_documents, list_collections, get_collection appear twice). Tools have descriptions and basic input schemas, but quality is inconsistent. Descriptions vary from verbose (search_documents, 1/3 of tools) to minimal (truncated versions of duplicates). Parameter descriptions are present but often generic. No input validation constraints (enums, ranges). Output schemas are not documented. Error handling is not visible in tool definitions. Naming follows verb_noun convention but tool duplication reduces clarity. The server is STDIO-only, which is a hard cap at 50 for protocol readiness.
Add a text document to a collection.
Create a new collection.
Delete a collection and all its documents.
Get details of a specific collection.
Get details of a specific collection. This function retrieves detailed information about a specific document collection. It's useful for verifying collection details, checking metadata, or confirming that a collection exists before performing operations on it. The function provides basic information about the collection including its name and unique identifier.
List all available document collections.
Tool duplication: search_documents, list_collections, and get_collection are defined twice with inconsistent descriptions. The second definitions have truncated descriptions (27-28 chars, below effective threshold). This creates ambiguity for LLM tool selection and indicates a registration/configuration bug.
No output schemas documented for any tool. LLMs cannot plan downstream calls or extract fields without knowing response structure. E.g., search_documents should document it returns {results: [{document_id, content, metadata, relevance_score}], total_count}.
search_type parameter in search_documents lacks enum constraint in schema. Description lists options ('semantic', 'keyword', 'hybrid') but schema should formalize as enum to prevent LLM hallucination of invalid types.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 57 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 55 | - | v1 |
List all available document collections. This function retrieves and displays all document collections that are available in the system. It's typically the first step in the RAG workflow to identify which collection contains the relevant documents for a user's query. The function returns structured information about each collection including names, IDs, and metadata. Use this function to discover what collections are available before performing searches or other operations.
List documents in a collection.
Search documents in a collection using semantic, keyword, or hybrid search. This function is used to find relevant documents within a specific collection based on a search query. It supports multiple search types to provide flexible document retrieval capabilities. The function returns structured search results with document content, metadata, relevance scores, and document IDs.
Search documents in a collection using semantic, keyword, or hybrid search.
filter_json parameter (search_documents, add_documents) accepts raw JSON strings instead of object types. This forces LLMs to manually construct JSON, increasing errors. Should be typed as object with schema or clearly document expected JSON structure.
delete_collection is destructive but lacks confirmation/dry-run support. No mention in description of consequences or recovery. Per Arcade pattern, irreversible operations should support confirmation step.
Numeric parameters (limit in search_documents, list_documents) lack explicit min/max bounds. search_documents states 'default is 5, maximum allowed is 100' in description, but schema should enforce via maxItems/min/max constraints. Unbounded numbers let LLMs pass absurd values.
add_documents parameter naming is misleading: tool is 'add_documents' (plural) but 'text' parameter accepts single document. Should be 'add_document' (singular) or accept array of texts. Current naming suggests batch capability that doesn't exist.
No error handling guidance in tool definitions. What happens if collection_id doesn't exist? If search query is empty? If delete fails? Error responses should guide LLM recovery (retry, lookup, ask user, etc.).
No idempotency guarantees documented. If add_documents or create_collection are called twice with same input, what happens? Agents retry on failures, non-idempotent tools risk duplicate side effects.
No chaining IDs in responses. E.g., if add_documents returns a document_id, and later tool needs both collection_id and document_id, the response should include both to enable one-call chains. Current output schema is undocumented; cannot verify.