AI-powered vector database for semantic document search, clustering, and knowledge management. Supports PDF, DOCX, PPTX, images (OCR), code files, and more.
Vector Knowledge Base has 22 tools with generally adequate naming and schema structure, but suffers from incomplete parameter descriptions, missing output schema documentation, and vague descriptions that don't clearly differentiate similar tools. Tool names follow verb_noun convention well (search_documents, delete_document, create_folder), but many parameter descriptions are minimal (e.g., 'Optional filters', 'Optional tags') without explaining expected format, constraints, or when to use them. No tool descriptions explain when to call one vs. a similar tool (e.g., when to use upload_file vs upload_folder_batch vs mcp_create_document). Output schemas are not documented in any tool. Error handling and recovery guidance are absent. The server uses fastapi-mcp on HTTP transport, which is current, but lacks tool annotations (destructiveHint, readOnlyHint) despite having 4 destructive operations and 13 read-only operations.
Perform clustering on all document embeddings.
Create a new folder in the hierarchy.
Delete a document and all its chunks by filename.
Delete a folder and optionally its contents.
Export documents and metadata.
Return list of allowed file extensions for frontend validation.
Get cluster information and assignments.
Parameter descriptions are minimal and non-actionable. Examples: 'Optional filters' (what structure?), 'Optional tags' (comma-separated format not stated), 'Optional folder path' (absolute or relative? what if path doesn't exist?). LLMs cannot infer parameter format from vague descriptions.
No output schemas documented for any tool. LLMs cannot plan downstream tool chains or extract necessary fields (e.g., what fields does search_documents return? Does it include embedding_id, score, chunk_id?). Without documented outputs, agents resort to trial-and-error.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 60 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 48 | - | v1 |
Get 3D reduced embeddings for all documents for visualization.
Get all files organized within the folder structure.
Get the folder hierarchy structure.
Get the status of an async job.
Get files that have not been organized into folders.
Health check endpoint to verify the API is running
List all uploaded documents.
Create a document directly via MCP (AI agent only).
Move a file to a different folder.
Reset the database (destructive operation).
Search for documents using vector similarity.
Transform a query embedding into 3D space for visualization.
Update a folder's name or parent.
Upload a file for ingestion into the vector database.
Upload multiple files from a folder in batch.
Three similar upload tools (upload_file, upload_folder_batch, mcp_create_document) lack differentiation guidance. Descriptions don't explain when to use one vs. another. 'upload_file' says 'Upload a file' and 'mcp_create_document' says 'Create a document directly via MCP (AI agent only)', but what's the functional difference? When would an LLM choose one?
Destructive operations (delete_document, delete_folder, reset_data) lack confirmation or dry-run support. reset_data takes an admin_key but no description explains what this key is, how the agent obtains it, or whether the operation is logged. An LLM could accidentally reset the entire database.
Tool annotations missing. 13 read-only tools (get_*, list_*, search_*) lack readOnlyHint. 4 destructive tools (delete_*, reset_data) lack destructiveHint. This metadata enables agents to reason about side effects and avoid unsafe operations.
Parameter 'filters' in search_documents and export_data is described as 'Optional filter criteria' (object type) with no schema. What fields are valid? What operators (eq, gt, lt, contains)? LLMs will hallucinate filter syntax.
Error handling and recovery guidance completely absent. No tool describes what errors might occur, how to recover, or what the LLM should do next. Example: if upload_file fails due to unsupported file type, should the LLM call get_allowed_extensions? This is undocumented.
Naming ambiguity: 'cluster_documents' suggests it returns clustered results, but does it perform clustering or just trigger async work? Description says 'Perform clustering on all document embeddings' (imperative, sounds stateful), but get_job_status suggests async behavior. Unclear if cluster_documents blocks or returns immediately.
Parameter 'folder_path' in multiple tools (upload_file, upload_folder_batch, mcp_create_document) is vague. Does it expect a filesystem path, a database ID, or a navigation path? What's the expected format? Can agents pass nested paths like 'documents/2024/january'?
Pagination and result limits undocumented. list_documents, get_folders, get_clusters, get_embeddings_3d return entire collections with no mention of limits or pagination. Returning thousands of embeddings in get_embeddings_3d could exhaust context windows.