Document retrieval and search API with semantic similarity and metadata filtering using vector embeddings
SemanDoc has 15 tools with basic schema definitions and descriptions, but exhibits significant gaps in naming clarity, parameter documentation, and output schema specification. Tool names follow action-verb conventions (create_, delete_, get_, search_, list_), but several are ambiguous or lack proper parameter type constraints. Descriptions exist but are often generic (10-80 chars) and fail to explain WHEN to use a tool, what happens on invocation, or what fields are returned. Input schemas are present but incomplete: parameters lack required type definitions in several tools (e.g., update_api_key_activation's is_active parameter type is documented in description but may not be formally typed). No output schemas are documented, making it impossible for LLMs to plan downstream tool calls. Error handling is minimal, most tools return generic HTTP 500 errors without recovery guidance. Security concerns exist: API key operations expose key_id and is_active in URLs/parameters without clear permission gating. The webhook tool (webhook_create_document) is particularly problematic, it bypasses standard metadata structure but lacks description of its unique purpose.
Create a new document in the vector store
Create multiple documents in a single batch operation
Create a new API key
Delete a document by its ID
Export all documents to CSV format
Force an immediate save of the vector store to disk
Retrieve a specific document by its ID
Missing output schema documentation for all 15 tools. LLMs cannot plan downstream calls or extract required fields without knowing response structure. No fields like 'created_at', 'document_id', 'total_count', 'next_cursor' documented.
Generic error handling returning HTTP 500 with minimal context. Error responses should guide recovery: 'Document not found. Try search_documents() with keywords.' or 'API key already exists, try list_api_keys() to see existing keys.'
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 52 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 41 | - | v1 |
Get statistics about documents in the vector store
Get all API keys
List all documents with optional filtering by tag and category
Rebuild the FAISS index from the current docstore state
Delete API key
Search documents using semantic similarity and optional metadata filters
Update API key status
Webhook endpoint for quickly creating documents with minimal data
Incomplete parameter descriptions. 'force_save_vector_store' and 'rebuild_index' accept zero params but descriptions do not explain prerequisites (e.g., 'Auto-save must be enabled'), failure modes, or when to call (e.g., 'Call after bulk document ingestion'). 'get_document_stats' lacks any description of output fields.
webhook_create_document exists alongside create_document with overlapping functionality but different metadata handling. No description explains when to use webhook variant vs standard. Naming is ambiguous, 'webhook_' prefix signals origin, not intent. Should be named 'create_document_quick' or merged into create_document with optional simplification mode.
API key operations expose sensitive operations (update_api_key_activation, remove_api_key) without clear permission gating or confirmation step. No dry-run pattern. delete_document and remove_api_key are destructive but lack confirmation/dry-run support, agents can irreversibly delete data without safety check.
Parameter documentation lacks expected format and constraints. 'document_id' does not specify format (UUID, alphanumeric, slug). 'tags' and 'categories' are arrays but no max length or character restrictions documented. 'score_threshold' in search_documents lacks range (0-1? 0-100?) or default value implications.
Pagination inconsistency: list_documents and list_api_keys use skip/limit with no 'total' field or 'next_cursor' documented. search_documents uses 'k' parameter (offset-less) instead of standard pagination. Without total counts or cursors, LLMs cannot iterate safely over large result sets.
export_documents description is 1 word ('Export all documents to CSV format') but no specification of return format, file encoding, or max document count. LLM cannot determine if output is a file stream, URL, base64, or error.