Documentation and email search MCP server with RAG, semantic search, and hybrid search modes
Librarian MCP demonstrates good definition quality with strong parameter schemas and detailed descriptions for search tools. All 4 tools have explicit schemas visible in source code. Naming follows verb_noun convention (search_*, get_*). However, descriptions vary in quality, output schemas are partially documented, and error handling lacks actionable recovery guidance. The server implements production-grade RAG/search patterns but omits tool annotations (readOnlyHint/destructiveHint) and lacks per-parameter validation constraints.
Retrieve full document content by path with optional section filtering
Get the current status of the documentation index including document counts by product and component
Search across all documentation with enhanced filtering and advanced RAG modes. Args: query: Search keywords (space-separated) product: Filter by product name (e.g., "symphony") component: Filter by component (e.g., "PAM") file_types: Filter by file extensions (e.g., [".md", ".docx"]) doc_type: Filter by document type (api, guide, architecture, reference, readme, documentation) tags: Filter by tags (documents with at least one matching tag) modified_after: Only docs modified after this date (ISO format: 2024-01-01) modified_before: Only docs modified before this date (ISO format: 2024-12-31) max_results: Maximum number of results (default: 10, max: 50) mode: Search mode - keyword, semantic, hybrid, rerank, hyde, or auto (auto selects best mode) include_parent_context: Include parent document context (title, summary, headings) enhance_results: Include rich metadata and summaries in results include_full_metadata: Include full metadata payload in enhanced results max_per_document: Maximum chunks per document (default: 3, 0=unlimited) Returns: Dictionary with search results including file paths, snippets, relevance scores, metadata, and optionally parent context
Search emails with email-specific filters. This tool is optimized for searching EML files with preprocessing (quote removal, signature removal, thread grouping). Supports inline search operators (Outlook/Gmail style): from:sender - Emails from sender (partial match) to:recipient - Emails to recipient (partial match) cc:recipient - Emails with CC recipient (partial match) subject:text - Subject contains text in:folder - Emails in folder (inbox, sent, important, etc.) has:attachment - Emails with attachments after:YYYY-MM-DD - Emails after date before:YYYY-MM-DD - Emails before date thread:id - Emails in specific thread Args: query: Search text with optional inline operators (e.g., "from:kiraly project deadline") sender: Filter by sender email address (partial match) recipient: Filter by recipient email address (partial match) cc: Filter by CC email address (partial match) thread_id: Filter by email thread ID (groups related emails) subject_contains: Filter by subject line (partial match) folder: Filter by email folder (inbox, sent, important, etc.) has_attachments: Filter emails with/without attachments date_after: Only emails after this date (ISO format: 2024-01-01) date_before: Only emails before this date (ISO format: 2024-12-31) max_results: Maximum number of results (default: 10, max: 50) mode: Search mode - keyword, semantic, hybrid, rerank, hyde, or auto include_parent_context: Include parent document context enhance_results: Include rich metadata and summaries in results include_full_metadata: Include full metadata payload in enhanced results collapse_threads: Collapse results by thread, showing only the best match per thread (default: True). When enabled, adds thread_count metadata showing total emails in that thread. max_per_document: Maximum chunks per document (default: 3, 0=unlimited) parse_operators: Whether to parse inline operators from query (default: True) Returns: Dictionary with email search results including: - from, to, cc, subject, date - thread_id for grouping related emails - attachment metadata - cleaned content (quotes and signature removed) - parsed_operators (if parse_operators=True)
get_document description is under-specified (10 words). Does not differentiate from search_documentation or explain when to call. Minimal descriptions (20-65 chars) weaken LLM tool selection.
Output schemas are not formally documented. search_documentation and search_emails describe return structure in narrative (description field) but do not provide JSON Schema definitions. LLMs cannot parse expected fields reliably without formal schemas.
Tool annotations missing. No readOnlyHint declared for any tool. All 4 tools are read-only (no state mutation) but lack explicit hint, downstream clients cannot optimize request batching or caching.
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 62 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 48 | - | v1 |
Error handling in code (tools.py search_documentation, search_emails) returns generic {'error': 'Search failed', 'detail': str(e)}. Does not guide LLM on retry strategy, alternatives, or next steps.
Parameter enum constraints partially documented. 'mode' param lists valid options in description (keyword, semantic, hybrid, rerank, hyde, auto) but 'doc_type' and 'folder' enumerate examples without JSON Schema enum enforcement. LLMs may generate invalid values.
No pagination guidance for large result sets. search_documentation returns {'results': <list>, 'total': <count>} but does not document result truncation behavior or provide next_cursor/offset for continuation.