Calibre RAG for researchers: page-level citations and LLM access via MCP. Provides semantic search, annotation management, and bibliography tools for Calibre libraries.
Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
ARCHILLES provides 14 tools with explicit schemas, descriptions, and parameter definitions. Strengths: comprehensive parameter documentation with type constraints (min_length, defaults, enums for method/format), clear naming following verb_noun convention, and structured output guidance. Weaknesses: descriptions are functional but brief (average ~80 chars, below baseline 194), lack actionable guidance for LLM selection, missing output schema documentation, and insufficient error handling guidance. All tools are READ_ONLY or REVERSIBLE (appropriate for a Calibre metadata system), but descriptions do not explain when to use one tool vs. a similar one (e.g., search_annotations vs search_books_with_citations). Three tools have empty input schemas (list_annotated_books, list_tags conceptually, compute_annotation_hash has only one param), which lowers schema completeness score.
No output schemas documented for any tool. Callers cannot know what fields/types to expect in responses, forcing LLMs to guess and parse unstructured output.
Descriptions are below baseline (average ~68 chars vs. baseline 194 chars). Many lack differentiation guidance: search_annotations vs search_books_with_citations, get_book_details vs get_book_annotations, are not explained.
Document output schemas for all 14 tools. For each, specify: field names, types (string, integer, array, object), and example values. Use consistent field naming across tools (e.g., book_id, not id or bookId). This is THE single highest-impact improvement.
Expand tool descriptions to 100 - 200 characters explaining WHAT, WHEN, and how it differs from related tools. Example: 'search_annotations' → 'Full-text search across all book annotations and Calibre comments via semantic embeddings. Use this for exploratory search across your entire library; for targeted annotation extraction from a known book, use get_book_annotations() instead.'
Add 'error recovery' subsection to every tool description. Example: 'If the book is not found, the tool returns an error with available title suggestions. If annotations cannot be parsed, the tool skips malformed entries and returns a warning in metadata.'
Implement cursor-based or offset-based pagination for list_tags, list_books_by_author, search_annotations, search_books_with_citations. Return a next_cursor or total_count field in responses to enable result set traversal.
Clarify search tool differentiation: add to search_books_with_citations description that it searches book content with citation metadata, vs search_annotations which searches annotation text. Provide an example scenario for each.
Define a canonical 'book object' schema returned by multiple tools (get_book_details, list_books_by_author, detect_duplicates). Include: book_id, title, author, isbn, series, year, annotation_count, tags. Reuse across tools to reduce cognitive load.
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
No error recovery guidance. Tools do not document what errors might occur (e.g., book not found, annotation parsing failed) or how LLMs should respond (retry, ask user, skip).
Tools operate on Calibre-specific objects (books, annotations, tags) but descriptions use jargon without explanation. 'doublette' (duplicates), 'TOC markers', 'Calibre comments' assume user familiarity.
Pagination not explicitly supported in list/search tools. list_books_by_author, list_tags, and search_annotations accept max_results but no offset/cursor parameters. Large result sets risk context window exhaustion.
Tool composition unclear: search_books_with_citations and search_annotations both search but one is 'citation-aware'. No output example or guidance on when to use which.
search_annotationssearch_books_with_citations
Add domain context to descriptions: 'doublette' → 'duplicate (Calibre term)'; 'TOC markers' → 'table of contents highlights'; 'Calibre comments' → 'Calibre metadata comments (stored separately from inline annotations)'. Make domain-specific terms understandable to non-Calibre users.
For set_research_interests (WRITE tool), document side effects clearly: 'This updates the server-side research interest index and boosts future search_annotations and search_books_with_citations results matching these keywords.' Also provide a get_research_interests() tool to retrieve current state.
For watchdog_scan (REVERSIBLE tool), document output schema: 'Returns an object with fields: files_scanned (int), files_added (int), files_removed (int), errors (array of {file, reason}), and dry_run (bool showing if changes were applied).'
Add constraint documentation to parameters where implicit: book_path should document expected file formats (PDF, EPUB, MOBI, etc.); annotation_type should be in description (currently enum in schema only); source should list available Calibre libraries or 'default' value.