Search your research library with synthesis using PaperQA2. Enables Claude to query academic papers and documents with AI-powered literature search and context retrieval.
The PaperQA2 MCP server has inconsistent tool definitions with moderate quality issues. 4 of 6 tools have reasonable descriptions and schemas visible in paperqa-mcp/server.py, but 2 tools (search_literature from archive tests and get_library_status) lack proper schema visibility or complete registration patterns. Parameter descriptions are present but generic. Output schemas are not documented. Error handling exists but is minimal. Tools do not implement idempotent patterns or confirmation steps for destructive operations (remove_document). The server architecture suggests tools are defined across multiple files (archive/redundant-tests and paperqa-mcp), making it difficult to verify complete, canonical definitions.
Add a document to the library. Copies file and starts background indexing.
Check status of background document indexing jobs.
Get raw text chunks matching a query. Uses the search index directly for low-cost retrieval without full agent synthesis.
Check what's in your research library and system status.
Remove a document from the library.
Minimal literature search - no progress reporting, no Context usage. Just query -> answer
Search your research library with synthesis.
get_library_status and check_indexing_status have empty input schemas ({}), making them appear as tools with no parameters. The descriptions are vague ('Check what's in your research library and system status' / 'Check status of background document indexing jobs') and under 50 characters of actionable context. LLMs cannot determine what fields they'll receive or when to call them instead of similar tools.
remove_document is a DESTRUCTIVE tool (deletes data from library) but lacks a confirmation step or dry-run option. The description does not state 'This removes the document irreversibly' or suggest a confirmation pattern. Agents can accidentally delete research papers without safeguards.
Output schemas are not documented for any tool. Users cannot see what fields search_literature, get_contexts, or add_document will return. This forces LLMs to guess what data is available, risking failed downstream tool chains and missed opportunities to include chaining IDs (e.g., document_id after add_document).
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 0 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 36 | - | v1 |
Debug version with proper typing and error handling
Ultra minimal test - no PaperQA calls at all
Even simpler test with no parameters
search_literature parameter 'mode' description contains literal enum values in the description ('"fast" (3 sources, cheaper) or "thorough" (15 sources, full synthesis)') rather than as a JSON Schema enum constraint. This is harder for LLMs to parse and does not prevent hallucinated values like 'medium' or 'balanced'.
add_document parameter 'file_path' requires an absolute path. The description does not clarify expected format, whether relative paths are accepted, or what happens if the file does not exist. No validation guidance for LLMs.
Tool definitions appear distributed across archive/redundant-tests (search_literature, get_library_status, add_document) and paperqa-mcp/server.py (get_contexts, remove_document, check_indexing_status). This fragmentation makes canonical definitions hard to verify. The source code shows multiple debug/minimal versions (server_debug.py, server_minimal.py, server_ultra_minimal.py) but it's unclear which is the active production registration.