This is a MCP server for researching on PubMed, Google Scholar, and arXiv. It will read/update your Google Drive for context.
The server defines 14 tools with explicit schemas and descriptions visible in src/deepresearch/server.py. Naming follows verb_noun convention (search_papers, fetch_paper_metadata, download_fulltext, etc.), which is appropriate. However, several critical gaps reduce quality: (1) Tool descriptions are generally adequate (100-150 chars) but lack guidance on WHEN to use each tool vs. alternatives, e.g., search_papers vs. search_apis vs. test_api_connector are not clearly differentiated. (2) Parameter descriptions are present but generic; many lack format/constraint details (e.g., 'paper_id' examples show format but no explanation of how to obtain one). (3) Output schemas are NOT documented anywhere, tools return unspecified structure, forcing LLMs to guess field names and types. (4) Error handling is absent from descriptions; no guidance on retryability, user-fixable errors, or recovery paths. (5) Two WRITE tools (store_to_drive, download_paper) lack confirmation/dry-run patterns and do not describe permission requirements. (6) Citation graph depth parameter accepts integers with no bounds (defaults to 1 but could accept unbounded values). This is typical of community MCP servers, functional definitions but missing production-grade polish around composition, error guidance, and output documentation.
Analyze publication trends over time, identify emerging topics, and plot term frequencies for a research area.
Highlight key sentences and extract keywords from a document.
Compare multiple scholarly papers to highlight similarities and differences in methods, results, and limitations.
Retrieve or download the PDF/HTML full text for a given paper ID.
Download a paper by its ID from the appropriate source.
Extract relationships between concepts in a scholarly paper (e.g., 'X causes Y' or 'Algorithm A outperforms B').
Fetch detailed metadata (title, authors, abstract) for a given paper ID.
No output schemas documented for any tool. LLMs cannot determine what fields to expect in responses (e.g., does search_papers return paper_id, arxiv_id, or some other identifier?). This forces agents to guess field names and risks incorrect chaining.
Tool descriptions lack differentiation and composition guidance. Three search tools exist (search_papers, search_apis, test_api_connector) but descriptions do not explain when to use each vs. the others. LLMs will guess and potentially invoke the wrong tool.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 46 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 41 | - | v1 |
Return citation relationships for a given paper or set of papers.
Search across multiple scholarly API sources with a single query.
Search multiple scholarly sources for papers matching a query.
Save fetched PDFs and summaries to the user's Google Drive folder.
Generate a structured summary (background, methods, results, conclusions).
Generate a focused summary of a specific section in a scholarly paper.
Test a specific API connector and return diagnostics.
Two WRITE tools (store_to_drive, download_paper) lack error handling descriptions, permission requirements, or confirmation patterns. No guidance on retryability or destructive consequences. Agents cannot reason about whether to retry or ask the user.
Parameter 'paper_id' format is unclear across multiple tools. Examples given ('arxiv:2104.08935', 'pubmed:12345678') but no description of how to obtain valid paper_ids or what happens if format is wrong. Agents may hallucinate invalid IDs.
Numeric parameters lack bounds or defaults. depth in get_citation_graph accepts any integer with no stated min/max; max_results and max_citations similarly unbounded. LLMs may pass absurd values (e.g., max_results=999999) that break API calls or cause timeouts.
No documentation of required permissions or scope for tools. store_to_drive and download_paper imply Google Drive/file system access but do not declare permission requirements. Agents cannot verify authorization before executing.
Pagination not mentioned for tools returning lists (search_papers, search_apis, get_citation_graph, analyze_trends). No indication of cursor, offset, or total_count in responses. Large result sets could exhaust context windows.