MCP server for scholarly research across PubMed, ArXiv, Google Scholar, and Firecrawl-backed web search (five consolidated tools)
The server defines 5 tools with explicit schemas and descriptions visible in tests/tool-tests.config.js. Naming follows verb_noun pattern (research_search, paper_analysis, citation_manager, research_preferences, web_research). Descriptions are substantive (average ~250 chars), exceeding the 194-char production baseline. Input schemas are present with type definitions and enums. However, several critical gaps reduce the score: (1) output schemas are not documented, no description of what fields each tool returns, forcing LLMs to guess at response structure; (2) parameter descriptions, while present, lack actionable format constraints (ranges, regex patterns, examples of valid IDs); (3) no error handling guidance in tool descriptions, LLMs cannot know what to do if a search fails or an identifier is invalid; (4) research_preferences is a configuration tool that mixes READ and WRITE concerns, violating single-responsibility pattern; (5) no idempotency hints despite stateful operations (cache management, preference updates). Tool naming is clear and verb-driven, but the server prioritizes breadth (5 tools covering 4 different data sources) over depth of definition quality.
Manage citations: count citations for a paper, generate formatted citations in APA/MLA/Chicago/Harvard/Nature styles, or get citation network metadata. identifier can be PMID, PMCID, DOI, or ArXiv ID. action: count fetches citation counts from Google Scholar; generate produces formatted citations; network returns in-references and out-references.
Retrieve metadata and optional full text for one paper. Identifiers: PMID (digits), PMCID (PMC plus digits), DOI (10.... or doi:...), ArXiv (YYYY.NNNNN, arxiv:..., or legacy category/number with slash). basic returns metadata and abstract only. full-text loads PubMed PMC text when available, splits heuristic sections capped by maxSectionLength, and otherwise shows a whitespace-normalized excerpt up to maxFullTextChars. textContains optionally keeps sentences containing that substring (case-insensitive). Non-PubMed records return abstract-only text for full-text mode.
Get or set research preferences: cache settings (enabled, TTL in seconds), preferred search sources (preferPubMed, preferGoogleScholar, preferArXiv, preferFirecrawl), and Firecrawl API key configuration. action get retrieves current prefs; action set updates them. category all returns all settings; category search returns only search prefs; category cache returns cache settings; category firecrawl returns Firecrawl config.
Search PubMed, Google Scholar, and ArXiv. This tool passes overrideSources so the sources you list are queried even when a source is disabled in saved preferences. Date filters: startDate is passed to PubMed mindate as given (prefer YYYY/MM/DD). endDate supplies the publication year for ArXiv and Google Scholar (first path segment, e.g. YYYY from YYYY/MM/DD) and PubMed maxdate as that year string. sortBy citations mainly affects Google Scholar; ArXiv has no citation metadata here. Google Scholar may be blocked; partial failures appear under Source warnings when adapters throw. In-memory search caching applies when research_preferences cache.enabled is true (TTL from preferences).
No documented output schemas for any tool. LLMs cannot infer response structure (field names, types, arrays vs objects). research_search returns merged results from 3 sources but the shape is not documented. paper_analysis returns 'metadata and abstract' but field names are not specified. This forces LLMs to guess, increasing hallucination risk.
Parameter descriptions lack actionable format constraints. research_search's startDate says 'YYYY/MM/DD recommended' but this is a hint, not a formal constraint, the description does not enforce it as required format. paper_analysis's identifier accepts PMID, PMCID, DOI, or ArXiv but the format rules (e.g. 'PMID is digits only', 'DOI starts with 10') are buried in prose, not in enum or pattern definitions. LLMs frequently ignore prose constraints and pass invalid formats.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 36 | 2024-11-05+ | v1 |
Scrape web content or search the web via Firecrawl. action scrape retrieves full page text and metadata from a URL (requires Firecrawl API key); action search queries the web and returns top results with snippets. Both actions require Firecrawl API key from environment (FIRECRAWL_API_KEY) or set via research_preferences.
research_preferences mixes GET and SET operations with WRITE side effects (e.g. setting Firecrawl API key, cache settings). Single-responsibility principle violated, this should be split into research_preferences_get (read-only) and research_preferences_set (write). Per Agentic Tool Patterns, each tool should do exactly one thing.
No error handling guidance in tool descriptions. research_search notes 'Google Scholar may be blocked; partial failures appear under Source warnings' but does not tell the LLM how to handle this state (retry? use other sources? inform user?). paper_analysis accepts multiple identifier formats but does not document what error the LLM should expect if the identifier is not found, or how to recover. Missing recovery guides per pattern:recovery-guide.
research_search's maxResults parameter defaults to 20 but 3 sources are queried and merged, the token cost of returning merged results from PubMed, Google Scholar, and ArXiv could easily exceed context budgets. The description does not warn about result explosion or offer pagination guidance. No total count or next_cursor documented.
citation_manager's 'format' parameter is required only for the 'generate' action, but the schema does not make this conditional requirement explicit. LLMs may pass format=None for count or network actions, causing confusion or runtime errors.
web_research description states 'requires Firecrawl API key from environment (FIRECRAWL_API_KEY) or set via research_preferences', this couples the web_research tool to server-side secret management, but API key is not exposed as a parameter (correct for security). However, if the key is missing, the error message should explicitly guide the user to call research_preferences_set or set the environment variable, per pattern:recovery-guide.