MCP server for Unpaywall: DOI metadata, title search, OA links, and PDF text extraction
This server presents well-structured tool definitions with good naming conventions and detailed descriptions. All four tools use clear action verbs (unpaywall_get_*, unpaywall_search_*, unpaywall_fetch_*) and have substantive descriptions (150-400 chars). Input schemas are present and include type constraints. However, there are gaps in output schema documentation, missing parameter descriptions in some cases, and limited error handling guidance. The server demonstrates solid fundamentals but lacks the polish required for A-grade (80+).
Fetch and extract text from an open-access PDF. If DOI is provided, resolves best OA PDF via Unpaywall first; if pdf_url is provided directly, downloads from that URL. Text can be truncated via truncate_chars to avoid massive outputs (default 20000 chars).
Fetch Unpaywall metadata for a DOI (accepts DOI, DOI URL, or 'doi:' prefix). Requires an email address via env UNPAYWALL_EMAIL or the optional 'email' argument.
Fetch best OA fulltext links for a DOI via Unpaywall. Similar to unpaywall_get_by_doi but returns a simpler, focused response with only the OA link information. Requires UNPAYWALL_EMAIL env var or 'email' argument.
Search article titles and return Unpaywall-style open-access metadata for each hit. Supports optional is_oa filter and pagination (50 results per page). Backed by OpenAlex's /works endpoint because Unpaywall's own /v2/search has been returning HTTP 500 since its May 2025 rewrite; response shape still mirrors Unpaywall's documented search result (results[].response is a DOI record, results[].score, results[].snippet).
Output schemas not documented. No description of what unpaywall_get_by_doi, unpaywall_search_titles, unpaywall_get_fulltext_links, or unpaywall_fetch_pdf_text return. LLMs cannot plan downstream tool calls or extract the correct fields without knowing the response structure.
unpaywall_fetch_pdf_text has no required parameters (all optional) and lacks documentation about the mutual dependency between 'doi' and 'pdf_url'. The description states 'doi' and 'pdf_url' are optional, but then says 'email' is 'required if using doi', this circular dependency is not enforced in the schema and could lead to invalid calls (e.g., passing doi without email).
No error handling guidance in tool descriptions. Tools call external APIs (Unpaywall, OpenAlex, PDF download services) but descriptions do not explain what happens if the API fails, which errors are retryable, or what the LLM should do. Users are left guessing on transient vs. fatal failures.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 71 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
truncate_chars parameter in unpaywall_fetch_pdf_text lacks minimum/maximum bounds. The description says 'default 20000 chars' but does not specify valid ranges, memory limits, or consequences of very large values. An LLM could pass truncate_chars=999999999, causing memory exhaustion.
is_oa parameter in unpaywall_search_titles is a boolean but the description says 'If true, only return OA results; if false, only closed; omit for all'. This is confusing, omitting the parameter should also be supported, but the schema does not clarify null/undefined behavior. Should be explicit: 'true=OA only, false=closed-access only, null/omitted=all'.
page parameter in unpaywall_search_titles uses integer with minimum=1 but no maximum. OpenAlex API may accept arbitrary large page numbers. Should document: 'Page number (1-based, 50 results per page; valid range depends on result set size)'.