Academic paper retrieval aggregator that searches multiple sources simultaneously (arXiv, Semantic Scholar, Crossref, Scopus, ADSABS, Sci-Hub), automatically aggregates, deduplicates, and ranks results with advanced filtering capabilities.
The Scholar Aggregator MCP server has two well-intentioned tools for academic paper discovery, but falls short of production quality due to missing parameter descriptions, incomplete schema documentation, and lack of proper error handling guidance. Both tools have reasonable names and high-level descriptions, but the parameter-level definitions are sparse. The code shows logging but no structured error recovery patterns. The searchScholarPapers tool accepts 13 parameters but only 2-3 have meaningful constraints documented. Parameter types are visible in the Go structs but not all map clearly to the JSON Schema. Descriptions exist but are generic (e.g., 'optional' without explaining when to use each filter). Output schema is mentioned in code but not formally documented in the tool definition.
Get detailed information of a specific academic paper by any identifier (DOI, arXiv ID, PubMed ID, Semantic Scholar ID, etc.). Automatically determines the identifier type and queries the most appropriate data source.
Search academic papers from multiple sources simultaneously (arXiv, Semantic Scholar, Crossref, Scopus, ADSABS, Sci-Hub). Automatically aggregates, deduplicates, and ranks results. Supports advanced filtering by author, journal, year, citations, and open access status. Returns unified paper metadata with source attribution.
Parameter descriptions are minimal or missing context. 'optional' alone does not explain when to use author vs. title vs. journal filters, or how they combine. LLMs cannot infer usage from sparse labels.
No output schema formally documented. searchScholarPapers returns a complex result with TotalSources, ActiveSources, SearchTime, Papers array, SourceStatus, and Suggestions, but the tool definition does not expose this structure for LLM planning. The response is formatted as text + JSON in _meta, making it opaque to structured reasoning.
sort_by and sort_order parameters have no enum constraints documented. Code accepts 'relevance, citation_count, published_date, title' and 'asc or desc' but the schema does not declare these as enums, inviting LLM hallucination of invalid values like 'date' or 'ascending'.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 43 | - | v1 |
Error responses lack recovery guidance. When searchScholarPapers fails, the code returns fmt.Errorf() with a generic message. No hints like 'Try simplifying your query' or 'Increase timeout if sources are slow' are provided to guide LLM retry logic.
No pagination guidance. searchScholarPapers accepts offset and limit, but does not document the maximum result count, whether there is a next_cursor, or when pagination is required. LLMs cannot know if 100 results is safe or will exhaust context.
enabled_sources parameter accepts arbitrary strings but source availability is runtime-dependent. No enum of valid sources (arXiv, Semantic Scholar, Crossref, Scopus, ADSABS, Sci-Hub) is declared in the schema, yet the description lists them. LLMs will guess at valid values.
getScholarPaper identifier parameter is overly generic. Description says 'DOI, arXiv ID, PubMed ID, Semantic Scholar ID, etc.' but does not specify format constraints or examples for each type. LLMs cannot distinguish a valid arXiv ID (e.g., '2301.12345') from a Semantic Scholar ID without examples.