MCP server for academic bibliographic search across arXiv, Semantic Scholar, and other sources
Two tools with decent naming and visible schemas, but descriptions lack depth and error handling is minimal. Schemas are present but output documentation is incomplete. Inline descriptions in code are adequate but do not follow the LLM-optimized format (50-200 chars target). Tool names follow verb_noun convention (search_papers, get_paper_details), which is good. However, parameter descriptions are generic and do not include format constraints, ranges, or valid value guidance. No enum declarations for the 'sort_by' parameter despite having a known set of valid values. Output schemas are inferred from code rather than formally documented. Error handling is present (try/catch) but errors are generic strings that do not guide the LLM on recovery. No pagination support despite tools that could return many results. Tool composition is reasonable, two separate tools for search vs. detail retrieval, but the search tool returns plain text rather than structured JSON, forcing the LLM to parse unstructured output.
Get detailed information about a specific paper. Args: paper_id: Paper ID (arXiv ID or Semantic Scholar ID) source: Source ('arxiv' or 'semantic_scholar') Returns: JSON string containing paper details
Search for academic papers across multiple sources. Args: query: Search query string (supports advanced syntax like title:"machine learning" author:smith) sources: Comma-separated list of sources ('arxiv', 'semantic_scholar', 'google_scholar') max_results: Maximum number of results to return (default: 10) sort_by: Sort criterion ('relevance', 'date', 'citations', 'combined', 'impact') Returns: String containing search results with advanced deduplication and ranking
Output format is plain text, not structured JSON. Tool returns a formatted string with paper results concatenated, forcing the LLM to parse unstructured text. This violates the response-shaper pattern and wastes tokens on parsing.
'sort_by' parameter accepts free-form strings ('relevance', 'date', 'citations', 'combined', 'impact') but is not declared as an enum. This invites hallucinated values; LLM may pass invalid sort criteria. Parameter description lists valid values in prose rather than as a formal enum constraint.
'sources' parameter is a comma-separated string requiring the LLM to format it correctly ('arxiv,semantic_scholar' vs 'arxiv, semantic_scholar'). Should be an array of enum values instead. Current design invites format errors.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 50 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 17 | - | v1 |
No pagination support. Tool accepts max_results up to an unbounded integer; no indication of what happens if 1000 is requested. Large result sets risk context window exhaustion and degraded LLM reasoning. Should implement limit + offset or cursor-based pagination.
Error messages are generic strings ('Search failed: <error>') that do not guide the LLM on recovery. Should categorize errors (retryable vs. fatal) and suggest next steps. E.g., 'arXiv search timed out, try with a simpler query or fewer max_results.'
get_paper_details has no documentation of output schema. Tool description says 'Returns JSON string' but does not specify fields (authors, abstract, doi, etc.). LLM cannot plan downstream usage without knowing the response structure.
Parameter descriptions are generic and lack format/range guidance. E.g., 'paper_id: Paper ID (arXiv ID or Semantic Scholar ID)' does not clarify format (e.g., '2301.12345' for arXiv, alphanumeric string for Semantic Scholar). Should include regex pattern or example.