MCP server for academic reference verification against Crossref, Semantic Scholar, arXiv, Scopus, and IEEE Xplore
refcheck provides three well-named academic reference tools with solid descriptions and functional schemas. Tool names follow verb_noun conventions (verify_reference, search_references, get_bibtex). All three tools have non-trivial descriptions in the 150-300 character range, meeting the 10-1024 baseline. Parameter schemas are present with type definitions and descriptions. However, there are notable gaps: (1) output schemas are not documented in the visible code, the tools return .model_dump() of Pydantic models, but the expected structure is not declared in tool definitions; (2) error handling lacks recovery guidance, validation failures (e.g. 'At least one of title or doi must be provided') are returned as generic discrepancy lists rather than actionable error messages; (3) parameter constraints are underspecified, max_results lacks bounds, year_from/year_to lack range constraints; (4) no tool annotations (readOnlyHint/destructiveHint/idempotentHint) despite all tools being read-only operations. The tools are functionally sound and well-composed (each has a single responsibility), but lack the polish and error-guidance expected of production grade.
Get BibTeX entries for one or more papers by DOI, title, or Semantic Scholar ID.
Search for real, verified academic papers on a topic. Returns only real references from Crossref, Semantic Scholar, arXiv, and optionally Scopus/IEEE. Use this to find legitimate citations instead of letting the AI fabricate them.
Verify an academic reference against real publication databases. Checks whether a citation is real by looking it up in Crossref, Semantic Scholar, arXiv, and optionally Scopus/IEEE. Returns a confidence verdict: "verified", "partial_match", or "not_found". At least one of `title` or `doi` must be provided.
Output schemas not declared in tool definitions. Tools return Pydantic model_dump() output (e.g. VerifyResult, BibtexResult), but the MCP tool definitions do not include outputSchema. LLMs cannot infer expected response structure, limiting their ability to parse results and chain tools.
Error handling lacks recovery guidance. When validation fails (e.g. 'At least one of title or doi must be provided'), the tool returns a generic VerifyResult with confidence=0.0 and discrepancies=[...]. No actionable error message guides the LLM on what to do next (e.g. 'Provide a title or DOI and retry').
Parameter constraints underspecified. search_references has max_results with no bounds (could accept 999999). year_from/year_to lack range clarification (valid years: 1900-2100?). These unbounded parameters invite LLM errors and API abuse.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 36 | - | v1 |
No tool annotations. All three tools are read-only (no side effects), but lack readOnlyHint annotations. This metadata helps clients (e.g. hosted agents) categorize tools and optimize execution planning.
Mutually exclusive parameters not documented. get_bibtex accepts doi, dois, title, or semantic_scholar_id. The tool implementation (not shown in detail) likely enforces exclusivity, but parameter descriptions do not state 'provide exactly one of: doi, dois, title, or semantic_scholar_id'. LLMs may pass multiple parameters.