Model Context Protocol server for EBI BioStudies API. Provides reliable tools to interact with the BioStudies API including getting comprehensive information about specific studies with rich metadata extraction, validating study accession numbers and checking study existence, and batch processing multiple studies efficiently.
The BioStudies MCP server presents three READ_ONLY tools with clear naming conventions (verb_noun pattern) and complete input schemas. However, several significant gaps prevent a higher score: (1) parameter descriptions are minimal or absent in some cases, (2) output schemas are not documented at all, (3) error handling is basic with no recovery guidance, and (4) no pagination or result-limiting mechanisms are visible despite batch operations. The tool set is focused and non-overlapping, which is positive, but the implementation lacks the rigor expected of production-grade tools. Tool names are appropriately specific (get_study_details, validate_study_accession, batch_get_studies) and align with verb_noun conventions. Input schemas are properly typed with JSON Schema format. However, the descriptions, while present, lack the depth and structure (WHAT, WHEN, prerequisites) expected by LLM-based agents.
Retrieve information for multiple studies in a single request (maximum 50 studies). Efficiently processes multiple accession numbers and provides success/failure status for each.
Get comprehensive information about a specific biological study by its accession number. This tool provides rich metadata including study attributes, section details, external references, associated files, and subsections.
Validate a study accession number format and check if the study exists. Supports all BioStudies accession formats including S-BSST, E-MTAB, EMPIAR, S-BIAD, and others.
Output schemas are completely undocumented. No evidence in source code of documented return types, field definitions, or response structure for any of the three tools. LLMs cannot plan downstream calls or extract relevant data without knowing what fields are returned.
Parameter descriptions are sparse. The 'accno' parameter in get_study_details and validate_study_accession has a description, but no constraints on format, length, or valid patterns beyond examples. The 'accessions' parameter in batch_get_studies has a basic description but lacks guidance on what happens if an accession is invalid or doesn't exist.
No error recovery guidance. Error handling in src/index.ts catches errors but returns generic text responses ('Error: <message>'). There is no classification of errors as retryable, user-fixable, or fatal. No guidance on what the LLM should do next (e.g., 'try with a different accession format').
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 68 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 49 | - | v1 |
batch_get_studies lacks per-item error reporting. The schema states 'success/failure status for each' in the description, but the actual response structure is not documented. If one accession out of 50 fails, does the entire batch fail? Or does it return partial results with per-item status?
No pagination or result limits documented for get_study_details. The description mentions 'rich metadata' but does not state how large responses can be or whether results are truncated. This risks blowing the context window if a study has thousands of files or subsections.
Tool descriptions lack WHEN to use them. The get_study_details description states WHAT it does but not when to call it vs. validate_study_accession. A more complete description would say: 'Use this after validate_study_accession confirms the study exists, or directly if you already know the accession is valid.'
No documentation of tool composition or chaining. It is unclear whether get_study_details output contains references (e.g., related study IDs) that could chain into another get_study_details call. The response structure is opaque to the LLM.