MCP server providing access to the BioStudies database API for retrieving and searching studies
BioStudies MCP server has 3 tools with clear naming (verb_noun convention: get_study, search_studies, get_study_info). However, schema and parameter documentation is minimal. Input schemas are present but lack proper type constraints and descriptions for complex parameters. The search_studies tool accepts a free-form 'params' string rather than structured parameters, forcing LLMs to construct query strings manually, this violates the constrained-input and parameter-description patterns. Descriptions are present but brief (averaging ~50-100 chars for tool descriptions, 20 chars for parameter descriptions). No output schemas are documented. Error handling returns raw HTTP status codes with minimal guidance. Composite description for search_studies is extensive but embedded in docstring rather than structured. No tool annotations (readOnlyHint). The server is read-only by design (all tools are READ_ONLY risk), which simplifies safety but doesn't excuse the structural gaps.
Get a study from the BioStudies database with the given accession
Get additional information for a study with the given accession. Returns information such as FTP link and relative path. Parameters: - accession: The BioStudies accession ID Returns: - JSON string with additional study information
Search for studies in the BioStudies database. Parameters: - params: A string containing all search parameters in the format "param1=value1¶m2=value2" Supported parameters include: - query: Searches for the provided text in all submissions * Each word is treated as a separate term unless in double quotes * Boolean operators (AND, OR, NOT) and brackets can modify behavior: e.g., "Leukemia AND (mouse OR human)" * Wildcards: * matches any sequence of characters, ? matches any single character * Regular expressions supported using /pattern/ syntax * Special characters (+, -, &&, ||, !, (), {}, [], ^, ", ~, *, ?, :, \, /) need to be escaped or quoted * DOIs and paths should be in quotes: e.g., "10.1371/journal.pone.0127346" or "eeg/fmri" - accession: Searches for a specific BioStudies accession (wildcards allowed after first character, e.g. S-EPMC*) - title: Searches for presence of the parameter in the title of the study - author: Searches for presence of the parameter in the name of the author(s)/submitter(s) - release_date: Searches for a specific release date (format: yyyy-mm-dd) Wildcards and ranges supported. For example: 2009* will search for experiments released in 2009 and [2008-01-01 2008-05-31] will search for experiments released between 1st of Jan and end of May 2008. - content: Free-text search in any part of the study content (including file names and links) - links: Number of links in the study - files: Number of files in the study - orcid: Searches for the ORCID of any authors of the study (if available) - type: Study type (supported: 'study', 'array', 'collection') - link_type: Searches for a specific type of link to external databases - link_value: Searches in the value of the link type field - page: Result page number (default: 1) - pageSize: Number of results per page (default: 20, max: 100) - sortBy: Sorting key (works only for numeric fields) - sortOrder: Sorting order ('ascending' or 'descending', default: 'descending') - collection: Optional collection name to limit search to a specific collection Collection-specific fields can also be used depending on the collection: For ArrayExpress collection, functional genomics experiments data: - experimental_design: Experiment design - study_type: Study type - experimental_factor: Experimental factor - experimental_factor_value: The value of an experimental factor - source_characteristics: Sample attribute values / Source Characteristics - source_characteristics_value: Sample attribute category / Source Characteristics value - technology: Technology - organism: Species/organism of the experiment, study or sample - gxa: Presence ("true") / absence ("false") in Expression Atlas - raw: Experiment has raw data available - processed: Experiment has processed data available - assay_count: The number of of assays - sample_count: The number of samples - experimental_factor_count: The number of experimental factors - miame_score: The MIAME compliance score - minseqe_score: The MINSEQE compliance score For BioModels collection: - domain: Domain of the model - curation_status: Curation status of the model - modelling_approach: Modeling approach used - model_format: Format of the model - model_tags: Tags associated with the model - organism: Organism in the model Returns: - JSON string with search results
search_studies accepts unstructured 'params' string instead of individual structured parameters. LLMs must manually construct 'param1=value1¶m2=value2' format, increasing error likelihood and violating the constrained-input pattern. Should decompose into discrete parameters: query, accession, title, author, release_date, content, links, files, orcid, type, link_type, link_value, page, pageSize, sortBy, sortOrder, collection.
Parameter descriptions are minimal or absent. Input schema shows accession param with description 'The BioStudies accession identifier' for get_study (good), but search_studies params only says 'A string containing all search parameters...' without specifying valid parameter names, format constraints, or expected behavior.
No output schemas documented for any tool. Rubric requires that tool responses must document expected fields (e.g., search_studies returns JSON with what structure?). LLMs cannot plan downstream calls or extract fields without knowing the response schema.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 42 | - | v1 |
Error handling returns raw HTTP status codes and response text. Example: 'Error: Request failed with status code 404. Response: {...}'. Does not provide recovery guidance, categorization, or actionable next steps. Try search_studies() to find the correct accession.').
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). Per current spec (2026-07-28), all tools should declare readOnlyHint: true since they are read-only. This enables clients to present safe vs. unsafe tools differently to users.
search_studies has complex collection-specific parameters (ArrayExpress fields like experimental_design, organism; BioModels fields like curation_status) embedded in docstring but not exposed in the input schema. LLMs cannot discover these optional parameters programmatically. Schema should include optional properties for collection-specific fields or separate tools per collection.
No pagination result documentation. search_studies sets defaults (page=1, pageSize=20) but response schema doesn't specify what total count or next_cursor fields are returned.