MCP server for fetching biological and chemical compound information from multiple public databases including ChEMBL, PDB, HUGO, PubChem, and UniProt
mcp-omics defines 5 tools for querying omics databases (ChEMBL, PDB, HUGO, PubChem, UniProt). While all tools have names starting with action verbs (get_*) and include descriptions with examples, the implementation has significant quality gaps: (1) Input schemas are minimal, each tool accepts only a single string parameter with minimal description; (2) Output schemas are completely undocumented, callers cannot know what fields to expect from API responses; (3) Error handling is basic, returns generic 'not found' errors without recovery guidance; (4) Descriptions, while present, rely on examples (e.g., 'Example: CHEMBL25') rather than formal constraints, which LLMs tend to reuse literally. The tools are read-only and narrowly scoped (no composition challenges), but the lack of structured output documentation and minimal parameter validation make these tools difficult for LLMs to use reliably.
Fetch compound metadata using a ChEMBL ID. Example: 'CHEMBL25' This function retrieves the compound's preferred name, molecular weight, structure (SMILES), mechanism of action, and target information.
Fetch gene information from HUGO. Example: 'BRCA1'. This function retrieves the gene's name, description, and other IDs.
Fetch information about a protein from the Protein Data Bank using its PDB ID. Example: '7WRL'. This function retrieves the protein's title, experimental method, resolution, and release date.
Fetch basic compound information using a PubChem CID. Example: '2244' for caffeine. This function retrieves the compound's name, molecular weight, and SMILES representation.
Fetch basic protein information using a UniProt ID. Example: 'P43220' for the GLP-1 Receptor. This function retrieves the protein name and function information.
Output schemas are completely undocumented. Code returns raw API responses (compound_resp.json(), response.json(), gene_info, etc.) without documenting what fields the LLM should expect. LLMs cannot plan downstream calls or extract relevant data without knowing the response structure.
Parameter descriptions lack formal constraints. All tools include example values in descriptions (e.g., 'Example: CHEMBL25', 'Example: 7WRL') which LLMs tend to reuse literally in subsequent calls. No format, length, or pattern constraints are documented.
Error handling is generic and provides no recovery guidance. All tools return {'error': 'X not found'} without suggesting alternatives, retryability, or next steps. Baseline requires errors to be categorized and include recovery hints.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 36 | - | v1 |
Input schemas are minimal and lack type/format metadata. FastMCP infers schemas from function signatures (single string param), but no min/max length, regex patterns, or enums are declared. Descriptions repeat only the parameter name ('ChEMBL compound identifier') without actionable format details.
No pagination or result limiting. Tools dump entire API responses (especially get_chembl_info which may include full target lists, get_gene_info which may return many IDs). Large responses dilute signal, waste tokens, and increase hallucination likelihood. Baseline requires limit/offset and total count.
API responses are not stripped. get_pdb_info manually extracts 4 fields, but get_chembl_info, get_gene_info, get_pubchem_info, and get_uniprot_info return raw API JSON including audit fields, metadata, and irrelevant structures. This wastes tokens and dilutes signal.