Search the PubChem chemical database for compounds, properties, safety data, bioactivity, cross-references, and entity summaries via MCP. STDIO or Streamable HTTP.
PubChem MCP server demonstrates solid definition quality with well-structured tool schemas, comprehensive parameter documentation, and clear descriptions. All 5 tools have proper JSON Schema definitions with type specifications and parameter descriptions. Tool names follow verb_noun conventions (get_*, search_*, list_*). Descriptions are substantive (150-350 chars average), explaining WHAT the tool does, WHEN to use it, and key behaviors. Parameters include enums for constrained inputs (e.g., outcome filters, format choices), min/max guidance for numeric parameters (offset, maxResults ranges 1-100), and proper typing. Output schemas are documented via description text and example field names. Error handling guidance is implicit through clear parameter constraints. Main gaps: (1) Output schemas lack formal JSON Schema definitions in the tool registration (only implied through description text); (2) No explicit error recovery guides in descriptions; (3) Missing idempotent/readonly hints in tool metadata (though descriptions indicate read-only risk classification). Overall, this is a well-designed set of tools suitable for production use by agentic callers.
Get a compound's bioactivity profile: which assays tested it, activity outcomes (Active/Inactive/Inconclusive), target identifiers (NCBI Gene ID, UniProt/GenBank accession), and quantitative values (IC50, EC50, Ki, etc.). Filter by outcome and/or a specific molecular target (NCBI Gene ID or protein accession) to focus the profile — e.g. "is this compound active against target T?".
Get a compound's default 3D conformer — atomic coordinates and bonds — for one CID. format="json" (default) returns atoms and bonds parsed into structured fields; format="sdf" returns the raw V2000 SDF text for passthrough to docking, rendering, or conformer tools. Optionally lists alternate conformer IDs. Not every compound has computed 3D coordinates (large molecules, mixtures, and some salts do not).
Get detailed compound information by CID. Returns physicochemical properties (molecular weight, SMILES, InChIKey, XLogP, TPSA, etc.), optionally with a textual description (pharmacology, mechanism, therapeutic use), known synonyms, drug-likeness assessment (Lipinski/Veber rules), and/or pharmacological classification (FDA classes, MeSH classes, ATC codes). Accepts up to 100 CIDs per call.
Fetch a 2D structure diagram (PNG image) for a compound by CID.
Output schemas lack formal JSON Schema structure in tool registration. Descriptions reference output fields (e.g., 'atomCount', 'nextOffset', 'targetGeneId') but no explicit outputSchema property is visible in tool definitions. LLMs must infer response structure from description text alone.
No explicit error recovery guidance in tool descriptions. While parameter constraints are clear, descriptions do not state 'If the compound has no 3D coordinates, the tool returns an empty result' or similar error scenarios. This leaves LLMs guessing about failure modes.
Tool annotations (readOnlyHint, destructiveHint, idempotentHint) are not explicitly declared in tool definitions. Risk field in source indicates 'READ_ONLY' but this should be formalized as toolAnnotations in the MCP schema.
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | C | 69 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 12 | 2025-03-26+ | v1 |
Get a compound's interaction data: drug-drug interactions (DrugBank), drug-food interactions, and chemical-target interactions (binding/activity from BindingDB, ChEMBL, and others). Each entry carries its originating source. Results are paged per kind, with the source-record total and the next offset reported for each. Richest for approved drugs; many compounds have no deposited interaction records.
pubchem_get_compound_image description is brief (75 chars) and lacks context about when to use it vs alternatives or error scenarios (e.g., 'image not available for mixture compounds'). Minimum actionable description is 50-200 chars; this is at lower end.
Parameter 'outcomeFilter' in pubchem_get_bioactivity uses enum ['active', 'inactive', 'all'] but description suggests 'all' is default behavior, the enum should clarify whether 'all' returns all outcomes or if null/omission achieves the same. Currently LLM must infer this distinction.