An MCP server that provides chemical compound search and retrieval capabilities from PubChem database using compound names, SMILES strings, molecular formulas, and CID identifiers.
The PubChem MCP server provides 4 tools with partial schema definitions and basic descriptions. All tools have readable names starting with action verbs (search_, get_), and each tool has a description present. However, there are critical gaps in schema completeness, parameter documentation, and error handling guidance. The descriptions are present but generic (averaging ~130 chars), lacking actionable context for LLM tool selection. While output structure (compound_to_dict) is well-defined internally, the tool response schema is not formally documented for LLM consumption. Error handling returns error dictionaries but does not guide the LLM on recovery steps. Parameter descriptions are present but lack constraints, ranges, or format specifications (e.g., no min/max for max_results, no guidance on SMILES format). The advanced search tool has a logical design issue: it documents a 'formula' parameter but the implementation doesn't validate that at least one search criterion is required before calling the API.
Fetch detailed information about a chemical compound using its PubChem CID. Args: cid: PubChem Compound ID (CID) Returns: Dictionary containing compound information
Perform an advanced search for compounds on PubChem. Args: name: Name of the chemical compound smiles: SMILES notation of the chemical compound formula: Molecular formula cid: PubChem Compound ID max_results: Maximum number of results to return (default: 5) Returns: List of dictionaries containing compound information
Search for chemical compounds on PubChem using a compound name. Args: name: Name of the chemical compound max_results: Maximum number of results to return (default: 5) Returns: List of dictionaries containing compound information
Search for chemical compounds on PubChem using a SMILES string. Args: smiles: SMILES notation of the chemical compound max_results: Maximum number of results to return (default: 5) Returns: List of dictionaries containing compound information
Parameter descriptions lack format constraints and ranges. 'max_results' is documented as 'Maximum number of results to return (default: 5)' but lacks min/max bounds, example valid range (1-100?), or guidance on performance implications of large values. 'name', 'smiles', and 'formula' parameters have no format specification, allowed character sets, or length limits. This invites LLMs to pass invalid inputs.
Tool descriptions are generic and lack LLM-selection guidance. Descriptions state WHAT the tool does but not WHEN to use it or how it differs from similar tools. For example, search_pubchem_by_name and search_pubchem_advanced both search by name, the descriptions don't explain when to call one vs the other. Missing: dependency hints (e.g., 'If you only have a compound structure, use search_pubchem_by_smiles first').
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 8 | - | v1 |
Error handling does not guide LLM recovery. When an error occurs (e.g., compound not found, invalid SMILES), the tool returns {'error': 'message'} but provides no actionable recovery guidance. Example: 'An error occurred while searching' tells the agent nothing about whether to retry, adjust parameters, or try a different search tool. Error messages should follow the pattern: 'CID 999999 not found. Try search_pubchem_by_name() or search_pubchem_by_smiles() instead.'
Output schema is not formally documented. While compound_to_dict() internally defines a rich structure (23 fields: cid, iupac_name, molecular_formula, etc.), the tool docstrings do not document the expected response schema for LLM consumption. LLMs cannot reliably extract fields like 'canonical_smiles', 'inchikey', or 'tpsa' without explicit schema declaration. This forces the LLM to parse responses via trial-and-error, increasing errors and token waste.
search_pubchem_advanced has a logical design flaw: parameters (name, smiles, formula, cid) are all marked optional, but the tool requires at least one to function. The implementation returns an error if none are provided, but the schema does not enforce this constraint. An LLM could call this tool with no parameters and receive a generic error. Fix: add a description note stating 'At least one of name, smiles, formula, or cid is required' and validate at the schema level or with a clearer error.
No pagination or result limit enforcement. While max_results defaults to 5, searches can return large result lists depending on PubChem database size. The tool description does not state an upper bound, performance implications, or guidance on handling large result sets. If an LLM searches for a common compound name with max_results=100, the response could exceed context limits. The tool should enforce a reasonable upper cap (e.g., max_results ≤ 50) and document it.