MCP server for molecular docking operations, including fetching PDB structures, converting SMILES to PDBQT format, and running molecular docking simulations using smina
The Smina Docking Server provides three molecular biology tools with basic input schemas and descriptions, but suffers from significant quality gaps. All three tools have registered schemas and non-empty descriptions, placing them above the baseline. However, descriptions are relatively brief (40-80 chars), parameters lack detailed constraints and validation guidance, and output schemas are only partially documented in docstrings rather than formally. The tools follow a clear verb_noun pattern (dock_ligand, fetch_pdb, smiles_to_pdbqt) which is positive, but parameter documentation is minimal and error handling lacks recovery guidance. Most critically, no tool descriptions explain prerequisites, dependencies, or when to call each tool instead of alternatives, critical for multi-step molecular workflows.
Dock a ligand to a receptor using smina.
Fetch a PDB structure from the RCSB PDB database.
Convert a SMILES string to PDBQT format using OpenBabel.
Parameter descriptions lack detail and constraints. 'center_x', 'center_y', 'center_z' have minimal descriptions; no guidance on units (Angstroms? meters?), valid ranges (0-1000? negative allowed?), or typical values. 'exhaustiveness' mentions it's optional but not what valid range is (1-1000? 1-16?).
Output schemas documented only in docstrings, not formally in tool registration. LLMs and client tools cannot introspect return types programmatically. 'docked_ligand' is documented as a 'str' but should specify format (raw PDB? partial?). 'all_scores' should specify if it's a list of floats.
Error responses return raw exception strings ('Error during docking: <str(e)>') without recovery guidance. No indication whether the error is retryable, requires user intervention, or is fatal. No suggested next steps (e.g., 'check ligand format' or 'try lower exhaustiveness').
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 56 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 48 | - | v1 |
Tool descriptions do not explain when to use each tool or dependencies between them. A user does not know whether to call fetch_pdb first, or what the workflow is. No guidance like 'Call fetch_pdb to get receptor_pdb, then smiles_to_pdbqt to prepare ligand, then dock_ligand.'
No input validation or type coercion for numeric parameters. If an LLM passes 'center_x' as a string '10.5' instead of float, the subprocess call may fail silently or with cryptic error. Validation should happen early with clear feedback.
dock_ligand returns 'score' as the first score if scores list is non-empty, else None. No documentation of what happens if docking produces zero poses or log parsing fails. Silent None return invites downstream errors.
No idempotency guarantees. If dock_ligand is retried with identical inputs, does it produce identical output? If temporary files or subprocess state cause variance, retries risk duplicate/conflicting docking results without agent awareness.