A conversational MCP server interface for PyMOL protein visualization
MCPymol demonstrates solid tool definition quality with comprehensive descriptions and well-structured parameters. All 9 tools have clear, detailed descriptions (averaging 250+ characters) that explain WHAT the tool does, WHEN to use it, and important caveats. Input schemas are present for all tools with proper type definitions and parameter annotations. However, there are notable gaps: no output schema documentation, missing error handling guidance, no tool annotations (readOnlyHint/destructiveHint), and limited actionable error messages. The server correctly identifies which tools are read-only vs write operations via Risk labels, but this metadata is not exposed in MCP tool definitions. Tool composition is excellent, each tool has a single, clear responsibility and natural names (occupancy_view, altloc_view, etc.) that convey intent. Parameter naming is precise with type suffixes where needed (e.g., obj_name, validation_path). Descriptions are notably strong for a scientific tool, with domain-specific context (e.g., 'This is occupancy in the crystallographic sense, the fraction of unit cells...' in occupancy_view).
Show every altloc group at once, one colour per group, occupancies labelled. A multiconformer model says "at this site, these discrete alternatives, in these proportions". PyMOL shows one of them by default, which quietly turns a statement about heterogeneity into a single structure.
Lists the residues in contact across two selections, with distances and types. This is the numeric counterpart to ligand_view and interface_view: instead of drawing the interactions, it reports them — which residue pairs touch, how close they get, and whether the contact is a salt bridge, hydrogen bond, hydrophobic packing or pi-stacking. Use it to answer "what holds this ligand in the pocket" or "which residues form this interface". Classification uses heavy-atom distance criteria, since crystal structures usually have no hydrogens: salt bridge <= 4.
Colour and thicken a multi-state object by how much its states disagree. Per-residue RMS deviation across states, pushed into the B-factor column and spectrum-coloured blue (rigid) to red (variable). Spread is a description of how much the deposited members differ. It is NOT a calibrated uncertainty and NOT an error bar: the number of members and the refinement protocol both shape it, so a wider tube means these models disagree more, not that the true position is less well determined.
Report an MRC/CCP4 map's geometry, above all its voxel size. Reads only the 1024-byte header — the map data is never loaded, PyMOL is never touched, and nothing goes over the network. Voxel size is not stored in the header, it is derived, and the derivation has a trap: it is cella/m (the grid sampling), not cella/n (the stored extent). Those differ on any boxed or cropped map, and dividing by n gives a wrong answer of the right order of magnitude. The nominal value is only expected to be accurate to ±5–15% in the first place, which at 1.2 Å is a systematic stretch of every distance in the model. Anisotropy, cropping and axis permutation are all flagged loudly.
No output schema documentation. Tools return strings (error handling pattern or structured report), but the documentation does not specify what fields LLMs should expect in successful responses. contact_report and map_info likely return structured data that agents need to parse, but the schema is invisible.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) in MCP definitions. The server tracks Risk labels internally (WRITE, WRITE, READ_ONLY), but these are not exposed via MCP's tool annotations feature, limiting agent planning and safety.
Error handling returns unstructured strings ('Error: <exception>'). While the _report() wrapper catches PortError, ValueError, and OSError, error responses do not include recovery guidance or categorization. Agents cannot distinguish retryable errors from user-fixable input errors.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 69 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 23 | - | v1 |
Interpolate between states, refusing when interpolation is ill-posed. The value here is the refusal, not the interpolation — PyMOL morphs natively. A morph is only meaningful when every state shares a topology, so that atom i in state 1 is the same atom as atom i in state 2. That holds for a deformation-model ensemble (3DFlex, DynaMight). It does not hold for volumes reconstructed and modelled independently, and a morph across those animates a correspondence nobody established. Note that cmd.morph is Incentive-only. On open-source PyMOL the topology check still runs and reports; only the interpolation is unavailable.
Colour and scale a structure by per-atom crystallographic occupancy (q). Alternates at q<1 are visually de-emphasised in proportion, so a half-occupied conformer looks half-there instead of looking like a confident model. This is occupancy in the crystallographic sense — the fraction of unit cells in which an atom sits at this position. It is NOT the fraction of imaged particles containing a subunit; a model can be q=1.0 everywhere while the subunit is present in half the particles. Both are true and they are different questions, so this tool never reports the other sense.
Colour a model by per-residue Q-score parsed from its wwPDB validation report. No network, no map, and no computation — Q-scores are already published, and this reads the numbers rather than re-deriving them. Two things to expect from real data: entries deposited before the September 2023 validation rollout carry no Q-scores at all, which is the common case for older EM structures; and real Q-scores go negative, so the published 0–1 framing is the intended range and not the observed one.
Restore the B-factors a Wiggles view overwrote. occupancy_view, ensemble_spread_view and qscore_view all push their values into the B-factor column so PyMOL can spectrum-colour by them. The originals are stashed first; this puts them back in one pass.
Superposes two structures and colors the mobile one by per-residue shift. An RMSD alone tells you that something moved, not where. This superposes ``mobile`` onto ``target``, measures how far each residue's CA ended up from its counterpart, and colors the mobile structure blue (unchanged) through white to red (most shifted). The target is left as a grey reference cartoon. Best on two states of the same protein — apo vs holo, open vs closed, a mutant against wild type. Residues are paired by chain and residue number, falling back to residue number alone when the two use different chain IDs. Both structures have to be loaded first, and ``fetch_structure`` clears the session by default — so fetch the second one with ``replace=False`` or it will replace the first. If either entry is a multimer, compare single chains (``create`` one object per chain): superposing one dimer onto another fits the assembly rather than the fold, which inflates the RMSD dramatically. 4AKE against 1AKE gives 18.5 A as deposited dimers and 2.1 A chain-to-chain.
Parameter descriptions lack explicit constraint documentation in some cases. E.g., 'cutoff' in contact_report says 'default 4.0' but does not state the valid range (0 - ∞?) or the specific meaning in the context of heavy-atom distances. 'method' accepts 'super' or 'align' but this is documented as prose rather than an enum.
No pagination for potentially large result sets. contact_report defaults to 'max_pairs=40' with a note that omitted pairs are reported, but there is no mechanism to fetch additional results. Agents cannot paginate through large contact lists or handle >40-residue interactions elegantly.
STDIO transport only. The server is not remotely accessible and cannot be used by hosted MCP clients. This is a fundamental limitation for production deployment scenarios where the client runs on a cloud platform.