MCP Server for Boltz - Provides synchronous and asynchronous APIs for protein structure and affinity prediction using Boltz deep learning models
Boltz MCP provides 9 well-structured tools with clear naming conventions and reasonably detailed descriptions. All tools follow verb_noun naming patterns (get_, list_, cancel_, submit_). However, the server has significant gaps in schema documentation, error handling guidance, and parameter constraint specification. Tool descriptions are adequate (100-250 chars on average) but lack actionable detail on dependencies, prerequisites, and failure recovery paths. Input schemas are present and properly typed, but output schemas are undocumented. The server handles a specialized domain (protein structure prediction) well, but lacks the polish and defensive practices expected in production-grade tooling.
Cancel a running job.
Get log output from a running or completed job.
Get the results of a completed job.
Get the status of a submitted job.
List all submitted jobs.
Predict protein-ligand binding affinity and structure using Boltz (fast mode). Fast operation that completes in ~2-5 minutes with pre-computed MSA, ~8-15 minutes without. For screening multiple ligands against the same protein, compute MSA once and reuse via msa_path.
Generate protein structure predictions using Boltz (fast mode). Fast operation that completes in ~5-10 minutes. Use this for single structure prediction. Supports protein-only or protein-ligand complex prediction.
No output schemas documented. Tools return dictionaries with fields like 'status', 'error', and domain-specific results, but the expected structure is not formally specified. LLMs cannot plan downstream operations or validate responses without knowing field names and types.
Mutually exclusive parameters documented in descriptions only. For example, simple_structure_prediction has 'input_file' and 'sequence' marked as mutually exclusive, and 'ligand_smiles' as exclusive with 'input_file'. LLMs cannot parse prose constraints reliably. These should be enforced at the schema level (using oneOf or allOf + not) or returned as validation errors with clear guidance on which combination to use.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 65 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 31 | - | v1 |
Submit protein-ligand binding affinity prediction for background processing. Returns a job_id for tracking.
Submit protein structure prediction for background processing. This operation may take >10 minutes for complex sequences. Returns a job_id for tracking. Supports protein-only or protein-ligand complex prediction.
Error handling is minimal. The code catches exceptions and returns {'status': 'error', 'error': str(e)}, but provides no guidance on what the LLM should do next. For example, if 'sequence' validation fails or 'input_file' is not found, the error message should be actionable: 'Sequence is invalid: contains non-standard amino acids X. Try: ...'. Current errors are raw exception strings.
No pagination or result limiting documented. list_jobs returns all jobs; if thousands exist, this will exhaust context and degrade reasoning. The tool description should state: 'Returns at most N jobs; use list_jobs(status=...) to filter.'
Parameter constraints are under-specified. For example, 'accelerator' accepts 'gpu', 'cpu', or 'tpu', but this is stated only in prose, not as an enum. Similarly, 'output_format' should be enum('pdb', 'cif'). Free-form string parameters invite LLM hallucination (e.g. passing 'GPU' or 'gpu_cuda').
'tail' parameter in get_job_log has a non-standard semantic. Docstring says '0 for all', but most APIs treat 0 as invalid or no result. Should clarify: 'Number of lines to return. Use -1 for all lines.' Or use a separate boolean parameter 'include_all: bool' to avoid ambiguity.
Job management tools lack failure modes. What happens if a job_id does not exist? Does get_job_result return 'not found' or a stack trace? Does cancel_job fail if the job is already completed? These edge cases should be documented with expected response structures.
No distinction between retryable and fatal errors. If submit_structure_prediction fails due to network timeout, should the agent retry? If it fails due to an invalid sequence, should it suggest alternatives? Current error response does not categorize the failure type.