Run reproducible bioinformatics pipelines from an AI assistant over MCP, with verifiable provenance
FlowProof MCP demonstrates solid definition quality with clear, descriptive tool names following verb_noun conventions and well-structured input schemas. All 6 tools are explicitly registered with descriptions and typed parameters. However, there are critical gaps in output schema documentation, parameter constraints, and error handling guidance. The descriptions are good (120-180 chars, above the 34 char p10 baseline) and parameters are consistently annotated. The main weaknesses are: (1) no documented output schemas for any tool, violating the pattern:tool requirement; (2) missing enum constraints on string parameters like pipeline_id; (3) no error recovery guidance; (4) incomplete composition chain documentation (e.g., get_results should state what fields are returned and how they chain to downstream tools). The server uses FastMCP with proper tool decorators and Pydantic Field annotations, which is correct for the framework. Tool naming is excellent (list_pipelines, describe_pipeline, run_pipeline, get_run_status, get_results, get_provenance), all start with action verbs and clearly signal intent.
Return the full manifest for one pipeline: its required inputs, tunable parameters, container image, and output file patterns. Call this before run_pipeline to learn exactly which inputs and params it expects.
Get the Workflow Run RO-Crate provenance for a run: workflow version, container digests, tool versions, parameters, and input/output SHA-256 checksums, the full record needed to reproduce the run byte-for-byte.
Get the outputs of a run: each output file with its SHA-256 checksum and size in bytes, plus the run status. Use the checksums to verify results independently.
Get the current status of a run (queued, running, succeeded, or failed) by its run_id, as returned by run_pipeline.
List the available bioinformatics pipelines. Returns each pipeline's id, a short description, and its read_type (short-read or long-read). Call this first to discover what can be run, then describe_pipeline for details.
No output schemas documented for any of the 6 tools. The code shows return statements but LLMs cannot see the response structure, field names, types, or nesting. This violates pattern:tool and pattern:response-shaper, forcing LLMs to infer output shapes from descriptions alone.
Missing error handling and recovery guidance. The code has validation logic (e.g., 'hosted_runnable' check, file path restrictions in hosted mode) but these errors are not documented in tool descriptions. LLMs do not know what errors are possible, whether they are retryable, or how to recover. For example, run_pipeline raises ValueError for hosted mode restrictions, but the description does not mention this condition.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 62 | 2026-07-28+ | v2 |
Start a pipeline run and return a run_id to track it. Provide inputs (a map of input name to file path or value) and params (a map of parameter name to value), as described by describe_pipeline. The run executes asynchronously; poll get_run_status, then fetch get_results and get_provenance when it completes.
String parameters lack enum constraints. pipeline_id is a free-form string with no validation hint, LLMs may hallucinate invalid pipeline names. The description should state 'Valid values: see list_pipelines output' or provide an enum constraint in the schema.
Composition chain documentation is incomplete. The tool descriptions reference related tools (e.g., run_pipeline says 'poll get_run_status, then fetch get_results') but do not specify the exact field names needed to chain calls. For example, 'run_pipeline returns run_id; pass this to get_run_status' is explicit, but 'describe_pipeline output includes inputs and params' lacks structure detail. This forces LLMs to reason about data flow instead of following clear field references.
Inputs and params parameters in run_pipeline are dict[str, str] | None but lack format guidance. The description states 'Map of input name to file path or value, per describe_pipeline' but does not clarify: (1) which values are file paths (absolute vs. relative?) and which are scalars; (2) how to know if an input is required; (3) what happens if a required input is omitted. This forces the LLM to call describe_pipeline every time to understand input structure, increasing latency and token consumption.