A Model Context Protocol (MCP) server for managing Apache Beam pipelines across different runners (Dataflow, Spark, Flink, Direct)
The Apache Beam MCP server has 12 tools with moderate to poor definition quality. While tool names generally follow verb_noun conventions and schemas are visible in the registry, there are significant gaps in descriptions, parameter documentation, and output schema specification. Many tools have generic or minimal descriptions (e.g., 'List available runners' is only 24 chars). Parameter descriptions are sparse, the create_job tool accepts 'pipeline_options' as a free-form object with only a brief description 'Runner-specific pipeline options', providing no guidance on structure, constraints, or valid keys. Output schemas are not documented anywhere in the provided source, making it difficult for LLMs to reason about chaining tools. The text-tokenizer tool lacks description of its return format (is it an array of tokens? an object with metadata?). Error handling is minimal, no evidence of recovery guidance, categorization of errors, or actionable error messages. Several tools (particularly the tools_router CRUD operations) operate on abstract 'tool definitions' without clear schema documentation. The server lacks tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite having a clear risk matrix in the metadata.
Counts words in a text corpus using Apache Beam
Cancel a running job
Create a new pipeline job
Create a new tool with the provided definition
Delete a tool
Get job details
Get job metrics
Get a specific tool by ID
Output schemas are not documented for any tool. LLMs cannot reason about the structure of returned data, making tool chaining and downstream processing error-prone. For example, 'get_job' returns data but the response structure is not specified, does it return job_id, status, metrics, logs, or something else?
Descriptions are too generic or short. 'List available runners' (24 chars) does not explain when to call this vs other discovery tools or what the runner names mean to agents. 'Get job metrics' (15 chars) omits what metrics are returned and whether they update in real-time. Baseline for A+ tools is 50 - 200 chars with context on WHAT, WHEN, and WHY.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 50 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 50 | 1.0+ | v1 |
List available runners
List all available tools, optionally filtered by type
Tokenizes text into words or subwords
Update an existing tool
The 'pipeline_options' parameter in create_job is a free-form object with no schema constraints, type hints, or examples of valid keys. This invites hallucinated options from LLMs. Should be a schema object with documented keys (e.g., num_workers, machine_type, disk_size_gb) or reference documentation.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) are present, despite the metadata clearly marking risk levels (READ_ONLY, WRITE, DESTRUCTIVE). Tools like 'delete_tool' should declare destructiveHint:true; 'get_job' should declare readOnlyHint:true. This guidance helps agents reason about side effects and retry safety.
Destructive operations (cancel_job, delete_tool) have no confirmation pattern or dry-run mode. An agent invoking 'delete_tool' with an incorrect tool_id could permanently remove it. Should implement a confirmation step or require explicit confirmation in the tool description.
Error handling is not visible in the provided source. No evidence of actionable error messages, error categorization (retryable vs fatal), or recovery guidance. For example, what happens if create_job fails because the code_path does not exist? Does it return 'File not found' or a stack trace?
Tool definitions for list_tools, get_tool, create_tool, update_tool, delete_tool operate on generic 'tool' and 'tool_definition' objects with no schema documentation. What fields must a tool_definition contain? Are there required vs optional fields? This is too vague for reliable agent usage.
Parameters lack type clarity in descriptions. For example, 'job_type' in create_job accepts 'BATCH or STREAMING' but the parameter schema does not show this as an enum. 'runner' in beam-wordcount has an enum, but other tools with constrained values do not use enums, forcing LLMs to guess at valid options.
No pagination parameters (limit, offset, page_size, cursor) on list_* tools. If the runner or tool list grows large, agents cannot control result size or iterate through pages, risking context window exhaustion.
The 'runner' parameter in beam-wordcount uses an enum correctly, but the description is missing for the parameter itself, only values are listed. Should explain what a runner is and when to choose each option.