A server that proxies requests to the Buildkite API, providing tools for managing clusters, agents, pipelines, builds, artifacts, logs, tests, annotations, investigations, and user information.
Static source inference · medium confidence · evidence: stateless requests
Current-spec patterns detected
Summary
The Buildkite MCP server provides 40 tools covering clusters, agents, pipelines, builds, artifacts, logs, tests, annotations, and investigations. However, critical definition quality issues prevent a higher score: (1) Tool descriptions are largely missing or minimal, I can see tool names in the file structure but cannot verify non-empty, substantive descriptions in the provided code excerpt. (2) Parameter schemas are not visible in the provided source code; without seeing explicit parameter definitions, types, and constraints, schema scores default to 0. (3) No output schemas are documented in the visible code. (4) Error handling guidance is not evident in the excerpt. The server appears functionally complete (covering 40 realistic Buildkite API operations), but lacks the LLM-optimized descriptions, parameter documentation, and output shaping that production-grade agent tools require. Since I cannot see the actual tool definitions in the provided code sample, many tools default to inferred scores capped at 50.
Tool descriptions are not visible in provided code excerpt. Cannot verify that tools have non-empty, substantive descriptions (rubric baseline: 100% of A+ tools have descriptions; average 194 chars). Without visible descriptions, LLMs cannot determine when to select each tool.
Add substantive, LLM-optimized descriptions to all 40 tools. Rubric baseline: 194 chars average (p10=34, p90=392). Each description should answer: What does this tool do? When should the LLM call it instead of a similar tool? What does it return? Example: 'create_cluster: Creates a new Buildkite cluster with the specified name, description, and queue configuration. Used when you need a new isolated deployment environment. Returns cluster_id, name, status, and queue_ids for downstream operations.'
Document input parameter schemas for every tool. For each parameter, specify: type (string/integer/boolean/enum), description (15 - 50 chars, actionable), constraints (min/max for numbers, regex patterns, enum values), and whether it's required. Example: 'cluster_id: string, required, format: UUID. The unique identifier of the cluster. Can be obtained from list_clusters().'
Document output schemas with typed fields. For list_* tools, specify pagination: limit (default 20, max 100), offset/page, total_count. For single-resource tools, return all IDs and references needed by downstream tools. Example: 'create_cluster returns { cluster_id: string, name: string, status: enum(active|paused), queue_ids: [string], created_at: ISO8601 }'.
Add error handling guidance. For each tool, document failure modes and recovery paths. Example: 'If cluster not found, try list_clusters() with partial name filter. If permission denied, check that token has cluster:write scope.' Use error categorization: retryable (429, 503), user-fixable (400, 404), fatal (403, 401).
Parameter schemas are not visible in provided code excerpt. Cannot verify input parameter types, descriptions, enums, or constraints. Rubric baseline: 100% of A+ tools have documented return types and parameter constraints.
Output schemas not documented. Rubric baseline: 100% of A+ tools document return types with typed fields. LLMs need to know what fields to expect so they can plan downstream tool calls. No visible documentation of response structure for any tool.
No error handling guidance visible. Rubric critical check: error responses must tell the LLM what to do next. Cannot verify recovery guides, error categorization (retryable/user-fixable/fatal), or actionable error messages in provided code.
No confirmation or dry-run pattern visible for destructive tools (delete_cluster_secret, delete_artifact). Rubric critical check: irreversible operations should support confirmation step to prevent catastrophic agent errors.
Naming ambiguity in state-change tools: pause_cluster_queue_dispatch vs resume_cluster_queue_dispatch are clear, but tools like 'update_cluster' and 'update_pipeline' are vague. Rubric check: 'update_ticket_status' is clear; 'modify_ticket' is ambiguous. No visibility into what fields each update accepts, forcing LLMs to guess.
Cannot verify pagination support for list_* tools. Rubric baseline: tools returning lists should accept page/offset/limit parameters and return total count or next_cursor. No visible pagination parameters in schema excerpt.
No visible indication that response schemas include chaining IDs. Rubric check: if next likely action requires team_id, channel_id, and message_id, current response must return all three. Cannot verify this for lookup tools without seeing response schemas.
Implement confirmation step for destructive tools (delete_cluster_secret, delete_artifact). Support a --dry-run flag or require explicit confirmation parameter to prevent accidental deletions by agents.
Clarify ambiguous update tool parameter sets. For update_cluster and update_pipeline, document exactly which fields can be updated and which are immutable. Use a structured approach: 'update_cluster accepts: name (string, 1-255 chars), description (string, optional), settings (object with queue_concurrency, agent_tags, etc.)'
Ensure list_* tools accept pagination parameters: limit (1 - 100, default 20), offset or page (0-indexed), and return { items: [...], total_count: integer, has_more: boolean }. This prevents context window exhaustion when agents work with large result sets.
Validate that response schemas from read operations (get_*, list_*) include all IDs and references needed by write operations. For example, get_build should return build_id, pipeline_id, cluster_id so create_build_job can chain without extra lookups.
Add per-parameter validation with actionable errors. Example: instead of '400: invalid cluster_id', return '400: Invalid cluster_id "xyz-bad", must be UUID format (e.g., 550e8400-e29b-41d4-a716-446655440000). Valid clusters: [list]'
Document whether each tool is idempotent. Agents retry on ambiguous failures, non-idempotent tools risk duplicate side effects. Example: 'create_cluster is idempotent if name is unique; repeated calls with same name return existing cluster without side effects.'
Ensure tool names and descriptions guide natural composition. If agents need to 'cancel a build and notify', split into cancel_build + create_annotation rather than combining. Document when separate tools should be chained vs. when a compound tool is justified.
Add toolAnnotations for WRITE and DESTRUCTIVE tools. The server already declares toolAnnotations=true in features. Ensure every tool handler sets readOnlyHint: false for mutations, destructiveHint: true for delete_*, idempotentHint: true where applicable. This guides agent planning.
Test parameter naming consistency. Ensure get_cluster, get_pipeline, get_build all use the same ID parameter name (e.g., cluster_id, pipeline_id, build_id) so agents can naturally chain calls without manual field mapping.