MCP server for Kueue job scheduling support and debugging on OpenShift clusters
The server defines 2 tools, but only 1 is actually registered and functional (collect_must_gather_ocp_kueue; cluster_queue_status). Three additional tool definitions exist in code (explain_job_scheduling, explain_local_queue, cluster_queue_status) but only cluster_queue_status is registered in RegisterTools(). Tool descriptions are adequate but lack actionable detail for agent planning. Parameter descriptions are minimal and sometimes copy-pasted incorrectly (dest_dir descriptions copied into unrelated parameters). Schemas are present and properly typed, but parameter descriptions are insufficient for LLM reasoning. No pagination, output schema documentation, or error recovery guidance. No tool annotations (readOnlyHint, etc.) despite READ_ONLY risk classification visible in the spec.
Provides information about why no workloads are admitted in the ClusterQueue.
Runs "oc adm must-gather" to capture debugging information. Create a temporary directory and pass it using the --dest-dir option to store the output in a single place. oc adm must-gather can scoop up almost every artifact engineers or support need in a single shot: it exports the full YAML for all cluster-scoped and namespaced resources (Deployments, CRDs, Nodes, ClusterOperators, etc.); captures pod and container logs as well as systemd journal slices from each node to trace runtime crashes or OOMs; grabs API-server and OAuth audit logs for security or compliance forensics; collects kernel, cgroup, and other node sysinfo plus tuned and kubelet configs for performance tuning; optionally runs add-on scripts such as gather_network_logs to archive iptables/OVN flows and CNI pod logs, or gather_profiling_node to fetch 30-second CPU and heap pprof dumps from both kubelet and CRI-O for hotspot analysis; and, through plug-in images, can extend to operator-specific data like storage states or virtualization metrics, ensuring one reproducible tarball contains configuration, logs, network traces, performance profiles, and security audits for thorough offline debugging. Use "oc adm must-gather -h" for available options.
Dead tool definitions: Three tools (explain_job_scheduling, explain_local_queue, cluster_queue_status defined twice) are defined in main.go but only cluster_queue_status is registered in RegisterTools(). This creates confusion, agents cannot invoke explain_job_scheduling or the first explain_local_queue definition despite appearing in code.
Parameter description copy-paste errors: cluster_queue_status's only parameter 'cluster_queue_name' has description 'Name of the ClusterQueue' (generic), yet explain_local_queue's 'namespace' parameter incorrectly says 'Directory to write gathered data', clearly copied from mustGatherTool's dest_dir. This misleads LLMs about what parameter does what.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 46 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 33 | - | v1 |
Insufficient parameter descriptions: collect_must_gather_ocp_kueue's 'extra_args' parameter description is 'Additional arguments passed directly to oc adm must-gather', unclear what format, safety constraints, or examples. Parameter descriptions should state format/range/allowed values (rubric §C). LLM cannot reason about when to use extra_args without examples of valid arguments.
No output schema documentation: Both tools return unstructured text (mcp.NewToolResultText). No documentation of what the output contains, structure, or fields. LLMs cannot plan downstream steps without knowing return structure. Pattern baseline: 100% of A+ tools document return types.
No error recovery guidance: Both tools return mcp.NewToolResultError(err.Error()) with raw error strings like 'oc adm must-gather failed: <underlying error>: <stderr>'. These give LLMs no actionable next step. Pattern baseline: error responses must say 'If X, try Y'. No guidance on retryability, user-fixability, or alternatives.
Missing tool annotations: Schema declares READ_ONLY risk for both tools but no readOnlyHint/destructiveHint annotations present in tool definitions. mcp-go supports mcp.WithReadOnlyHint(), should be used to signal to agents these are safe repeated calls (pattern: tool-annotations).
Naming clarity: cluster_queue_status is not verb-first. Should be 'describe_cluster_queue' or 'get_cluster_queue_status' to match verb_noun convention (rubric baseline: 90% of A+ tools start with action verb). 'status' is vague, does it return state, diagnostics, capacity, or conditions? Describe vs Get signals different intent.
Incomplete parameter type documentation: collect_must_gather_ocp_kueue's 'dest_dir' parameter has no constraint documentation. Should specify: is it a relative or absolute path? Does it need to exist? Size limits? 'Extra_args' is an array of strings but no guidance on escaping, special characters, or dangerous flags to avoid. Rubric §C: 'describe expected format, range, and allowed values directly in parameter description'.
No idempotence or side-effect clarity: collect_must_gather_ocp_kueue calls 'oc adm must-gather' which is a I/O-heavy operation (creates a tarball, writes to dest_dir). Description does not state if calling twice with same parameters is safe or if it overwrites/appends. Pattern: tools should declare side effects so agents know if retries are safe.