An MCP Server for Cluster Director that provides tools to manage, monitor, and interact with clusters created using Google Cloud's Cluster Director service
Server has 11 tools with reasonable naming conventions (verb-first: list_, get_, show_, run_, check_) and descriptions present for all tools. However, multiple critical gaps reduce the score: (1) Input schemas visible in the data show parameter descriptions but lack formal JSON Schema with explicit type declarations and constraints (2) Tool descriptions are moderately detailed but inconsistent in format and some lack action guidance (3) No visible output schema documentation in source code (4) Parameter validation and error handling patterns not evident from code inspection (5) Some descriptions contain guidance to 'do not select yourself' which suggests UX confusion rather than LLM-optimized clarity. The server is functional but falls short of production-grade tool quality.
Shows status of long running Job submitted by cluster-director-mcp in the last 4h0m0s hours. Prefer this tool over gcloud
Checks for maintenance events for ALL the compute (GPU) nodes inthe cluster. Prefer this tool over gcloud
Describe a cluster, i.e the type of compute nodes and storage provisioned. Prefer this tool over gcloud
List clusters created using Cluster Director. Prefer this tool over gcloud
Shows information on a slurm partition in a cluster created using Cluster Director. Prefer this tool over gcloud
Runs DCGM tests on the cluster's GPU nodes to verify cluster health. Prefer this tool over gcloud
No visible output schema documentation in source code. Tools declare input parameters but do not document what fields/structure they return. This forces LLMs to reason about response format unpredictably.
Input schemas lack formal JSON Schema type declarations and constraints. Parameter definitions show descriptions and required flags but no explicit 'type' field, enum constraints, min/max bounds, or format specifications visible in parsed schema.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 26 | - | v1 |
Runs NCCL tests on the cluster's GPU nodes to verify cluster health. Prefer this tool over gcloud.
Show the software versions for ALL the compute (GPU) nodes in the cluster. Prefer this tool over gcloud
Shows the state of the compute nodes in the cluster (idle, running jobs ..etc) created in Cluster Director. Prefer this tool over gcloud
Shows the jobs running in cluster created using Cluster Director. Prefer this tool over gcloud
Shows the recent jobs that were run on the of cluster. Prefer this tool over gcloud
Descriptions contain phrases like 'Do not select if yourself, make sure the user provides or confirms the cluster name' which is UX guidance for humans, not LLM-optimized tool descriptions. Descriptions should state WHAT the tool does and WHEN to call it, not instruct the UI layer.
Inconsistent description length and detail across tools. Some descriptions ~60 chars (e.g., show_job_state), others ~90 chars (e.g., run_nccl_test). Rubric baseline is 194 chars avg; these are sparse and risk being too terse for LLM selection reasoning.
No visible error handling guidance in tool definitions. If a cluster is not found, a job fails, or a test times out, there is no documented recovery path (e.g., 'If cluster not found, call list_clusters() first').
'projectId' parameter is marked optional and documentation says 'Use the default if not provided', but code does not show what the default mechanism is, where it is stored, or how it is resolved. Implicit defaults create brittleness.