Agentic MCP server for Kubeflow - High-level semantic AI training orchestration on Kubernetes
The Kubeflow MCP server provides 33 tools with basic schemas and descriptions, but exhibits significant gaps in description quality, parameter documentation, and output schema clarity. Tool naming is generally clear and verb-driven, following conventions (get_, list_, create_, delete_). However, descriptions are often generic (e.g., 'Get information about the current Kubernetes cluster') and lack guidance on WHEN to use the tool or WHAT output to expect. Many parameters lack complete type information, some tools have parameters with descriptions but no explicit schema (e.g., 'config' in validate_training_config is typed as 'object' with no sub-schema). Output schemas are not documented, making it impossible for LLMs to plan multi-step workflows reliably. Error handling is absent from descriptions, and there is no guidance on which operations are destructive vs. safe to retry. The schema for tools like 'manage_checkpoints' is overly broad (action parameter accepts free-form strings instead of an enum: list|save|restore|delete). No tool annotations (readOnlyHint, destructiveHint, idempotentHint) are present despite 4 destructive operations (delete_training_job, delete_resource) and multiple reversible operations (suspend/resume). Security concerns: setup_hf_credentials accepts a 'token' parameter directly, which violates secret-injection patterns. Schema scores are capped at 30-40 for tools with incomplete parameter documentation (e.g., 'training_config' object with no schema definition).
Adapt a training script for Kubeflow execution
Analyze and provide insights on training job failures
Build a training runtime specification
Check if prerequisites are met for training
Create a training runtime
Delete a Kubernetes resource
Delete a training job
Estimate resource requirements for training
Missing output schema documentation for all 33 tools. LLMs cannot infer what fields to expect in responses, preventing reliable multi-step workflows and field extractions for downstream tool calls.
Credential exposed as tool parameter: setup_hf_credentials accepts 'token' directly. This violates secret-injection pattern and risks credential leakage in logs and traces.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). Four destructive tools (delete_training_job, delete_resource, etc.) and multiple reversible operations lack risk classification, preventing LLMs from reasoning about consequences.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 44 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 41 | - | v1 |
Fix permissions on PersistentVolumeClaim resources
Get algorithm parameters for a training framework
Get information about the current Kubernetes cluster
Get resource usage and availability in the cluster
Get events associated with a training job
Get the specification of a training job
Get details about a training runtime
Get available packages in a runtime
Get the specification of a training runtime
Get details about a specific training job
Get logs from a training job
Get progress information about a training job
List training jobs in a namespace
List available training runtimes
Manage training job checkpoints
Resume a suspended training job
Run a container-based training job
Setup HuggingFace credentials for model access
Setup NFS storage for training
Setup a training runtime with specified configuration
Setup storage for training jobs
Suspend a running training job
Submit a training job to Kubeflow
Validate training configuration
Wait for a training job to complete
Incomplete parameter schemas for object-typed parameters. Tools like validate_training_config, create_runtime, setup_training_runtime, and train accept 'config'/'runtime_spec'/'storage_config' as generic objects with no sub-schema definition, forcing LLMs to guess the structure.
manage_checkpoints uses a free-form 'action' string parameter instead of an enum constraint. Accepts 'list, save, restore, delete' per description, but no enum enforcement lets LLMs hallucinate invalid values.
Generic, uninformative descriptions. Many tools describe what they do (e.g., 'Get information about the current Kubernetes cluster') but omit WHEN to use them, WHAT output to expect, or dependencies. Average description length ~50 chars, below the 194-char baseline.
No error handling guidance. No tool description explains what to do if a job is not found, credentials are invalid, or resources are insufficient. LLMs receive no recovery path.
Missing required parameter descriptions. All tools appear to have parameter descriptions, but they are often minimal (e.g., 'Training configuration object' for a complex config param) and lack constraints, expected format, or examples of valid values.
No idempotency guidance. Tools like resume_training_job and suspend_training_job lack information on whether they are idempotent, critical for agents deciding whether to retry on failure.