BreezeML MCP server: lets AI agents (Claude, GPT, etc.) train, compare, explain, export, and deploy ML models through the Model Context Protocol.
BreezeML MCP server has 11 well-intentioned tools with complete parameter schemas and clear descriptions. However, there are significant gaps in output schema documentation, error handling guidance, and parameter constraint specification. Tool names are action-oriented and appropriate (inspect_data, audit, train, compare, predict, explain, report, model_card, export, deploy, save), but lack verbosity that would make them self-documenting at a glance. Descriptions are present for all tools (averaging ~150 chars) and reasonably detailed, but several lack guidance on when to use them vs. similar tools or what to do on failure. Parameter descriptions are adequate but often lack format constraints, ranges, or validation rules. No output schemas are documented, LLMs cannot see what fields to expect from responses. Error handling is minimal: tools raise exceptions with messages but provide no recovery guidance or error categorization.
Audit a CSV for data-quality problems and target leakage before training: ID columns, constants, duplicates, label noise, and features that predict the target suspiciously well on their own.
Benchmark all built-in models on a CSV and return a ranked leaderboard.
Write a complete FastAPI + Docker serving directory for a trained model.
Plain-English explanation of a trained model's pipeline decisions.
Export a trained model as a standalone scikit-learn script with zero breezeml imports.
Profile a CSV: rows, features, missing values, class balance.
Generate a markdown model card. Optionally write it to a file.
No output schemas documented. LLMs cannot see what fields to expect from tool responses. tool_train() returns {model_id, task, estimator, report, decisions}, but this is invisible in the tool definition. tool_predict() returns {predictions: list}, tool_report() returns rep.to_dict(), etc., all inferred from code, not declared in schema.
Parameter descriptions lack format constraints and ranges. 'csv_path' is described as 'Path to the CSV file...' but does not specify file existence validation, encoding expectations, or size limits. 'task' accepts 'auto', 'classification', or 'regression' but the description does not list these enums formally, LLM must infer from text.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 58 | 2026-07-28+ | v2 |
Run predictions with a trained model. records_json is a JSON list of feature dicts.
Run the full honesty gauntlet on a trained model and return a single SHIP / WARN / STOP verdict: cross-validated performance vs a naive baseline, data audit (leakage/quality), class-imbalance severity, and an optional fairness check. Agents SHOULD call this and confirm a SHIP verdict before calling deploy() or export().
Persist a trained model to a .joblib file.
Train a model on a CSV. Returns model_id, metrics, and an explanation of every pipeline decision.
Error handling provides no recovery guidance. _get(model_id) raises ValueError with message 'Unknown model_id...', but LLM does not know what to do next: retry? list available models? call train() first? Similarly, _load_df() raises FileNotFoundError, but description does not hint at how to resolve it.
Tool descriptions do not clarify when to use them vs. alternatives. 'explain' and 'report' both analyze a model, but LLM cannot determine from descriptions which to call first or whether to call both. 'model_card' vs 'report' distinction is unclear.
records_json parameter in predict() is a string, not a structured array. LLM must manually serialize to JSON, which is error-prone and increases hallucination risk. No format example provided, no schema for the inner objects.
No idempotency markers. train(), export(), and deploy() are destructive (write files, register models), but tool definitions do not declare idempotentHint: false. Agents may retry on timeout without knowing side effects.
Session-based model storage uses in-memory _MODELS dict. Models are lost when server restarts. No persistence mechanism, no warning in tool descriptions that models are ephemeral.