An MCP server for machine learning model training, dataset management, and AutoML capabilities
This server has significant quality gaps across naming, descriptions, schemas, and error handling. While tool definitions are explicitly visible in main.py, they lack the rigor expected of production-grade agent tooling. Tool names are reasonably clear (upload_dataset, train_model, clean_dataset) but descriptions are minimal (10-70 chars, well below the 194-char baseline for A+ tools). Input schemas are present but lack proper typing metadata (no 'type' fields visible in JSON Schema format, only parameter names and descriptions in docstrings). Parameter descriptions are present but generic, they state WHAT a parameter is named, not WHY an agent should choose a particular value or what constraints apply. Output schemas are completely undocumented: agents have no way to know what structure train_model or clean_dataset return. Error handling is weak, tools return free-text error strings rather than structured error objects with recovery guidance. The server lacks critical security patterns (no credential isolation, no permission gates, no audit trails). No tool demonstrates confirmation-request or dry-run patterns despite several tools performing irreversible operations (train_model modifies model_cache, clean_dataset modifies the dataset in-place). Parameter interdependencies (e.g., model_type=classification determines valid model_name values) are documented in code but not exposed to the agent. This is a D-grade server, better than completely broken, but not production-ready.
Clean the dataset: - Impute missing values (mean for numeric, mode for categorical) - Remove duplicate rows - Encode categorical variables (optional)
Downloads a dataset from Kaggle and loads the first CSV file into the dataset cache.
Preview a few rows of the dataset.
Train a model from the uploaded dataset.
Upload a dataset and cache it for future use.
Generate histograms for each numeric column in the dataset.
No output schemas documented for any tool. Agents cannot infer what structure train_model, clean_dataset, or visualize_data_distribution return, forcing them to make assumptions or fail on type mismatches.
Parameter enums are missing. model_type should be enum: ['classification', 'regression']. model_name valid values (logistic_regression, random_forest, svm, knn, decision_tree) are hardcoded in implementation but not exposed to agent, creating an invisible dependency.
Descriptions are too short and lack LLM-optimized context. Baseline is 194 chars; most tools are 35-115 chars. They state WHAT but not WHEN/WHY or prerequisites. For example, download_kaggle_dataset omits the requirement for Kaggle API credentials.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 39 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Irreversible operations (train_model, clean_dataset) lack confirmation or dry-run support. clean_dataset modifies dataset_cache in-place with no undo mechanism. An agent calling clean_dataset then realizing it needed the original data cannot recover it.
Error handling is unstructured. Tools return free-text error strings (e.g., 'Dataset not found') with no recovery guidance. Pattern requires errors to be categorized as retryable/user-fixable/fatal and include actionable next steps.
Parameter interdependencies are not documented. model_type determines which model_name values are valid (e.g., KNN is classification-only), but this constraint is not described to the agent. Agents may pass invalid combinations and fail.
No numeric constraints (min/max) on parameters. rows parameter in preview_dataset has no bounds, agent could request 1 billion rows. model_name and other string params accept any input; validation happens at runtime, not schema-level.
Destructive operations not marked. clean_dataset modifies the cached dataset in-place; train_model modifies model_cache. Agents need to know these have side effects. Pattern requires destructiveHint tool annotations.
Credential isolation missing. download_kaggle_dataset relies on Kaggle API credentials (code: api.authenticate()), but no documentation of where these are stored or how to inject them securely. Pattern requires server-side secret injection, not agent-side credential passing.