AutoML and Statistical Analysis Platform with Model Context Protocol (MCP) Integration. Provides tools for data cleaning, statistical analysis, machine learning training, and job management.
The AutoML Stat MCP server has 10 tools with varying quality. Most tools have descriptions (8/10 non-empty), but descriptions are verbose and often include example values that LLMs may reuse literally. Tool naming is mostly action-verb based but some names are generic (smart_analyze, handle_missing_values). Parameter schemas are present but inconsistently structured. No output schemas are documented. Error handling guidance is absent. Security considerations for file paths and dataset isolation are underdeveloped. The server uses HTTP+SSE transport (current standard) but lacks tool annotations and modern MCP patterns like error classification or recovery guidance.
Cancel a pending or running training job. Cannot cancel jobs that are already completed or failed.
🔄 Convert a column to binary (0/1) format. Essential for propensity score analysis and many ML algorithms that require binary treatment indicators.
Delete a registered dataset. Note: This only removes the registration, not the file in MinIO.
🏷️ Encode categorical column to numeric. Methods: label (assign integers 0, 1, 2, ... to each category), onehot (create binary columns for each category), ordinal (like label but preserves order).
Get the status of a training job. Use this to check if training is complete.
🔧 Handle missing values in dataset. Strategies: auto (Numeric → median, Categorical → mode), mean (fill with column mean), median (fill with column median), mode (fill with most frequent value), constant (fill with specified value), drop_rows (remove rows with missing values), drop_columns (remove columns with any missing values).
No output schemas documented. Tools return responses but LLMs cannot plan downstream tool calls without knowing the response structure (e.g., what fields register_dataset returns, what smart_analyze produces). This violates the schemas & output pattern.
Generic tool names reduce clarity. 'handle_missing_values' is vague (does it impute, drop, or flag?). 'smart_analyze' is extremely vague and masks multiple concerns. Rename to action_verb + specific object: 'fill_missing_values', 'drop_missing_rows', 'compute_statistics_and_correlations'.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 56 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
List all datasets registered by the user.
List all training jobs for the user. Returns jobs in all states (pending, running, completed, failed).
Register a CSV dataset from MinIO for use in training. The file must already exist in MinIO. This validates the file and registers it with the AutoML service.
🎯 Smart Analyze - Complete data analysis in one call. Automatically resolves file paths (host → container), gets quick statistics, generates Table One (if group_column provided), analyzes correlations (if requested), and compiles results into a summary.
Descriptions contain example values (e.g., '/data/sample_data/file.csv', {"200": 0, "400": 1}) that LLMs may reuse literally in real calls, causing failures or data corruption. Replace examples with constraints: use enum/pattern for valid formats, describe requirements in prose without concrete examples.
No error handling guidance. Tools like delete_dataset and cancel_job are destructive but lack recovery instructions. No distinction between retryable errors (network timeout) vs user-fixable (invalid dataset_id) vs fatal. Error responses should tell LLMs: 'Dataset not found. Call list_datasets() to see available IDs.'
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). The MCP spec supports declaring tool risk profiles, list_datasets is read-only, delete_dataset is destructive, cancel_job is destructive. These annotations inform agent planning and guardrails.
Parameter descriptions lack format/constraint details. E.g., 'strategy' enum is documented but no explanation of when to use 'auto' vs 'median' vs 'drop_rows'. 'user_id' defaults to 'default', is this secure? How do agents know whether to pass their user ID or the default?
Destructive tools lack confirmation or dry-run support. Agents make mistakes, delete_dataset and cancel_job should support a dry-run mode or require explicit confirmation to prevent irreversible actions.
Implicit response filtering and chaining. Tools like register_dataset and list_datasets do not document what fields are returned (name, id, size, created_at?). Without clear response fields, agents cannot extract IDs for downstream calls (e.g., get_job_status needs job_id returned from a prior call).
'user_id' is required on nearly every tool but defaults to 'default'. This is a security anti-pattern, agents may forget to pass user_id, defaulting to a shared namespace. Require user_id explicitly or implement server-side user resolution via auth headers.
Path traversal risk. CSV file paths like '/data/sample_data/file.csv' are agent-provided strings with no validation visible. Sanitize against directory traversal (../../../etc/passwd). Validate that paths stay within allowed data directory.
'smart_analyze' combines multiple concerns: file stats, stratified analysis, correlation analysis, and report generation. This violates the 'one tool, one job' principle. Split into: compute_statistics, compute_correlations, generate_table_one, generate_report. Agents can compose these in order.