A data processing and analysis MCP server with tools for loading, cleaning, validating, profiling, and analyzing datasets using pandas, polars, and topological data analysis.
Data-Forge MCP provides 15 tools with explicit schemas and descriptions visible in src/server.py and src/session_manager.py. However, critical gaps significantly reduce quality: (1) Most tool descriptions lack specificity about when to use each tool vs alternatives (e.g., generate_chart vs generate_map vs scan_semantic_voids all visualize data but lack comparative context). (2) Parameter descriptions are present but often minimal, many lack format constraints, examples of expected input, or guidance on invalid inputs. (3) Output schemas are documented as return types in docstrings but lack structured JSON Schema definitions. (4) Error handling is generic (try/except returning error strings) without recovery guidance or error categorization. (5) No tool annotations present (readOnlyHint, destructiveHint, idempotentHint). (6) Several tools accept complex nested inputs (operations in clean_dataset, schema in validate_dataset) where the structure is described informally rather than formally schematized. (7) The scan_semantic_voids tool has a context-conditional signature that creates two different tool registrations based on environment variable, complicating discoverability. Average per-tool definition score: 52/100. Most tools score 45-60 due to present but minimalist descriptions and incomplete parameter guidance.
Applies a sequence of Pyjanitor cleaning functions to the dataset.
Returns the System Prompt / SOP for the Data-Forge Agent. Use this to understand how to operate the server's tools effectively.
Extracts time-series signals (features) using tsfresh.
Extracts tables from a web URL using pandas (requires 'lxml' or 'html5lib').
Generates a chart from the dataset and saves it as an image.
Generates a geospatial map (scatter plot) from a dataset.
Returns the schema and summary (df.info()) of a loaded dataset.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) present. Tools with side effects (load_data, clean_dataset, generate_chart, generate_map, start_explorer, load_hf_dataset) should be marked as destructiveHint or non-idempotent. Read-only tools (get_dataset_info, list_active_datasets, validate_dataset, get_dataset_profile, run_sql_query, extract_signals, extract_tables) should declare readOnlyHint:true.
Complex parameter schemas (clean_dataset.operations, validate_dataset.schema) use informal Union or object types with example-based documentation instead of formal JSON Schema. LLM must reverse-engineer valid structure from code examples. validate_dataset.schema lacks specification of required/optional keys within the schema object.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 61 | 2026-07-28+ | v2 |
Generates a statistical profile of the dataset using YData Profiling. Returns a JSON summary of key insights (alerts, variable list).
Lists all currently loaded datasets and their IDs.
Loads a dataset file (CSV/Parquet) into the server's working memory.
Loads a dataset from the Hugging Face Hub (requires 'datasets' library).
Executes a SQL query on your datasets using DuckDB. Use this to filter, aggregate, join, or reshape data. You can reference datasets by their ID (e.g., 'SELECT * FROM df_a1b2'). If you provide a 'dataset_id' argument, you can refer to it as table 'this'.
Performs Topological Data Analysis to find "semantic voids" or gaps in the dataset's text column. Useful for identifying missing research topics, unaddressed customer complaints, or concept holes. Generates a persistence barcode and 3D manifold plot.
Launches an interactive D-Tale explorer for the dataset.
Validates the dataset against a provided schema using Pandera.
Output schemas are documented as prose in docstrings (e.g., 'Returns JSON summary', 'Returns image path') but not formally specified. LLM does not know structure of get_dataset_profile output (what keys in 'alerts'? What is 'variable list' structure?). No pagination guidance for tools returning lists.
Error handling is generic (all tools use try/except with basic error strings like 'Error loading data: {str(e)}'). No error categorization, recovery guidance, or actionable next steps. Stack traces are returned as strings, useless to LLM. Pattern: recovery-guide not implemented.
scan_semantic_voids has a conditional signature (two @mcp.tool() registrations based on DATA_FORGE_NO_CONTEXT env var). This creates stateful tool registration, violates stateless request handling. Protocol should not depend on startup environment for tool availability. Tool should have a single, consistent signature.
Many parameter descriptions are minimal or missing action context. Example: extract_signals.id_column is 'optional' but description does not explain what happens if omitted (auto-detect? single series assumed?). generate_chart.x and generate_chart.y lack constraints on valid column names or type compatibility with chart_type.
start_explorer may not be compatible with STDIO-only transport. It 'launches an interactive D-Tale explorer', implies UI/server startup. STDIO is stateless, non-interactive, and cannot serve HTTP or WebSocket. Tool is included but may be non-functional in the deployment model.
No distinction between similar visualization tools. generate_chart, generate_map, and scan_semantic_voids all produce visual outputs. Descriptions do not clarify when to use each (scatter plot vs geographic map vs topological manifold). LLM may select wrong tool.
No idempotency guarantees documented. load_data, load_hf_dataset may have side effects (memory allocation). If called twice with same args, does it create duplicate datasets or reuse? clean_dataset may or may not be idempotent (does it modify in-place or return new dataset?). Agents retry on ambiguous failures, need idempotency guarantees or clear non-idempotent declarations.