A Model Context Protocol server for data science and exploratory data analysis with support for CSV loading, Python script execution, and interactive data exploration prompts
This server has significant definition quality gaps. Both tools have descriptions, but they are verbose (500+ chars), include prohibited example values, and lack the crisp structure needed for LLM action selection. Input schemas are present but incomplete: 'load_csv' omits the required field explicitly, and 'run_script' lacks type information for the 'save_to_memory' array items. Parameter descriptions are present but vague, they explain the obvious without actionable format constraints or dependency hints. Output schemas are entirely undocumented. Error handling guidance is absent. The tool names are reasonably clear (verb_noun), but the descriptions meander through usage notes and prohibited actions rather than stating WHAT, WHEN, and WHY concisely. This pattern suggests the descriptions were written for human reference, not LLM-optimized tool selection.
Load CSV File Tool Purpose: Load a local CSV file into a DataFrame. Usage Notes: • If a df_name is not provided, the tool will automatically assign names sequentially as df_1, df_2, and so on.
Python Script Execution Tool Purpose: Execute Python scripts for specific data analytics tasks. Allowed Actions 1. Print Results: Output will be displayed as the script's stdout. 2. [Optional] Save DataFrames: Store DataFrames in memory for future use by specifying a save_to_memory name. Prohibited Actions 1. Overwriting Original DataFrames: Do not modify existing DataFrames to preserve their integrity for future tasks. 2. Creating Charts: Chart generation is not permitted.
Tool descriptions are verbose (500+ chars) and include prohibited example values ('df_1, df_2'). LLMs latch onto example values and pass them literally in real calls, causing failures. Descriptions should be 10 - 200 chars, crisp, and example-free.
Output schemas are entirely undocumented. LLMs cannot plan downstream calls or extract the right data without knowing what fields to expect. Both tools must document their return structure (what fields, types, and meaning).
Input schema for 'load_csv' omits explicit 'required' declaration for 'csv_path' (should be true). The schema marks 'df_name' as required=false, which is correct, but the structure is incomplete and inconsistent with JSON Schema conventions.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 43 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 37 | - | v1 |
'run_script' schema lists 'save_to_memory' as array of strings, but lacks minItems, maxItems, or description for array items. An LLM cannot infer whether to pass one, two, or ten dataframe names, or what naming rules apply.
No error handling guidance. If 'load_csv' receives a non-existent path, or 'run_script' raises a Python exception, the error response does not guide the LLM on what to do next (retry, ask the user, fallback, etc.). Error responses must categorize as retryable, user-fixable, or fatal.
'run_script' description emphasizes 'Prohibited Actions' (do not modify DataFrames, no chart generation) as a warning. These constraints should be enforced server-side, not delegated to LLM compliance. If the tool truly forbids these actions, implement validation; if not, document what IS allowed instead of threats.
Parameter descriptions lack actionable format constraints. 'csv_path' should specify: is it an absolute path or relative to a working directory? Does it accept URLs? What file sizes are supported? 'save_to_memory' names: are they case-sensitive? Do they match Python variable naming rules?
Tool names 'load_csv' and 'run_script' are action verbs, but 'run_script' is generic. 'Execute_python_script' or 'run_analytics_script' would disambiguate from other potential script runners. Generic names force LLMs to reason through descriptions rather than inferring intent from the name alone.