Static source inference · medium confidence · evidence: Streamable HTTP
Current-spec patterns detected
Summary
This MCP server exhibits pervasive definition quality gaps across all 11 tools. While tool names follow verb_noun conventions (load_data, analyze_data, etc.), descriptions are present but generic and lack actionable context for LLM selection. (2) Parameter descriptions exist but are extremely minimal (10-25 chars) and lack constraints, format hints, or LLM-actionable guidance (e.g., 'Format' for format parameter provides no enum, no examples, no constraint). (3) No output schemas documented anywhere. (4) Error handling is not visible, no recovery guidance, no categorization of failures, no actionable messages. (5) The repository structure (Makefile referencing 'mcp-servers/python/data_analysis_server' and 'mcp-servers/python/graphviz_server') suggests these are placeholder/example tool definitions, not production implementations. The 11 tools span two unrelated domains (data analysis + graphviz), suggesting this may be a template repository rather than a cohesive MCP server. File paths in tool definitions ('mcp-servers/python/data_analysis_server/Makefile', 'server_fastmcp.py') do not match the Rust/Cargo.toml structure shown, indicating misalignment between declared language (Rust) and actual tool implementation (Python). This inconsistency prevents confident assessment of actual tool implementations.
NO input schemas visible in source code. Tools declare parameters (path, format, dataset_name, etc.) but no JSON Schema definitions with types are shown in Cargo.toml, Makefile, or provided code samples.
Parameter descriptions are under 20 characters and lack actionable constraint information. Examples: 'File path or URL to the dataset' (35 chars, acceptable) but 'Data format: csv, json, excel, parquet' lacks constraint framing; 'Type of analysis: descriptive, correlation, regression' lacks guidance on which to choose when. No enums, no format declarations, no LLM-selection guidance.
Provide actual JSON Schema definitions for all 11 tools. For each tool, create a detailed input schema with types (string, number, array, object), required fields, and constraints (enum, pattern, minLength, maxLength). Example for load_data: {"type": "object", "required": ["path", "format"], "properties": {"path": {"type": "string", "description": "File path or URL (local file, s3://, http(s)://)", "minLength": 1}, "format": {"type": "string", "enum": ["csv", "json", "excel", "parquet"], "description": "Data format of the source file"}, "preview_rows": {"type": "number", "description": "Number of rows to preview (1-1000, default 10)", "minimum": 1, "maximum": 1000}}}.
Expand tool descriptions to 100-200 characters and include 'WHEN to use' guidance. Example rewrite for load_data: 'Load and preview data from CSV, JSON, Excel, or Parquet files. Use this first to inspect a new dataset structure, see sample rows, and infer column types before calling analyze_data or query_data. Supports local paths and remote URLs (S3, HTTP).'
Add output schema documentation for every tool. Example for load_data: 'Returns {data: [{row: object}], schema: {column_name: {type, nullable, sample_value}}, row_count: number, preview_limited: boolean}.'
Implement error handling with actionable messages. Example: if format parameter is invalid, return 'Invalid format: got "xlsx", must be one of: csv, json, excel, parquet. File appears to be Excel; use format="excel".' If file not found, return 'File not found at path. Call list_datasets() to see available files or verify the path.'
Tool descriptions are generic and do not answer 'WHEN to use this tool' or distinguish from similar tools. Example: 'Load and preview datasets from various formats' (51 chars) is action-focused but does not explain prerequisites, alternatives, or when to choose load_data vs query_data. Pattern requires 50-200 char descriptions explaining what, when, and any dependencies.
No output schemas documented. No tool definition shows what fields are returned, what structure is produced, or what the LLM should expect downstream (e.g., does load_data return {data: [...], schema: {...}, row_count: N} or just {rows: [...]}?). This prevents LLMs from chaining tools.
No error handling or recovery guidance visible. Pattern requires: error responses must tell the LLM what to do next (e.g., 'File not found. Check the path or call list_datasets() to see available files.'). No tool definition includes error categorization (retryable vs user-fixable vs fatal) or actionable error messages.
Mismatch between declared language (Rust/Cargo.toml) and source file references (Python paths like 'mcp-servers/python/data_analysis_server/...'). Source code only shows Rust Cargo.toml and Makefile; no actual Rust implementation visible. Tool definitions reference Python files that are not provided. This prevents verification of actual schema implementation and suggests this may be a template or example repo, not a production MCP server.
Composition issue: two unrelated tool domains (data_analysis + graphviz) in a single server without coherent purpose. Tools like load_data, analyze_data, and query_data share no context with create_graph, add_node, add_edge. This violates the principle that a cohesive MCP server should serve a single well-defined purpose or strongly related domain.
No tool annotations visible. Modern MCP pattern: tools should declare readOnlyHint, destructiveHint, and idempotentHint. Several tools modify state (transform_data, create_graph, add_node, add_edge, render_graph marked as WRITE risk) but no tool definition declares this via annotations.
No pagination support documented. Pattern requires: tools returning lists (e.g., analyze_data with multiple columns, visualize_data with multiple charts) should accept limit and offset/cursor parameters and return total_count or next_cursor. No tool definition mentions result limits or pagination.
Parameters lack type information and enum definitions. Pattern: when a parameter accepts one of a known set of values (e.g., format: csv|json|excel|parquet, analysis_type: descriptive|correlation|regression, chart_type: line|bar|scatter|histogram), it must be formally declared as an enum in the schema, not just mentioned in the description.
Add tool annotations to state side effects. Mark transform_data, create_graph, add_node, add_edge, render_graph with destructiveHint: true or readOnlyHint: false. Mark all READ-ONLY tools (load_data, analyze_data, query_data, visualize_data, statistical_test, time_series_analysis) with readOnlyHint: true.
Reconcile the Rust/Cargo.toml structure with actual tool implementations. Either (a) provide the Rust source code showing how tools are registered and schemas are defined, or (b) clarify that this is a template/example repo and provide actual production implementation. Currently, the mismatch (Rust declared, Python files referenced, no actual tool registration visible) prevents confident assessment.
Separate data_analysis and graphviz tools into two focused MCP servers, or document a coherent use case that justifies combining them (e.g., 'analyze data and generate diagrams'). If combined, add a server description explaining when/why an agent would use both tool groups together.
For parameters accepting enumerated values, formally declare enums in the schema instead of only mentioning them in descriptions. Example for visualize_data: {"chart_type": {"type": "string", "enum": ["line", "bar", "scatter", "histogram"], "description": "Type of visualization. Use 'line' for time-series trends, 'bar' for categorical comparisons, 'scatter' for correlation exploration, 'histogram' for distribution analysis."}}.
Add pagination support. Example for analyze_data: accept limit (1-100, default 20) and offset (default 0) parameters. Return {results: [{analysis}], total: number, has_more: boolean, next_offset: number}. Document that large result sets are capped at 100 items.
Add dry-run or confirmation support for destructive operations. Example for transform_data: accept dry_run: boolean parameter. If true, return what would be changed without modifying the dataset. This prevents agent mistakes.
Document parameter interdependencies. Example: if operations array in transform_data contains certain operation types, explain which operation_params are required. Example: 'If operations includes "normalize", provide normalize_method ("min-max"|"z-score"|"log") and optionally normalize_columns array.'
Consider adding a discovery/list tool (e.g., list_datasets, list_charts) to help agents understand available resources before calling load_data or query_data.
Add validation examples to descriptions. Example for query_data: 'SQL syntax: SELECT column_name FROM dataset WHERE condition. Example: "SELECT age, salary FROM employees WHERE department = \'sales\' AND salary > 50000."'
Clarify whether this is a real production server or a template. If production, provide actual source code. If template/example, mark it as such in the README and link to a real implementation example.