Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
This server provides 7 data quality tools with basic structure, but significant gaps prevent higher scoring. Tool names follow verb_noun convention appropriately (analyze_data_file, suggest_quality_rules, validate_data_quality). However, descriptions are generic and lack LLM-optimization guidance. Parameter schemas are present but incomplete, many lack proper type definitions, minimum/maximum bounds, or enum constraints. Output schemas are largely undocumented. Error handling is minimal with no recovery guidance. The codebase shows analysis logic but tool definitions are partially inferred from the server.py code rather than explicitly registered with complete metadata. Tool composition is reasonable (each has one primary responsibility), but parameter documentation needs improvement to guide LLM behavior.
Tools (7)
analyze_data_fileread onlysource verified65/100
Analyze a data file and return comprehensive quality assessment
analyze_sample_dataread onlysource verified62/100
Analyze sample data rows and return quality assessment
Output schemas are not documented for any tool. LLMs cannot predict what fields to expect or plan downstream calls. The code shows analysis logic returns dicts with 'shape', 'columns', 'dtypes', 'column_analysis', 'issues' keys, but this is not formalized as a documented return schema.
Parameter descriptions lack actionable constraints and format guidance. For example, 'confidence_threshold' in generate_expectation_suite is described as '(0.0-1.0), default 0.8' in the tool definition, but this is insufficient, should explicitly state 'Numeric value between 0.0 and 1.0 representing minimum confidence level for generated rules. Values below 0.5 may produce low-quality rules.' Similar gaps exist for file_path parameters, no guidance on supported formats, required permissions, or handling of relative vs absolute paths.
Document the output schema for each tool as a JSON Schema object. For analyze_data_file, explicitly define: returns object with keys {shape: [int, int], columns: [string], dtypes: {string: string}, column_analysis: {string: {data_type: string, null_count: int, null_percentage: float, unique_count: int, unique_percentage: float, stats: {mean?, median?, std?, min?, max?, q25?, q75?}}}, issues: [{type: string, column: string, severity: string, description: string, percentage: float}]}. Similar formalization needed for all 7 tools.
Enhance parameter descriptions to include format constraints, valid ranges, and examples of correct vs incorrect input. For file_path: 'Absolute or relative path to a data file. Supported formats: .csv, .xlsx, .xls, .json, .parquet. Example: /data/customers.csv or data/Q4_sales.xlsx. File must exist and be readable.' For confidence_threshold: 'Numeric threshold 0.0 - 1.0; rules below this confidence are excluded. Recommended: 0.7 for permissive rules, 0.9 for strict rules. Default 0.8.'
Add error recovery guidance to tool descriptions. Example for validate_data_quality: 'If validation fails with "expectation suite not found", ensure you first generated a suite via generate_expectation_suite(). If validation fails with "missing columns", verify the input file contains the same columns as the expectation suite was created for.'
Implement per-request result limits and pagination. Add optional parameters: limit (default 20, max 100) and offset (default 0) to get_data_profile and suggest_quality_rules. Document in the description: 'Returns up to [limit] results. Use offset to retrieve subsequent pages. Large datasets may be summarized, request only the most relevant suggestions via the limit parameter.'
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Error handling is minimal. The code shows file validation via _validate_file_path() and _load_data_file() with basic try-catch blocks, but tool responses do not include recovery guidance. If a file is not found or unsupported format is provided, the LLM receives a generic error with no suggestion of what to try next (e.g., 'File not found. Verify the path is correct and the file exists.' or 'Unsupported format. Supported formats: CSV, XLSX, JSON, Parquet.').
Parameter enums are missing where they should be present. The 'minimal' parameter in create_comprehensive_profile is boolean, but no guidance is provided on when to use true vs false. The profile_name and suite_name parameters accept any string with no constraints, should enforce naming rules (alphanumeric, underscores, max length) to prevent downstream issues.
Tool definitions are inferred from code rather than explicitly registered with complete metadata. The server.py file shows function decorators and docstrings, but full parameter schema registration details are truncated (code sample ends mid-schema). This violates verifiability, cannot confirm all tools have complete JSON Schema definitions with all required fields.
No pagination support documented for tools that may return large result sets. If suggest_quality_rules or get_data_profile return many items, there is no limit parameter, page offset, or cursor mechanism visible in the schemas. This risks context window exhaustion when processing large datasets.
Global mutable state (_current_schema, _current_profile, _current_rules, _current_data_source) is used to persist analysis context between tool calls. This is problematic in concurrent or multi-user scenarios and violates stateless request handling principles. The server should return complete results from each tool call without relying on server-side state.
Tool descriptions do not distinguish when to use analyze_data_file vs analyze_sample_data. Both perform analysis and return quality assessments, the LLM cannot easily determine which to call without deeper reasoning. add clarifying language: 'Use analyze_data_file for files on disk; use analyze_sample_data for in-memory data structures or when you have a representative sample.'
analyze_data_fileanalyze_sample_data
Remove global mutable state (_current_schema, _current_profile, etc.) and ensure each tool call returns complete, self-contained results. If analysis context must be shared between calls, require explicit passing of analysis_results object (as already done for suggest_quality_rules and generate_expectation_suite). This ensures stateless handling and prevents state corruption in concurrent scenarios.
Add tool annotations (readOnlyHint, destructiveHint) via FastMCP to signal tool intent. All 7 tools are read-only; annotate them as such: @mcp.tool(description='...', hints={'readOnly': True}). This helps the LLM understand which tools are safe to call repeatedly.
Clarify tool distinctions in descriptions. Revise: analyze_data_file: 'Analyze a data file on disk and return comprehensive quality assessment. Use this for files stored locally.' analyze_sample_data: 'Analyze in-memory data rows and return quality assessment. Use this for small representative samples or data from other tools.'
Add enum constraints where applicable. For minimal parameter in create_comprehensive_profile: change type to enum with values ['true', 'false'] and add description: 'Set to true for a faster, minimal profile (omits correlations, complex visualizations); false for a detailed profile (slower but more thorough).'
Validate all parameter inputs before processing and return field-level error messages. Example: 'Invalid confidence_threshold: got 1.5, expected float 0.0 - 1.0.' Example: 'File format not supported: got .txt, expected one of: csv, xlsx, xls, json, parquet.'
Add usage hints in tool descriptions to guide multi-step workflows. Example for suggest_quality_rules: 'Typically called after analyze_data_file() or analyze_sample_data(). Pass the analysis_results object to generate initial rules. Refine with generate_expectation_suite() to create formal validation expectations.'
Consider adding a batch variant or a consolidate tool (e.g., analyze_and_generate_suite) that combines analyze_data_file + generate_expectation_suite in one call for common workflows. This reduces round-trips and token overhead.
Document whether file paths must be absolute or can be relative, and if relative, relative to what directory. Update file_path description: 'Path to the data file. Relative paths are resolved from the current working directory. For reliability, use absolute paths.'