MCP server for reading and analyzing CSV files with schema and metadata extraction
CSVInfo provides 4 tools for CSV inspection. All tools are READ_ONLY and have basic structure, but suffer from significant gaps: parameter descriptions are minimal (just 'Path to the CSV file' repeated identically across all tools), output schemas are undocumented, error messages are generic, and tool descriptions lack context about when to use each tool or what they return. The naming convention is consistent (verb_noun) and appropriate, but descriptions are too brief to distinguish tools effectively. No input validation hints, no output shape documentation, and no guidance on tool interdependencies. This is typical of a minimal proof-of-concept server rather than production-ready tooling.
Count the number of columns in a CSV file
Count the number of rows in a CSV file
Get the schema of a CSV file
Read a CSV file and return its column names
Parameter descriptions are identical and uninformative across all tools ('Path to the CSV file'). They do not explain format constraints, encoding, size limits, or path resolution behavior (e.g., relative vs absolute, symlink handling).
Tool descriptions are too brief (14-32 chars) and lack essential context: what data structure is returned? When should the LLM call this tool vs. another? Are there prerequisites (e.g., file must be readable)? Example: 'Get the schema of a CSV file' does not explain that it returns a dict of column→type mappings or when to use it instead of read_csv_columns.
Output schemas are not documented. Code shows tools return dict, int, and list[str], but no schema structure is declared for the LLM. LLMs cannot reliably extract fields or plan chained calls without explicit output shape documentation.
Inferred effective spec: 2025-06-18+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 40 | 2025-06-18+ | v2 |
| 2026-03-09 | F | 39 | - | v1 |
Error messages are generic catch-alls ('Error getting CSV schema: {e}', 'Error counting CSV rows: {e}'). They do not distinguish between file-not-found, permission-denied, malformed-CSV, or other recoverable errors. LLM cannot determine whether to retry, ask user, or abandon.
Roots capability is used (deprecated since ~2026-07). Code includes RootsCapability check and list_roots() call. While fully functional through ~2027-07, new implementations should prefer passing directories via tool parameters, environment variables, or resource URIs instead of relying on client-initiated Roots.
Tool descriptions do not explain interdependencies or common workflows. LLM cannot infer that read_csv_columns is often called before get_csv_schema, or that count_csv_rows is lightweight vs. loading the full dataframe.
No input validation hints in parameter descriptions. No mention of valid file extensions (.csv), encoding (UTF-8 assumed?), max file size, or delimiter handling. LLM may pass invalid paths or unsupported formats without guidance.
Tool descriptions do not clarify return values or structure. 'Count the number of rows' returns an int, but does it include or exclude the header row? Does it match df.height or len(df.columns)? Ambiguity forces LLM to guess.