MCP Server for interacting with Delta Lake tables via Spark
The server provides 8 tools with consistent naming patterns and READ_ONLY operations, but significant gaps exist in parameter descriptions and output schema documentation. Tool names follow verb_noun convention (list_, get_, count_, sample_, query_) which is appropriate. However, parameter descriptions are minimal, many lack detail about expected formats, constraints, and ranges. Output schemas are not documented in the visible code, making it difficult for LLMs to plan downstream tool calls. Error handling guidance is absent. The schema coverage is moderate; all tools visible have input schemas with types, but descriptions for parameters are sparse. Notably, tools like sample_delta_table and query_delta_table accept user-provided input (WHERE clauses, SQL queries) without documented validation or injection safeguards, which is a security concern for an agent-facing interface.
Gets the total row count for a specified Delta table.
Gets the complete structure of all databases, optionally including table schemas.
Gets the schema (column names) of a specific table in a database.
Returns the health status of the API.
Lists all tables in a specific database, optionally using PostgreSQL for faster retrieval.
Lists all databases available in the Hive metastore, optionally using PostgreSQL for faster retrieval.
Output schemas not documented. LLMs cannot see what fields are returned (e.g., does list_databases return names only, or also metadata like owner, creation date, size?). This forces agents to make tool calls blind and risks context loss when downstream tools need IDs/references.
Parameter descriptions lack detail on format and constraints. 'use_postgres' is documented as boolean with brief rationale, but tools like sample_delta_table have parameters (limit: integer, where_clause: string) without stated ranges, validation rules, or format requirements. E.g., limit says 'default 10, max 1000' in description but no min is specified; where_clause is free-form SQL with no injection guidance.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 72 | 2026-07-28+ | v2 |
| 2026-03-09 | C | 65 | - | v1 |
Executes a SQL query against a specified Delta table.
Retrieves a small sample of rows from a specified Delta table.
query_delta_table accepts arbitrary SQL SELECT queries as a free-form string with minimal validation guidance. No mention of SQL injection risks, query timeout, or result limits. An agent could accidentally pass malicious or expensive queries. This violates the untrusted-input principle and pattern:tool-gateway.
No error handling guidance. Descriptions do not explain what happens on failure (database not found, table not found, permission denied, query timeout). LLMs cannot plan recovery or understand retry eligibility. No pattern:recovery-guide implementation visible.
Tool descriptions are generic and under-optimized for LLM selection. E.g., 'Gets the schema (column names) of a specific table' does not explain WHEN to call this vs get_database_structure, or what the returned schema looks like. Baseline is 50-200 chars for A-grade tools; most descriptions here are 60-90 chars but lack depth.
No pagination parameters on list_* tools. list_databases and list_database_tables could return hundreds of items; no offset/limit/cursor parameters visible. Unbounded lists blow context windows. Baseline pattern:paginated-result calls for limits and cursors.
where_clause parameter in sample_delta_table is SQL-free-form with no examples, constraints, or SQL injection guidance. LLMs may pass malformed WHERE syntax or attempt injection. Should document expected syntax and validate strictly server-side.