MCP server for Apache Hadoop Hive that provides tools to query tables, retrieve columns, and execute SQL queries against Hive data sources via JDBC
This server defines three read-only tools for querying Apache Hadoop Hive with reasonable structure. Tools are clearly named with action verbs (get_, run_) and have descriptions present. However, several quality gaps limit production readiness: (1) Parameter descriptions lack detail about constraints, formats, and valid values. For example, 'The table name' does not explain whether table names are case-sensitive, what characters are allowed, or how to escape special names. (2) Output schemas are completely undocumented, the tools claim to return CSV format, but there is no structured schema definition of what fields or rows to expect, forcing LLMs to parse unstructured text. (3) Error handling is minimal, the run() method throws RuntimeException with a generic message, giving agents no guidance on retry logic, user-fixable errors, or recovery steps. (4) No pagination support visible despite tools potentially returning large result sets. (5) Tool composition is reasonable, hive_get_tables and hive_get_columns are discoverable prerequisites to hive_run_query, but cross-tool chaining IDs (catalog, schema, table) are not explicitly documented in descriptions.
Retrieves a list of fields, dimensions, or measures (as columns) for an object, entity or collection (table). Use the `hive_get_tables` tool to get a list of available tables. The output of the tool will be returned in CSV format, with the first line containing column headers.
Retrieves a list of objects, entities, collections, etc. (as tables) available in the data source. Use the `hive_get_columns` tool to list available columns on a table. Both `catalog` and `schema` are optional parameters. The output of the tool will be returned in CSV format, with the first line containing column headers.
Execute a SQL SELECT statement.
No output schemas documented for any tool. Tools claim CSV format but provide no structured schema of returned fields, data types, or row counts. LLMs must parse unstructured text, risking errors and wasting tokens.
Parameter descriptions are generic and lack actionable constraints. Example: 'The table name' does not specify case sensitivity, character restrictions, escaping rules, or maximum length. LLMs cannot validate input or format identifiers correctly.
Error handling is minimal and non-actionable. RuntimeException with generic 'ERROR: <message>' provides no guidance on retry logic, user-fixable errors, or recovery steps. Agents cannot distinguish between transient failures and unrecoverable problems.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 48 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
No pagination or result limiting strategy visible. Large table lists or query result sets could blow context windows. Tool descriptions do not mention maximum rows returned or pagination parameters.
hive_run_query accepts free-form SQL string with no validation or injection protection visible at the tool interface. SQL injection risk if agent is compromised or prompted maliciously.
SQL dialect constraints mentioned in description ('FROM, INNER JOIN, LEFT JOIN, GROUP BY, ORDER BY, LIMIT/OFFSET') but not enforced or documented formally. Agents may attempt unsupported clauses (subqueries, CTEs, DML) and receive cryptic errors.
Tool descriptions mention CSV output but do not specify handling of special characters, quoting, escaping, or header row format. Ambiguous output format risks parsing errors in downstream LLM reasoning.