Model Context Protocol server for Databricks Unity Catalog and SQL execution
The Databricks MCP server has 16 tools with mixed quality. Schemas are present and reasonably structured with type declarations, but descriptions vary significantly in quality and language consistency (English and Portuguese mixed). Parameter descriptions are present but often generic. Critical gaps: no error handling guidance, no tool annotations, no output schema documentation, and descriptions often lack context on WHEN to use tools and WHAT they return. The server lacks idempotency guidance for write operations and has no permission/scope declarations. About 60% of tools have adequate descriptions; 40% have minimal or language-inconsistent documentation.
Cria um novo catálogo no Databricks Unity Catalog. Parâmetros: name (str): Nome do catálogo (obrigatório) comment (str, opcional): Descrição do catálogo connection_name (str, opcional): Nome da conexão externa options (dict, opcional): Propriedades customizadas (ex: {"property1": "valor"}) properties (dict, opcional): Propriedades customizadas (ex: {"property1": "valor"}) provider_name (str, opcional): Nome do provedor Delta Sharing share_name (str, opcional): Nome do share storage_root (str, opcional): URL raiz de armazenamento
Create a new schema in a catalog. Args: catalog_name (str): The name of the parent catalog. name (str): The name of the new schema. comment (str, optional): A comment for the schema. properties (dict, optional): A dictionary of key-value properties.
Create a new table in a schema. Args: table_info (Dict): A dictionary containing the table information.
Deleta um catálogo no Databricks. Use force=True para forçar a exclusão mesmo se houver dependências. Parâmetros: name (str): Nome do catálogo a ser deletado force (bool, opcional): Se True, força a exclusão (default: False)
Language inconsistency: descriptions mix Portuguese and English. 'create_catalog' docstring is entirely in Portuguese ('Cria um novo catálogo...'), while 'list_catalogs' is in English. This breaks LLM parsing and makes the server appear unprofessional. All descriptions must be in a single language (recommend English for international tooling).
Minimal or missing descriptions on discovery tools. 'get_catalog_info', 'get_schema_info', 'get_table_info', and 'list_sql_warehouses' have descriptions under 35 characters ('Get catalog information', 'Get information about a specific schema', 'Get information about a specific table', 'Lists all available SQL Warehouses...'). These lack context: WHEN should the LLM call this? WHAT fields are returned? Descriptions must be 50 - 200 chars and explain purpose, prerequisites, and return structure.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 68 | <=2025-11-25 | v2 |
Delete a schema. Args: full_name (str): The full name of the schema (e.g., 'catalog_name.schema_name').
Delete a table. Args: full_name (str): The full name of the table (e.g., 'catalog_name.schema_name.table_name').
Executa uma query SQL em um SQL Warehouse e aguarda o resultado. Esta ferramenta submete a query, monitora o status da execução e, apos o sucesso, busca os dados (inclusive de links externos) e os retorna. Utilize a ferramenta 'list_sql_warehouses' para obter um 'warehouse_id' válido. Args: warehouse_id (str): O ID do SQL Warehouse onde a query será executada. sql_query (str): A instrução SQL a ser executada. timeout_seconds (int): O tempo máximo em segundos para aguardar a conclusão da query. Returns: Dict: Um dicionário contendo o resultado, com a seguinte estrutura: - 'schema' (dict): Descreve as colunas do resultado. - 'data_array' (list): Uma lista de listas, onde cada lista interna representa uma linha. - 'row_count' (int): O número total de linhas retornadas.
Get catalog information
Get information about a specific schema.
Get information about a specific table.
List all catalogs in the Databricks workspace, with pagination support. Args: page_token (str, optional): Token for next page of results Returns: Dict: Response from Databricks API, including 'catalogs' and 'next_page_token'
List all schemas in a specific catalog. Args: catalog_name (str): The name of the catalog. Returns: Dict: Response from Databricks API, including 'schemas'.
Lists all available SQL Warehouses to find a 'warehouse_id' for running queries.
List all tables in a specific schema. Args: catalog_name (str): The name of the catalog. schema_name (str): The name of the schema. Returns: Dict: Response from Databricks API, including 'tables'.
Update an existing schema. Args: full_name (str): The full name of the schema (e.g., 'catalog_name.schema_name'). new_name (str, optional): A new name for the schema. comment (str, optional): A new comment for the schema. properties (dict, optional): A new set of key-value properties.
Update an existing table. Args: full_name (str): The full name of the table (e.g., 'catalog_name.schema_name.table_name'). updates (Dict): A dictionary containing the updates to the table.
No output schema documentation. Tools like 'execute_sql_query' return complex nested structures (schema dict, data_array list, row_count int) but the MCP server does not expose these in the tool's registered return type. LLMs cannot plan downstream steps without knowing what fields to expect. All tools must declare their output schema in the tool registration (via FastMCP return type hints or response schema field).
No tool annotations (destructiveHint, readOnlyHint, idempotentHint). The MCP server identifies tools as READ_ONLY, WRITE, or DESTRUCTIVE in the evaluation metadata, but these hints are NOT exposed in the tool definitions themselves. FastMCP and the MCP spec support tool annotations, use @mcp.tool(hideFromUI=false) and return structured hints so clients can warn users before executing destructive operations.
No error recovery guidance. The code shows basic try-except blocks (e.g., catalogs.py: 'except requests.exceptions.RequestException, ValueError as e: raise Exception(...)'), but errors are re-raised as generic Exceptions. LLMs receive no guidance on whether errors are retryable, what to try next, or if they require user intervention. Per pattern:recovery-guide, errors must be categorized and include next steps (e.g., 'Catalog not found. Call list_catalogs() to see available options.').
No scope/permission declarations. Tools like 'delete_catalog' and 'delete_table' perform destructive operations, but the tool definitions do not declare required permissions (e.g., 'requires: write:catalog, delete:catalog'). Agents deploying this server cannot configure least-privilege access, and audit logs lack permission context.
Vague parameter descriptions on generic parameters. 'create_table' accepts 'table_info' (dict) with description 'A dictionary containing the table information.' This is not actionable, what keys does the dict expect? What is the schema? LLMs cannot construct valid input without seeing the expected structure. Inline the full schema or provide an example pattern.
No idempotency guidance for write operations. 'create_catalog', 'create_schema', 'create_table' do not state whether calling twice with the same name returns an error or succeeds idempotently. Agents retry on transient failures, idempotent operations prevent duplicate records; non-idempotent ones risk silent duplicates.
Missing pagination limits and result truncation. 'list_catalogs', 'list_schemas', 'list_tables' accept pagination tokens but do not document the default result limit or maximum page size. Per pattern:paginated-result, tools should state: 'Returns up to 50 items per page. Pass next_page_token to fetch additional results.' Without this, agents may request huge result sets that exhaust context.
Credentials exposed in per-request headers (X-Databricks-Host, X-Databricks-Token). While the code uses a middleware to extract them from headers (not parameters), this approach is fragile: headers can be logged in access logs, cached by proxies, or leaked in error traces. The MCP spec 2026-07-28 recommends server-side secret injection via environment variables or vault. Consider migrating to a single authenticated HTTP transport or server-side token issuance.