MCP server connecting Claude to Metabase for data analysis
The Metabase MCP server provides 26 tools with mostly complete schemas and descriptions. Tools are well-named with action verbs (get_, list_, create_, update_, delete_, execute_). However, several critical gaps reduce confidence: (1) Output schemas are NOT documented, the code shows response formatting happens in utilities but the tool registration does not publish what fields clients should expect. (2) Many parameter descriptions lack constraint details (ranges, allowed formats, dependencies). (3) Error handling is present but rarely guides recovery ('Tool is disabled by server policy' vs 'Tool is disabled. Enable with allow=[...]'). (4) Batch and workflow tools (batch_execute, run_workflow) have complex nested schemas but limited guidance on when to use them vs sequential calls. (5) Write operations (create_card, delete_dashboard) lack confirmation/dry-run patterns despite being destructive. Average tool description ~120 chars (baseline 194), which is adequate but on the short side. Schemas are present and typed but lack explicit output documentation.
Answer questions about your data using natural language
Execute multiple operations in a single call. Supports up to 20 operations run in parallel.
Compare two metrics or time periods
Create a new question/card in Metabase
Create a new dashboard
Delete (archive) a question/card
Delete (archive) a dashboard
Output schemas are not documented for any tool. Code shows results are formatted (formatQueryResult, formatSchemaResult) but tool registration does not declare what fields clients should expect. This forces LLMs to guess at response structure and breaks tool chaining.
Error handling does not guide recovery. batch_execute and run_workflow return {success: false, error: string} but errors like 'Tool is disabled by server policy' do not suggest next steps. Per pattern:recovery-guide, errors should include actionable guidance ('Try enabling via server config' or 'Call check_permissions first').
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 68 | <=2025-11-25 | v2 |
Execute an existing question/card and get results
Execute a custom SQL query against a database
Explain a SQL query in plain English
Auto-generate insights from an existing card/question
Get details of a specific question/card
List all collections in Metabase
Get dashboard details including its cards
Get tables and columns for a database
List questions/cards in Metabase. Returns up to 100 cards. Use search_content for discovery.
List all dashboards in Metabase
List all connected databases
Convert a natural language question to SQL query
Get suggestions to optimize a SQL query
Execute a multi-step workflow pipeline in a single call. Steps run sequentially and can reference previous step results using "$stepName.path" syntax.
Search for dashboards, cards, and other Metabase content
Identify trends in time-series data
Update an existing question/card
Update an existing dashboard
Check SQL query for syntax issues and security concerns
Destructive tools (delete_card, delete_dashboard) lack confirmation or dry-run patterns. An agent can irreversibly delete dashboards without a confirmation step. Per pattern:confirmation-request, critical operations should offer a dry-run or require explicit user confirmation.
run_workflow tool description and schema are complex but lack detailed documentation. No guidance on: max workflow depth/complexity, how to reference nested step results ($stepName.path.to.deep.field syntax), behavior when on_error='continue' but all steps depend on prior results, limits on step count. This tool risks misuse due to cognitive load.
Parameters lack constraint documentation. Many accept enums or have ranges but descriptions do not always state constraints (e.g., compare_metrics and trend_analysis accept 'database_id' but do not clarify if only valid DB IDs are accepted, or what error occurs for invalid ones). Per pattern:constrained-input, enums and ranges should be explicit in descriptions.
Examples in descriptions (e.g., 'Compare sales in Q1 vs Q2' in compare_metrics) violate best practice, LLMs latch onto example values and reuse them literally. Per pattern:tool-description, use formal constraints (enums, patterns) instead of prose examples.
Parameter dependencies and mutual exclusivity are undocumented. For example, nlq_to_sql accepts both 'question' and 'tables', but it is unclear: are 'tables' optional constraints, or required? What if 'question' references a table not in 'tables'? Per pattern:tool-description, dependencies must be explicit.
Pagination is not consistently offered or documented. list_cards and list_dashboards lack offset/next_cursor parameters despite potentially large result sets. Per pattern:paginated-result, tools returning lists must support pagination and document the behavior.
Tool descriptions are generic or under-detailed. 'List all dashboards' and 'List all collections' do not explain when to call them, what data they reveal, or why an agent would choose one over the other. Per pattern:tool-description, descriptions should answer WHAT, WHEN, and WHY.
Tool composition is incomplete. Writing tools (create_card, create_dashboard) do not clearly state what happens next. Does create_card return card_id (for use with execute_card)? Does create_dashboard allow adding cards in the same call or require a separate step? Broken chaining forces discovery detours per pattern:tool-chain.