This dbt CLI MCP server has well-structured tool definitions with consistent naming, solid descriptions, and proper parameter schemas. All 9 tools follow verb_noun convention (dbt_run, dbt_test, dbt_ls, etc.). Descriptions are action-oriented and explain WHEN to use each tool, with most ranging 150-280 characters, within the productive 10-1024 range. However, there are notable gaps: (1) Output schemas are not explicitly documented in the source code or tool descriptions, forcing LLMs to reason about what fields to expect. (2) Error handling guidance is absent, no recovery hints like 'if compilation fails, try dbt_debug first'. (3) Parameters lack explicit enums for constrained values (e.g., output_format accepts 'json|name|path|selector' but is declared as free-form string). (4) No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite clear read/write risk distinction. (5) Parameter descriptions for some tools could be more precise about valid input formats (e.g., dbt_ls's 'models' parameter just says 'dbt selection syntax' without explaining what that means). The tools are well-designed for agent composition, each does one clear thing, and naming makes intent obvious. Risk levels are correctly assigned (WRITE vs READ_ONLY), but they are not expressed as tool hints in the schema.
Run the build command. An AI agent should use this tool when it needs to run a comprehensive build of the dbt project that includes models, tests, seeds, and snapshots in dependency order. This is useful for ensuring the entire project is built and validated in a single operation.
Compile dbt models. An AI agent should use this tool when it needs to generate the SQL that will be executed without actually running it against the database. This is valuable for validating SQL syntax, previewing transformations, or investigating how dbt interprets models before committing to execution.
Debug dbt project configuration. An AI agent should use this tool when it needs to verify that the dbt project is properly configured, including checking database connectivity, profile configurations, and other setup issues.
Install dbt package dependencies. An AI agent should use this tool when it needs to install or update packages specified in the dbt project's packages.yml file. This is typically done before running models that depend on external packages.
List dbt resources. An AI agent should use this tool when it needs to discover available models, tests, sources, and other resources within a dbt project. This helps the agent understand the project structure, identify dependencies, and select specific resources for other operations like running or testing.
Output schemas not documented. Tool descriptions state 'Returns: Output from the dbt X command as text' but do not specify what fields, structure, or format agents should expect. LLMs cannot plan downstream calls or extract chaining IDs without knowing the response structure.
Constrained parameters use free-form strings instead of enums. 'output_format' accepts only 'json|name|path|selector' but is declared as string type with no enum constraint. 'resource_type' accepts 'model|test|source|etc.' but lacks enum. This invites LLM hallucination of invalid values like 'json-compact' or 'yaml'.
No error handling or recovery guidance. All tool descriptions omit what errors are possible and how to recover (e.g., 'If dbt_compile fails with a syntax error, check your Jinja templates', or 'If dbt_run fails with a database connection error, try dbt_debug first'). LLMs have no guidance on next steps when calls fail.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 28 | - | v1 |
Run dbt models. An AI agent should use this tool when it needs to execute dbt models to transform data and build analytical tables in the data warehouse. This is essential for refreshing data or implementing new data transformations in a project.
Load CSV files as seed data. An AI agent should use this tool when it needs to load static data from CSV files into the data warehouse. Seeds are useful for reference data, lookup tables, or test data that doesn't change frequently.
Preview model or source results. An AI agent should use this tool when it needs to see the actual data produced by a model or source without having to query the database directly. This is useful for validating that transformations are working correctly or understanding the data structure.
Run dbt tests. An AI agent should use this tool when it needs to validate data quality and integrity by running tests defined in a dbt project. This helps ensure that data transformations meet expected business rules and constraints before being used for analysis or reporting.
Tool annotations missing. All tools carry explicit Risk levels (WRITE vs READ_ONLY) in the specification, but these are not exposed as readOnlyHint or destructiveHint in the tool schema. LLMs cannot programmatically determine which tools are safe to call speculatively vs. which require confirmation.
Parameter descriptions lack format and constraint details. 'models' parameter description says 'using the dbt selection syntax (e.g., "model_name+")' but does not explain what syntax is valid (operators like +, @, *; logic like 'tag:daily'). LLMs cannot construct valid selectors without understanding the grammar.
Default values for project_dir are risky. dbt_run, dbt_test, dbt_ls, etc. all default project_dir to '.' (current directory), but the description explicitly warns against this: 'ABSOLUTE PATH ... (e.g. '/Users/username/projects/dbt_project' not '.')'. This contradiction may cause silent failures if the LLM does not specify an absolute path.
dbt_show requires 'models' but describes it as optional. The parameter definition shows 'models: Optional[str]' with default None, but the docstring says 'Specific model or source to show (required)'. This type-description mismatch forces LLMs to guess whether the parameter is mandatory.
No pagination or result limits specified. dbt_ls can return hundreds of resources. The tool description does not mention a limit parameter or state how many results are returned. LLMs cannot predict result size, large responses may exhaust context or timeout.