AI-driven quality & governance MCP Server for dbt projects. Audit coverage, profile data, detect schema drift, and auto-generate documentation.
dbt-doctor has 12 tools with basic descriptions and no visible input schemas in the provided source. Most tool descriptions are present but lack LLM-optimization guidance (e.g., many exceed 200 chars and include implementation details rather than WHEN to use them). Parameters are documented in descriptions but lack formal schema definitions with types. No visible error handling patterns, no output schema documentation, and no tool annotations (readOnlyHint/destructiveHint). The server is READ_ONLY-focused with one WRITE operation (update_model_yaml) that lacks confirmation/dry-run safeguards. Tool composition is reasonable (each does one thing), but parameter documentation relies entirely on inline text rather than structured schemas. The codebase shows professional organization but definition quality lags production standards.
Analyze the dbt project DAG (Directed Acyclic Graph) structure. Detects: - Orphan models (no downstream consumers or exposures) - Root models (no model dependencies, only sources) - High fan-out models (one model feeding many downstream consumers) - Maximum chain depth (long lineage chains) Use this to find structural issues and potential refactoring opportunities.
Run a comprehensive quality audit of the entire dbt project. Checks: - Model-level documentation coverage (% of models with descriptions) - Column-level documentation coverage (% of columns with descriptions) - Test coverage (% of models with at least one dbt test) - Naming convention adherence (snake_case, layer prefixes) Returns an overall health score (0-100%) and a list of specific issues. Use this as the FIRST step before any documentation or test generation work.
Show detailed test coverage statistics for all models. Returns a ranked list of models sorted by test coverage, from worst to best. Use this to identify which models need tests the most, then call `suggest_tests` on the worst offenders.
Detect schema drift between dbt manifest and actual database table.
Execute a custom read-only SQL query against the dbt database.
No visible input schemas (JSON Schema with types) for any tool. All parameter descriptions are inline text only. Parameters lack formal type declarations, constraints, and descriptions visible in a structured schema format.
No output schema documentation. Tool descriptions do not specify what fields are returned or how downstream tools should interpret responses. This forces LLMs to infer output structure from experience.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 46 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 29 | - | v1 |
Generate comprehensive model documentation including descriptions and column details.
Get full details for a specific dbt model: SQL code, columns, lineage, and config. Use this before writing documentation or suggesting tests for a specific model. Also use this to understand what a model does before profiling it.
Get a comprehensive health dashboard for the dbt project. This is the ENTRY POINT tool — call this first to get an overview of: - Overall quality score - Top coverage issues - DAG structural issues - Recommended next actions Think of it as a doctor's overall diagnosis before deciding on treatments.
List all dbt models in the project. Returns a summary of every model including name, materialization, description status, and documentation coverage. Use this as your starting point before any other tool.
Profile all columns in a dbt model, producing comprehensive statistics.
Generate recommended dbt tests for a specific model.
Generate and update schema.yml for a model with documentation and column info.
Tool annotations (readOnlyHint, destructiveHint, idempotentHint) are absent. The single WRITE operation (update_model_yaml) is not explicitly marked as destructive, and there is no confirmation/dry-run pattern to prevent accidental data modification.
Descriptions are not LLM-optimized. Many exceed 200 characters and include implementation minutiae (e.g., 'Returns a summary of every model including name, materialization, description status, and documentation coverage') rather than guiding WHEN to call the tool. Baseline for A+ tools is 50-200 chars with clear intent.
Parameter descriptions are minimal or missing context. 'profile_model' description says only 'Profile all columns in a dbt model, producing comprehensive statistics.', no guidance on what statistics are included or when to call this vs other profiling tools.
execute_query accepts raw SQL with no validation, injection protection, or scope limiting described. This is a potential SQL injection risk if untrusted input reaches the tool. Tool description lacks security guidance.
No error handling patterns visible. Tools reference RuntimeError states ('dbt-doctor not initialized', 'Database connector not available') but descriptions do not indicate to the LLM what these errors mean or how to recover.
update_model_yaml is a WRITE operation but has no confirmation/dry-run pattern. LLMs may call it unexpectedly and modify the project. No explicit warning in description that this modifies files on disk.
No pagination guidance. Tools like 'list_models', 'check_test_coverage', and 'analyze_dag' return potentially large result sets with no limit parameter, offset/page parameter, or total_count field documented.
Tool names could be more granular. 'get_project_health' and 'audit_project' appear to do similar work; descriptions should clarify the distinction or consider merging to reduce LLM confusion.