An LLM-powered agent for interacting with dbt projects via MCP protocol, exposing dbt models, metadata, and project information to Claude and other LLM clients
The server defines 4 tools with clear action verbs and reasonable descriptions. All tools have proper input schemas with type definitions and parameter descriptions. However, there are notable gaps: output schemas are not documented, error handling guidance is missing, and the descriptions, while adequate, lack the depth needed for LLM-optimized selection in complex scenarios. The tools follow a consistent naming pattern (verb_noun) and cover a coherent dbt exploration domain, but lack defensive details about constraints, limits, and recovery paths that production tools typically include.
Get detailed information about specific dbt models including SQL, documentation, and lineage
Get a summary of connected dbt projects and their models
List available dbt models in the knowledge base with optional filtering
Search for relevant dbt models using natural language queries with semantic similarity
Output schemas not documented. Users cannot see what fields are returned by each tool, forcing LLMs to infer structure from context or make failed downstream calls.
Error handling guidance missing. No recovery hints. If search_dbt_models returns no results or a threshold error, the tool description does not explain what the LLM should do next (e.g., 'lower similarity_threshold' or 'try list_dbt_models').
Limit parameters lack enforcement rationale. 'limit' defaults are stated (50, 10, etc.) but descriptions do not explain why or what happens if exceeded. Baseline rubric expects limits to include reasoning like 'prevents context window exhaustion' or 'API timeout risk'.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 71 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Dependency hints missing. get_model_details accepts model_names but does not hint that list_dbt_models or search_dbt_models can be called first to discover valid names. Forces agents to guess or trial-and-error.
Parameter constraints not fully documented. similarity_threshold accepts 0.0 - 1.0 but does not explain semantics ('0.0 = no match, 1.0 = perfect match'). Materialization enum is not explicitly listed in description.