The server provides 5 well-named tools with clear, substantive descriptions (150-250 chars each) and complete JSON Schema input definitions. However, there are significant gaps in schema completeness and parameter descriptions that prevent a higher score. All tools follow verb_noun naming conventions (guardrail_status, guardrail_review, etc.), which is excellent. Descriptions are action-oriented and answer WHEN to use each tool. However, parameter descriptions are inconsistent, some parameters lack detail about constraints, format, or valid ranges. No output schemas are documented, forcing LLMs to infer result structure. The server lacks error handling guidance, parameter validation rules, and recovery instructions. The codebase is readable and tool definitions are explicit and visible in guardrail/server.py, so no inference penalty applies.
Show generated SQL checks without executing them. Useful for reviewing or debugging the SQL that guardrail_review would run.
Return diff, raw SQL, and metadata for changed models. Use this to get everything needed to reason about semantic edge cases.
Full dbt model review pipeline. Parses manifest, detects changed models via git diff, generates SQL checks (grain, distribution, join, rowcount), executes against Snowflake, evaluates PASS/WARN/FAIL, and writes results. Returns structured JSON summary.
Execute semantic edge case SQL queries against Snowflake and store results. Submit edge cases that Claude Code identified from analyzing model diffs.
Quick metadata about the dbt project state: manifest age, model count, git branch, changed models, blast radius, and last review summary. Call this first to orient before running a full review.
No output schemas documented for any tool. LLMs cannot infer the structure, field names, or types of results returned by guardrail_status, guardrail_review, guardrail_checks, or guardrail_model_context. This forces LLMs to guess at which fields to extract and how to chain results to downstream calls.
Parameter descriptions lack detail on constraints, formats, and valid ranges. Example: 'dbt_project_dir' has description 'Path to dbt project root' but does not specify whether it must be absolute or relative, whether it must exist, or what error is returned if missing. 'base_branch' defaults to 'main' but does not explain what happens if the branch does not exist.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 68 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
guardrail_review has a 'checks' parameter accepting an array of strings with description 'Check categories to run: grain, distribution, join, rowcount' but does not declare it as an enum. Free-form strings invite LLMs to pass invalid values like 'density' or 'correlation'. Should use enum constraint with exactly ['grain', 'distribution', 'join', 'rowcount'].
No error handling guidance. Tools lack descriptions of how they fail, what errors LLMs should expect, and how to recover. For example: what happens if the Snowflake connection fails? What if a dbt project has syntax errors? What if git diff fails? LLMs receive no actionable recovery instructions.
guardrail_run_edge_cases accepts edge_cases array with nested objects containing 'model', 'description', and 'risk' fields, but the risk field is an enum ['HIGH', 'MEDIUM', 'LOW']. The 'description' field lacks constraints, should specify max length or format guidance (e.g. 'brief natural-language description of the edge case, 1-500 characters').
No pagination or result-limiting guidance documented. guardrail_review and guardrail_checks may return large result sets (dozens or hundreds of check results, model diffs, or SQL snippets). Tool descriptions do not specify limits, pagination parameters, or how to handle large responses. This risks context window exhaustion.
guardrail_status, guardrail_review, and guardrail_checks accept optional 'models' parameter but do not specify what happens if an invalid model name is passed. Should document: 'If a model name does not exist in the manifest, the tool skips it silently or raises an error with the list of available models.'