dbt-native development, evaluation, and observability for trusted data agents
Scoring was not performed
Missing input schemas for critical tools. 'init', 'serve', and 'list_manifests' have no visible input parameter definitions despite being valid tools. Per hard scoring rules, these must score 0 for schema.
Tool descriptions are uniformly too brief and lack context. 'Initialize a new project' (27 chars) and 'Start the Flask web server' (28 chars) do not explain WHEN to use the tool, what prerequisites exist, or what the LLM should expect afterward. Rubric requires 10-1024 chars with WHO/WHAT/WHEN/WHY context.
Output schemas are not documented for any tool. LLMs cannot determine what fields to expect from responses (e.g., does list_manifests return {manifests: []} or {data: {manifests: []}}?). This forces LLMs to guess and plan poorly for downstream tool calls.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 25 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 29 | - | v1 |
Parameter descriptions are missing entirely for several tools. 'serve' has a port parameter with a description, but 'apply' and 'query' have project_folder parameters whose descriptions are minimalist ('Path to project folder'). They do not explain format constraints, allowed values, or what the tool does when the folder is invalid.
No error handling guidance. Code shows HTTP error responses (404, 400, 500) but no recovery hints. E.g., if 'Manifest folder not found' is returned, the LLM has no actionable next step, should it call init? Check the path? Try a different folder?
Tool 'query' has confusing semantics. Description says 'Start an interactive session to query manifests with an LLM.' This is ambiguous, does it launch a REPL? Does it generate a query plan? Return a result? The name 'query' conflicts with 'query_endpoint', creating naming overlap that forces LLMs to reason about which to use.
Tool 'serve' is poorly described and risky. 'Start the Flask web server' does not explain that this is a READ_ONLY tool, an agent might think starting a web server is a destructive action. No mention of what the server does, where to access it, or why an agent should call this.
Parameters lack type specificity in descriptions. E.g., 'project_folder' says 'Path to project folder' but does not specify: relative vs absolute? Required to exist? What if it doesn't? What are valid folder structures?