The server provides 4 well-named, action-verb tools with clear descriptions and detailed input schemas using Zod validation. All tools have descriptions (range: 86-154 chars, within the 10-1024 baseline). All parameters are typed and mostly described. However, output schemas are undocumented, responses are returned as unstructured text rather than structured objects, forcing LLMs to parse free-form text. Error handling is basic (generic fallback messages like 'No results found'). No tool annotations (readOnlyHint, destructiveHint, idempotentHint) are present, missing critical metadata about side effects and retry safety. The generate_code tool modifies state locally but lacks explicit documentation of immutability and retry semantics. Parameter defaults are reasonable (maxResults=10), but some edge cases lack clear guidance (e.g., what happens if 'fields' is empty in generate_code?).
Get the right Duxt CLI command for a task. Describes exact commands, flags, and what they do.
Generate ready-to-use Dart code following Duxt framework patterns.
Look up API documentation for a specific Duxt component (from duxt_html, duxt_ui, or duxt_icons).
Search Duxt framework documentation. Returns matching pages with relevance scores.
Output schemas are not documented. All tools return unstructured text (ContentBlock with type='text') rather than structured objects with typed fields. LLMs must parse free-form text to extract actionable data, increasing token waste and hallucination risk.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). While all tools are READ_ONLY, this metadata is not declared in the schema. Agents cannot determine retry safety or side-effect scope from the definition alone.
generate_code description lacks clarity on state modification. Description says 'Generate ready-to-use Dart code' but does not clarify whether the code is returned only (no write to disk) or written to the file system. Missing: 'Immutable, returns code only; does not modify the project.' Without this, agents may assume the tool persists changes.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | A | 80 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 20 | - | v1 |
Error handling is generic and non-actionable. Example: 'No results found for query. Try broader terms like routing, orm, components, signals, or cli.' suggests terms but does not guide the agent on what to do next (retry? use a different tool? ask the user?). No error categorization (retryable vs user-fixable vs fatal).
Parameter 'fields' in generate_code is optional but semantics are unclear. If omitted, does the tool generate an empty model? A model with default fields (id, created_at)? The description does not explain the fallback behavior. For a code-generation tool, this ambiguity is risky.
get_component_api fallback behavior is vague. When a component is not found, the tool returns the full component API reference document as fallback. This could be thousands of lines, wasting context. Better: return a structured 'component not found' response with a list of available components in that package.