UI component library and examples for AI-powered Next.js applications, including prompt input, tools, agents, artifacts, attachments, audio players, and database index optimization scripts
Scoring was not performed
Complex object parameters (Supabase client, Drizzle DB, Anthropic client) have no schema definitions. Parameters like 'supabase: {type: object, description: Supabase client instance}' lack property constraints, required fields, and nested type info. LLMs cannot understand or construct these objects.
Output schemas are completely undocumented. Tools like searchDocuments return data from Supabase RPC calls, hybridSearch returns hybrid search results, and run_eval returns evaluation metrics, but there is no specification of the return structure, field types, or examples. LLMs cannot plan downstream tool calls without knowing what fields to expect.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 29 | 2025-06-18+ | v2 |
| 2026-03-09 | F | 49 | - | v1 |
No error handling guidance in tool descriptions. Tools that call external APIs (OpenAI embeddings, Supabase, Claude) or perform file operations can fail in multiple ways, but descriptions do not indicate error cases or recovery steps. Per pattern:recovery-guide, descriptions must say 'If X fails, try Y'.
Multiple tools combine concerns: generate_benchmark both loads results AND aggregates AND generates JSON (three separate operations). improve_description both evaluates AND reasons AND improves. Per pattern:tool, each tool should do exactly one thing so agents can compose them independently.
Parameter descriptions lack actionable constraints. 'threshold' params say they default to 0.7 but do not specify the valid range (0.0 - 1.0 implied). 'limit' params don't state minimum (1?) or maximum (100?). Per review:param-validation-rules, format, range, and allowed values must be explicit in descriptions.
Tool names using underscores (calculate_stats, load_run_results) instead of camelCase. Most MCP servers and LLM toolkits use camelCase (getEmbedding, runEval). Inconsistency across the repo (some tools are camelCase, others snake_case) signals weak tooling governance.
No indication of which tools are idempotent vs. which have side effects. package_skill has risk=WRITE but other tools (generate_benchmark, improve_description, run_eval) also write files or invoke external services. Per pattern:idempotent-operation, LLMs need to know what is safe to retry.
Some parameter descriptions are generic boilerplate ('Path to benchmark directory', 'Array of test items with query and should_trigger fields') and do not explain WHEN to call the tool or WHAT to do if it fails. Descriptions should be prompt-engineered per pattern:tool-description.