Ensemble testing for front-end web quality with accessibility and WCAG compliance testing
Kilotest exposes 10 tools with severe definition quality gaps. Tool names lack action verbs, descriptions are extremely brief (averaging 8-10 words), and input schemas are present but poorly structured. Most critically, parameters use free-form strings instead of enums for constrained inputs, descriptions lack guidance on when/why to use tools, and error handling is absent. The server treats web form handlers as MCP tools, blurring the boundary between HTTP handlers and agent-callable tools. Only listReports, listIssues, listRules, and listDiagnoses follow a sensible list_* pattern; others (ai0BalanceForm, deleteNotesForm, enqueue) expose implementation details (HTTP methods, form submission) that belong inside tool handlers, not in the tool signature.
Serves a form for recording the AI service 0 balance
Serves a form for deleting web tutorial comments, AI tutorial comments, and MCP feature requests
Implements a test request approval and returns a revised request page
Returns a test order form
Returns a form for deleting sole reports
Returns a form for hiding a report
Lists the diagnoses of a violator of an issue in a report
Tool names do not start with action verbs. Forms (ai0BalanceForm, deleteNotesForm, enqueueForm, expungeReportsForm, hideReportForm) expose HTTP form terminology instead of agent intents. Should be update_ai0_balance, delete_notes, request_test_enqueue, delete_reports, hide_report.
Descriptions are extremely brief (8-10 words average) and lack context on when/why to call the tool. Example: 'Lists the diagnoses of a violator of an issue in a report' does not explain prerequisites, what 'violator' means, or what data to pass. Critical for LLM tool selection.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 44 | 2026-07-28+ | v2 |
List the issues in a report
Lists all available reports
List the rules belonging to an issue
Parameters accept free-form strings where enums are required. Example: pageArgs accepts arbitrary 'colon-separated' values with no validation; search parameters are unvalidated query strings. LLMs will hallucinate invalid combinations.
HTTP method parameter ('method': 'HTTP method (GET or POST)') exposed in tool signature. This is implementation detail that should be hidden inside the tool, the agent should not need to choose between GET and POST.
No error handling guidance. Tools have risk levels (DESTRUCTIVE for deleteNotesForm, expungeReportsForm) but no confirmation pattern, dry-run mode, or error recovery instructions. Agents will blindly call destructive tools.
Output schemas not visible in tool definitions. ListDiagnoses, listIssues, etc. have no documented return type. LLMs cannot plan downstream tool calls or know what fields to extract.
Parameter 'search' is overloaded across tools with different query string structures (authCode, newBalance, note, report). No documentation of expected query parameter names or formats. LLMs will guess wrong.
Credentials passed via query strings. Parameters like 'authCode' in search/query params are security antipatterns, credentials appear in logs, browser history, and agent traces. Must use server-side secret injection.
Tool definitions inferred from code imports (answer functions imported as handlers) rather than explicit MCP tool registration visible in mcp.ts or index.ts. Cannot verify actual MCP schema registration without seeing handleMCP implementation.