MCP server for running accessibility audits using axe-core and Playwright, supporting both single-URL tests and multi-step browser scenarios
Two tools with well-structured schemas and adequate descriptions, but missing critical LLM-optimization details and error handling guidance. Tool names are clear action-based verbs (exec-a11y-test, exec-a11y-test-scenario), meeting basic naming requirements. Descriptions are present but generic (194 chars average baseline suggests these are underweight). Parameter descriptions exist in schema but lack dependency hints, constraint examples, and recovery guidance. The scenario tool exhibits complex composition (10 step types in oneOf) that is properly structured but underexplained. No error classifications, retry guidance, or output field documentation visible. Zod schemas provide runtime validation but LLM-facing descriptions are sparse.
Obtains a list of specified list of URL and a list of WCAG indicators and returns the results
Run a multi-step browser scenario (navigation, click, fill, etc.) and execute one or more accessibility audits at chosen points. Useful for authenticated pages, modal/menu open states, SPA route transitions, and any UI state not reachable from a single URL. Framework-agnostic: works with any rendered DOM (Vue/React/Svelte/Angular/Lit/plain HTML/Web Components).
Descriptions lack LLM-optimization and actionable context. 'Obtains a list of specified list of URL and a list of WCAG indicators and returns the results' (exec-a11y-test) is grammatically awkward and does not state WHEN to use this tool vs the scenario variant, WHAT format the output is, or prerequisites (browser availability, network access). Baseline expectation: 50-200 chars, clear intent, and dependency hints.
Parameter descriptions are minimal and missing constraint detail. For exec-a11y-test, 'wcagStandards' is documented as 'Optional list of WCAG standards/tags to apply (e.g., 'wcag2a', 'wcag2aa', 'wcag21aa')' but does NOT explain: what happens if an invalid tag is passed, which tags are valid beyond examples, whether the default is wcag2a+wcag2aa (from constants.ts), or what each standard tests for. LLMs need explicit constraint descriptions and defaults to make intelligent choices.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
No documented output schema or return value structure. Tools return {content: [{type: 'text', text: '...'}]} but the LLM has no visibility into what fields the accessibility test results contain. Does it include: violation IDs, impact levels, affected elements, remediation steps, pass/fail counts? The constants.ts and functions.ts show ViolationSummary, AccessibilityTestOutput structs, but these are NOT surfaced in tool descriptions or response documentation visible to the LLM. Pattern baseline: 100% of A+ tools document return types.
Error handling does not guide LLM recovery. No visible error classification (retryable vs. user-fixable), actionable error messages, or next-step suggestions. If a URL is unreachable, playwright times out, or axe-core fails, what does the LLM see? Expected: 'Network timeout on https://example.com. Try again with a shorter timeout, or verify the URL is accessible.' Actual: likely a generic error string.
Scenario tool's complex oneOf step union is under-documented. The input schema defines 9 step types (goto, click, fill, select, press, hover, waitFor, waitForUrl, waitForNetworkIdle, audit) but the description does NOT explain: the order/sequence semantics, what happens if a selector is not found (step failure?), whether steps are atomic or rollback on error, how timeout cascades work (globalTimeoutMs vs per-step), or which step types are mutually exclusive. An LLM cannot intelligently compose a scenario without this clarity.
No pagination or result-limiting guidance. The accessibility test could return hundreds of violations. The tool description does NOT state: max violations returned, pagination strategy, or guidance for the LLM to handle large result sets. Baseline: tools returning lists must state result limits and pagination approach.
Missing per-step feedback in scenario execution. If step 1 (goto) succeeds but step 5 (click) fails due to missing selector, does the response include which step failed, why, and at what point the audit results were captured? The description suggests 'audit steps' can be interspersed, but feedback clarity is undocumented.