The accessibility layer for AI coding agents — MCP server + CLI for WCAG 2.2: express static audit, rendered audit (contrast, 2.5.8 target size, focus), contrast math (pairs, CSS names, hsl/rgba, text-over-image pixel sampling), EAA/RD 1112/2018 declarations, snapshot regression diffs, aria-live monitor. Zero dependencies at core; es/en.
a11y-toolkit presents well-structured accessibility tools with strong semantic naming (verb_noun pattern) and comprehensive descriptions. All 23 tools follow consistent naming conventions (a11y_* prefix with clear action verbs: audit, suggest, generate, validate). Descriptions range 150-500 chars, well above the 10-char minimum and within the production baseline of 194 chars average. Input schemas are present for all tools with typed parameters and enum constraints where appropriate (lang: [es, en] throughout; browser enum in a11y_audit_dom; action enum in a11y_ledger). However, output schemas are NOT explicitly documented in the visible code, the server.py excerpt shows only tool registration with input schemas, not return types. This is a critical gap: LLMs cannot reason about downstream tool chaining without knowing what fields are returned. Error handling is minimal in the visible code; no recovery guidance or actionable error messages are apparent. Parameter descriptions are generally good (mode explanations, timeout defaults documented), but dependencies between parameters (e.g., informe being optional in a11y_disprove) lack explicit documentation. The ledger tool conflates three distinct actions (record/gaps/summary) into one tool, violating single-responsibility principle. Tools are well-composed overall, with good chaining potential (snapshot/diff workflow, audit/disprove loop), but the lack of visible output schema documentation and error handling prevents a higher score.
Injectable live-region monitor: returns a <script> that monitors aria-live announcements and logs them, useful for verifying screen-reader dynamic content announcements.
Deep RENDERED WCAG audit via local Playwright/Chromium: real computed text contrast against effective backgrounds with alpha compositing (1.4.3), minimum target size 24×24 (2.5.8, new in WCAG 2.2), visible focus indicator heuristic (2.4.7), plus rendered versions of the static checks (alt, accessible names, labels, headings, lang/title, tabindex, aria-hidden, captions, tables). Findings include remediation. Requires playwright: pip install playwright && playwright install chromium.
Express WCAG 2.2 audit of a URL or an HTML string: 20+ automated signals with a weighted 0-100 score — images without alt (1.1.1), controls without accessible names (4.1.2), form fields without labels (3.3.2), missing autocomplete on user-data fields (1.3.5), click handlers on non-interactive elements (2.1.1), unknown ARIA roles and broken aria-labelledby (4.1.2), duplicated unnamed landmarks, timed meta refresh (2.2.1), missing skip mechanism (2.4.1), lang/title (3.1.1, 2.4.2), heading structure (1.3.1), blocked zoom (1.4.4), captions (1.2.2), autoplay audio (1.4.2), generic/duplicated link text (2.4.4), target=_blank without warning (3.2.5), positive tabindex (2.4.3), aria-hidden on focusable elements, tables without th, duplicate ids, duplicate accesskeys. Each finding includes concrete remediation. Filter, not verdict: automation covers ~1/3 of WCAG; query a11y_criterion for what a criterion means.
Deterministic safe fixes: auto-corrects obvious HTML/CSS issues (missing alt, heading nesting, duplicate ids, aria-hidden on interactive, lang attribute, etc.) without changing intent or layout.
Output schemas not documented. The visible server.py shows input schemas for all tools but no return type documentation. LLMs cannot plan downstream tool chaining (e.g., what fields does a11y_audit_url return?) without knowing the response structure. This violates pattern:tool-description and forces LLMs to guess or perform discovery calls.
a11y_ledger conflates three distinct actions (record/gaps/summary) into a single tool. This violates single-responsibility principle (pattern:tool). The action enum parameter forces the LLM to know which mode to use, increasing cognitive load. Split into three tools: a11y_ledger_record, a11y_ledger_gaps, a11y_ledger_summary.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | 2026-07-28+ | v2 |
Honest SVG badge (accessibility score with date and scope, no false promises). Score colors: ≥90 green · 75-89 lime · 50-74 orange · <50 red. Role=img + <title>, zero dependencies.
TEXT OVER IMAGE contrast: pixel-level sampling of the real background behind the text box → worst/median/p95 ratio, % of area passing AA, and automatic hostile-zone detection on a 3×3 grid (zona_peor). What pair-only checkers cannot do. region="x,y,w,h" recommended.
Exact WCAG contrast ratio for a color pair: per-criterion verdicts 1.4.3 (AA), 1.4.6 (AAA), 1.4.11 (non-text). Accepts #hex, rgb(), hsl(), CSS color names; rgba/hsl with alpha is composited over the background. If AA fails, suggests the nearest passing color.
WCAG 2.2 criterion explained: name, level, what it requires, typical failures, manual verification steps, which a11y-toolkit tools help check it.
Regression diff between two snapshots: added/removed/renamed interactive elements, focus order changes, accessibility tree changes. Exit 0 ok / 2 regressions.
One-call regression diff: snapshot two URLs and compare them (staging vs production). Useful for pre-deploy verification.
Re-verify findings against the live page: marks each finding as confirmed/rejected based on whether it reproduces. Catches false positives and regressions. Can accept an existing audit report or perform a fresh audit + disprove.
Countersignature-ready evidence pack: aggregates audit findings with verification metadata for legal/compliance documentation and external auditor review.
Form errors (criteria 3.3.1 / 3.3.3): submits invalid data and judges error announcement. Requires Playwright.
Generate a legally-sound accessibility declaration (RD 1112/2018 Spain or European Accessibility Act) based on audit findings and remediation status.
Tooltip Escape dismissibility (criterion 1.4.13): checks that hover/focus-triggered content can be dismissed with Esc and is persistently hoverable.
W3C Nu parser view: HTML markup validation and structure checker for semantic correctness.
Keyboard-trap detection (criterion 2.1.2): real Tab walk, cycle detection, Escape-release test. Requires Playwright.
Persistent coverage ledger: track accessibility findings over time (record, gaps, summary). Compares baselines to detect new issues and regressions.
320px reflow check (criterion 1.4.10): verifies content reflows without horizontal scroll at minimum mobile viewport. Requires Playwright.
Infinite-scroll audit: detects unannounced content loading patterns that break screen-reader linearity. Requires Playwright.
Interactive elements + tab order: capture a URL snapshot for regression testing. Includes interactive element names, DOM order, real focus order (tab walk), and computed accessibility tree (what a screen reader announces).
What a blind user hears: linearized accessibility tree (screen reader output) for a rendered page. Requires Playwright.
Nearest opaque color (RGB distance) to fg that reaches the target ratio against bg (4.5 default). Returns the color, its ratio, and the adjustment vector.
Error handling not visible in code. No recovery guidance, actionable error messages, or error categorization (retryable vs fatal) documented. Tools that fail (e.g., a11y_audit_dom without Playwright installed, a11y_audit_url on network timeout) will return raw errors without guidance on what the LLM should try next.
Optional parameter dependencies not documented. Tools like a11y_disprove accept optional 'informe' (existing audit report path) but the description does not explain what happens if omitted (does it run a fresh audit?). Parameter relationships should be explicit per pattern:tool-description.
Tool a11y_audit_url and a11y_audit_dom both perform audits but with different scopes (static vs rendered). The naming does not make this distinction obvious to LLMs unfamiliar with accessibility tooling. Consider renaming to a11y_audit_static and a11y_audit_rendered for clarity, or document when to use each in the tool descriptions.
Prompts feature enabled (4 declared in docstring) but not visible in source code excerpt. Prompts should be registered in the MCP server with descriptions and arguments. Without visible prompt definitions, cannot verify they follow pattern:prompt-template or contain LLM-optimized guidance.