No automated accessibility scanner has detected more than half the barriers on a controlled test page in any published, independent study. The best result in the research literature is approximately 40%, recorded when the United Kingdom's Government Digital Service (GDS) tested 13 tools against 142 intentional barriers in 2018. The most recent controlled experiment, conducted by Equal Entry in partnership with Evinced in 2024, found detection rates between 3.8% and 10.6% across six scanners tested against 104 intentional defects. Adrian Roselli's 2023 comparison of free automated tools against a manually audited page found rates between 0% and 13.5%.
These figures might seem to contradict Deque Systems' widely cited finding that automated testing catches 57% of accessibility issues. They do not. The disagreement is not about which tools work. It is about what "detection rate" means, and understanding that disagreement is the most important insight this article can offer.
The published research spans more than a decade: government experiments, peer-reviewed academic studies, practitioner-run controlled tests, vendor analyses, and standards-body data. The evidence converges on a clear finding. Automated accessibility scanners are reliable within a narrow band of mechanical checks and structurally limited beyond it. The gap is not a function of tool maturity; it is a function of what accessibility requires.
13–31%
Criterion count
How much of the standard can a machine enforce?
W3C ACT; Accessible.org; Groves
57%
Issue volume
Of the bugs on real sites, how many are the kind tools find?
Deque, 2021
0–40%
Page-level detection
On this page, how many real problems did the tool catch?
GDS 2018; Roselli 2023; Equal Entry 2024
Figure 1. Three methods of measuring scanner detection, and the different question each one answers. The spread between them is a difference of denominator, not a disagreement about the tools.
The Denominator Problem: Three Ways to Measure, Three Different Answers
Ask "what percentage of accessibility issues do automated tools catch?" and you will receive answers ranging from 0% to 57%. These are not contradictions. They are answers to different questions, and understanding the difference determines whether a scanner purchase is a useful investment or a false assurance.
The criterion-count method (13–31%)
This approach asks: of all Web Content Accessibility Guidelines (WCAG) 2.2 success criteria, how many can automated tools evaluate? The World Wide Web Consortium's (W3C) Accessibility Conformance Testing (ACT) Task Force maintains approved automated testing rules for 17 of the 55 WCAG 2.2 Level A and AA success criteria, which is 31%. Accessible.org applied a stricter definition, "reliably flagged with minimal false positives," and found that only 7 of 55 criteria, or 13%, meet that bar. Karl Groves's foundational testability analysis found that at the AA level, 17% of accessibility best practices are fully automatable, 41% require human verification after machine flagging, and 41% require manual testing only.
This method measures theoretical coverage of the standard. It answers the question: how much of WCAG can a machine enforce?
The issue-volume method (57%)
Deque Systems, the developer of the axe-core accessibility testing engine, analyzed anonymized audit data from more than 2,000 first-time accessibility audits spanning 13,000 pages and nearly 300,000 individual issues. Their finding: automated testing with axe-core caught 57% of those issues by volume. The reason this number is higher is structural. The most common accessibility failures on real websites, meaning missing alternative text, insufficient color contrast, empty links, and missing form labels, are precisely the types that automated tools detect reliably. The WebAIM (Web Accessibility In Mind) Million 2026 report found that six error types account for 96% of all detected failures across one million homepages. High-frequency issues inflate the volume-weighted figure because the problems that happen the most are the problems machines catch the best.
This method measures practical impact in a typical audit. It answers the question: of the bugs that actually exist on real sites, how many are the kind tools can find?
The page-level detection method (0–40%)
This is what happens when you build a page with known barriers, or audit a page with expert review to establish a baseline, and then measure what a scanner actually flags. The GDS found a ceiling of approximately 40% in 2018. Roselli found 0% to 13.5% across six free tools in 2023. Notably, WAVE (Web Accessibility Evaluation Tool), one of the most widely deployed accessibility checkers, maintained by WebAIM, returned zero failures on a page that manual review found to have 37 distinct WCAG violations. Equal Entry found 3.8% to 10.6% across six scanners in 2024. Markel Vigo, Justin Brown, and Vivienne Conway, in a peer-reviewed 2013 study published by the Association for Computing Machinery (ACM), found maximum coverage of 50% of success criteria, with completeness, the proportion of actual violations detected within covered criteria, ranging from 14% to 38%.
This method measures real-world tool performance against a verified baseline. It answers the question: on this specific page, how many of the actual problems did the tool catch?
Why the disagreement matters
All three methods are valid. None is wrong. Any vendor, consultant, or procurement team citing a single percentage without specifying the denominator is, at minimum, incomplete and may be misleading.
Interpretation
The persistence of headline "57% coverage" claims in vendor marketing, without denominator context, functions as a form of statistical misdirection, even when the underlying data is legitimate.
The page-level detection method is the most conservative and the most directly relevant to a buyer asking: if my site has accessibility problems, how many will this tool actually find?
The Reliable Core: What Every Study Confirms Scanners Find
Across all published studies, automated scanners consistently detect the same categories of violations. These represent the reliable core of automated accessibility testing: the barrier types that every major tool catches, regardless of engine or methodology. They map to a small set of WCAG success criteria:
- 1.1.1 Non-text Content: missing
altattributes on images - 1.4.3 Contrast (Minimum): text-to-background luminance ratios below the 4.5:1 threshold for normal text (3:1 for large text)
- 1.3.1 Info and Relationships: form inputs without associated labels
- 3.1.1 Language of Page: missing
langattribute on the HyperText Markup Language (HTML)<html>element - 2.4.4 Link Purpose: empty links with no discernible text
- 4.1.2 Name, Role, Value: buttons and controls without accessible names
Figure 2. The six error types every scanner reliably detects, and the share of one million homepages carrying each. Source: WebAIM Million, 2026. Together these account for 96% of all detected failures.
These six criteria are overwhelmingly mechanical. They involve checking whether a specific attribute exists, whether a computed value exceeds a threshold, or whether a Document Object Model (DOM) element contains required child content. A machine can evaluate these reliably because the success or failure condition is fully specified in the code.
These are also, not coincidentally, the six error types that dominate the WebAIM Million year after year. The scanners are well-calibrated for the problems that occur most frequently. The reliable core is real, it is useful, and it is narrow.
Finding
Despite the scanners' ability to detect these six error types, the WebAIM Million 2026 report found detectable WCAG 2 A/AA failures on 95.9% of the top one million homepages, with an average of 56.1 errors per page, a 10.1% increase from the previous year. The web is getting worse on the metrics automation can measure.
Automated detection is necessary but clearly insufficient: the problems scanners can find are pervasive, and finding them has not been enough to reduce their prevalence.
The Automation Boundary: What Published Research Shows Scanners Miss
The Accessible.org criteria-level analysis classified 23 of 55 WCAG 2.2 AA criteria (42%) as entirely invisible to automation. The GDS experiment confirmed that all 13 tools missed reading order issues, lack of visible focus indicators, and context-dependent use of color. Roselli found that no automated tool detected a single AA-level violation on the page he tested, while his manual review recorded failures against 18 success criteria in total: 11 at Level A and 7 at Level AA. His expert review caught roughly seven and a half times more issues than the best automated tool, across three times as many success criteria.
- Reliable7 criteria
- Fully specified in code: contrast ratios, attribute presence, accessible names.
- Partial25 criteria
- A machine-checkable shell around a human-judgment core. Heading order can be checked; heading accuracy cannot.
- Invisible23 criteria
- Meaning, sequence, and interaction. Reading order, focus management, navigational consistency.
Figure 3. All 55 WCAG 2.2 Level A and AA success criteria classified by automation detectability. Source: Accessible.org, 2025, cross-referenced against W3C ACT rule coverage.
The barrier categories that no published study has found scanners detecting include:
- Meaningful sequence (WCAG 1.3.2): whether the visual order of content matches the DOM order, and whether Cascading Style Sheets (CSS) reordering creates a reading sequence that makes no sense when linearized. Both the GDS and Roselli experiments found zero automated detection for this criterion, regardless of whether the issue manifested as DOM-order divergence or CSS-driven visual reordering.
- Focus management (WCAG 2.4.3 Focus Order): whether focus moves to the appropriate element after modal dialogs open, content changes dynamically, or interactive widgets update. Deque's own data shows a 0% automated detection rate for focus order.
- Content on hover or focus (WCAG 1.4.13): whether tooltip content triggered by hover can be dismissed, hovered over, and persists until the user removes it.
- Error suggestion quality (WCAG 3.3.3 Error Suggestion): whether form validation messages are present is machine-checkable, but whether the suggestion actually helps a person fix the problem is not.
- Consistent navigation (WCAG 3.2.3): whether navigation elements appear in consistent positions across pages. No automated tool evaluates cross-page positional consistency.
These categories share a structural property: evaluating them requires understanding what the content means to a human user, not just what the DOM contains. A tool can verify that a focus indicator exists; it cannot verify that focus moves to the right place after a state change. A tool can confirm that an error message is present; it cannot judge whether the message helps a person fix the problem.
The challenge is growing, not shrinking. The WebAIM Million 2026 report found that pages using Accessible Rich Internet Applications (ARIA) attributes averaged 59.1 detectable errors, compared to 42 errors on pages without ARIA. This correlation does not mean ARIA causes accessibility problems; it more likely reflects the fact that ARIA-heavy pages involve more complex interactive patterns with more opportunities for failure. ARIA usage increased 27% year over year in the 2026 report. As web applications grow more complex, the proportion of barriers requiring human judgment to evaluate grows with them.
The Keyboard Gap
Deque's own 2021 study of nearly 300,000 issues found a 2.49% automated detection rate for keyboard accessibility (WCAG 2.1.1 Keyboard) and a 0% rate for focus order (WCAG 2.4.3). The GDS experiment confirmed that keyboard-interaction barriers were among the categories no tool detected. Roselli's manual review caught keyboard and focus issues that every automated tool missed.
Figure 4. Automated detection rates for keyboard-related criteria, against contrast as a reference point. Source: Deque, 2021, nearly 300,000 issues; the GDS 2018 experiment recorded zero keyboard detection across all 13 tools.
This is the most consequential gap in automated accessibility testing. Keyboard navigation is the primary interaction method for people who cannot use a mouse: users with motor disabilities, many screen reader users, and users with temporary injuries. When a website traps keyboard focus inside an invisible element, when a modal dialog opens without moving focus into it, when a "close" button cannot be reached by tabbing, the site is functionally broken for these users. Automated scanners, as currently constructed, cannot see the breakage.
Inference
Keyboard accessibility may be the single area where the gap between automated detection and real-world impact is widest. The barriers are severe, the affected user population is large, and the detection rate is near zero.
False Positives: The Other Side of the Ledger
Detection rate alone is an incomplete measure of scanner performance. A tool that flags real barriers but also generates hundreds of false positives creates operational costs that can outweigh its detection value.
Equal Entry's 2024 study documented this trade-off at scale.
| Scanner | Found (of 104) | Detection rate | False positives |
|---|---|---|---|
| D | 11 | 10.6% | 2 |
| B | 9 | 8.7% | 12 |
| E | 9 | 8.7% | 3 |
| C | 7 | 6.7% | 46 |
| A | 5 | 4.8% | 63 |
| F | 4 | 3.8% | 474 |
Scanner F generated 474 false positives while finding only 4 real issues, a ratio of nearly 119 false positives per true detection. Four of the six tools produced more false positives than true defects. Scanner D, the most accurate, found 11 real issues with only 2 false positives.
The axe-core engine, which powers both Deque's axe DevTools and Google Lighthouse, addresses this through a design principle: any finding that cannot be stated with certainty as a violation is categorized as "needs review" rather than reported as a failure. This means axe-core's precision is high, since every flag is actionable, but its recall is lower, leaving more real barriers unflagged. The design choice prioritizes trust over coverage. Microsoft Accessibility Insights, also powered by axe-core, follows the same principle.
For procurement decisions, this means detection rate cannot be evaluated in isolation. A scanner's value is a function of what it catches, what it misses, and how much time the team spends investigating findings that turn out to be wrong.
The AI Question: What "AI-Powered" Detection Adds
Several scanner vendors now market AI-powered detection capabilities, including Deque's Intelligent Guided Tests (IGT), Siteimprove's AI-enhanced scanning, and various overlay products. The marketing claims are substantial: Deque states that with IGT, a semi-automated, human-in-the-loop workflow, coverage reaches approximately 80%.
The published evidence for these claims is thin. No independent study has isolated the contribution of AI features by comparing AI-on versus AI-off detection rates against a controlled baseline. The Equal Entry study anonymized its scanners, making it impossible to attribute performance differences to specific technological approaches. Practitioners surveyed by Maria Korneeva for heise online estimated that 20% to 40% of barriers are technically detectable by automated tools, a range that has not shifted substantially with the introduction of AI marketing.
What is known: the 80% figure Deque cites for IGT includes human-in-the-loop testing, which is semi-automated rather than fully automated. The distinction matters. A tool that flags potential issues for human verification is fundamentally different from a tool that detects violations autonomously. The former is a workflow enhancement; the latter is what buyers typically understand "automated detection" to mean.
Inference
AI-powered analysis may incrementally extend the automation boundary. The structural constraint, however, is that many WCAG criteria are not code-checkable properties at all. They are evaluations of meaning, intent, and experience quality. Incremental AI improvements will compress the partially-detectable category without substantially reducing the undetectable one.
The Federal Trade Commission's Warning
In April 2025, the Federal Trade Commission (FTC) finalized a $1 million enforcement action against accessiBe for deceptively marketing its AI-powered overlay widget (accessWidget) as capable of making any website WCAG 2.1 AA compliant. The FTC found that accessWidget did not make websites WCAG-compliant. More damaging still, the FTC found that accessWidget created accessibility barriers for users with disabilities on some websites where it was deployed. AccessiBe was also found to have misrepresented third-party reviews as independent endorsements. The company was barred from making unsubstantiated compliance claims for 20 years.
This enforcement action concerned an overlay vendor, not a scanner vendor. The relevance is in what the FTC's finding implies about the broader market for AI accessibility claims. If the line between "our tool detects accessibility issues" and "our tool makes your site compliant" becomes blurred in marketing, the FTC has demonstrated its willingness to intervene.
The regulatory context extends beyond enforcement. The European Accessibility Act (EAA) took effect in June 2025, establishing accessibility requirements for specified categories of digital products and services in the European Union, among them e-commerce, banking, transport, telecommunications, e-books, and certain self-service terminals. The United States Department of Justice (DOJ) finalized an Americans with Disabilities Act (ADA) Title II rule requiring state and local government websites to conform to WCAG 2.1 Level AA. Both regulations reference WCAG as the compliance standard, and neither exempts organizations using automated scanning from the obligation to meet criteria that scanners cannot evaluate.
Interpretation
The FTC's action, combined with the detection-rate data from the published literature, suggests that "AI-powered accessibility compliance" is a marketing category, not a technical capability. The tools within that category vary from genuinely useful, meaning scanners with transparent detection rates and documented limitations, to demonstrably misleading. Buyers should evaluate the detection data, not the marketing language.
Why the Gap Is Structural, Not Temporary
It would be reassuring to treat these findings as a snapshot of tools that will improve. Tools will improve. The gap, however, is not primarily a function of tool maturity; it is a function of what accessibility requires.
WCAG does not only measure code quality. It measures whether a human can use a website. Success criterion 2.4.3 (Focus Order) does not ask whether a tabindex attribute is present; it asks whether the focus sequence "preserves meaning and operability." The criterion's reliance on meaning requires understanding what the content is trying to do. Success criterion 3.3.3 (Error Suggestion) does not ask whether an error message exists; it asks whether the suggestion helps the user fix the problem. These are conditional, context-dependent evaluations.
The W3C's ACT rules cover 17 of 55 WCAG 2.2 AA criteria. Karl Groves found that 41% of AA best practices require manual-only testing. Eric Eggert, a former W3C Web Accessibility Initiative (WAI) team member, has argued that accessibility involves diverse user interactions through various assistive technologies, and that automated tools cannot evaluate these interaction paths. The professional consensus is clear: the automation boundary exists because WCAG tests human experience, and testing human experience requires humans.
What This Means for Buyers
If you are evaluating accessibility scanners for enterprise procurement, the published evidence supports several concrete conclusions.
Every scanner is worth running; no scanner is sufficient
The worst tool in the GDS experiment still caught 13% of known barriers that would otherwise require manual discovery. The best caught approximately 40%. Automated scanning is a necessary layer, not a complete strategy.
Running multiple scanners improves coverage, with diminishing returns
The Vigo et al. study found that no single tool excelled across all evaluation dimensions: coverage, completeness, and correctness. Different tools catch different subsets of barriers, so running multiple tools will find more issues than any single tool alone. The coverage gain from each additional tool, however, decreases as the tools converge on the same reliable core, the six high-frequency error types that all tools catch, and diverge unpredictably on everything else.
Evaluate false positive rates alongside detection rates
Equal Entry's data shows that some scanners generate more false positives than true detections. A tool that produces 474 false flags while finding 4 real issues requires a team with the expertise and time to triage them. If your team lacks that capacity, the false positives will either consume disproportionate time or be ignored, taking real findings with them.
Budget for manual testing
The barriers that no scanner catches are not optional. They include keyboard navigation, focus management, error messaging quality, and navigational consistency. These barriers most severely affect users who rely on keyboards, screen readers, and other assistive technologies. If your accessibility program relies solely on automated scanning, the published evidence shows that 42% of WCAG criteria are entirely invisible to automation, and the barriers scanners do detect are not being eliminated, since the WebAIM Million shows them increasing year over year.
Distinguish scanners from compliance solutions
The FTC's action against accessiBe drew a regulatory line. A scanner that detects a subset of barriers is a useful diagnostic tool. A product marketed as a compliance solution at those detection rates is making a claim the published data does not support. The regulatory landscape, meaning the FTC, the EAA, and the DOJ's ADA Title II rule, is tightening, and none of these frameworks accept automated scanning alone as evidence of compliance.
What the Evidence Shows
Accessibility scanners work. Within their reliable core, they catch real barriers quickly, consistently, and at scale. The six error types they all detect are the same six that afflict the vast majority of production websites, and catching those errors in a continuous integration and continuous deployment (CI/CD) pipeline or a pre-launch audit is materially better than not catching them at all.
Accessibility scanners are also incomplete. They miss the majority of WCAG criteria, including the criteria that most severely affect users who rely on keyboards, screen readers, and other assistive technologies. This incompleteness is structural, not a deficiency that better algorithms or larger training datasets will resolve in the near term. The automation boundary exists because WCAG measures human experience, and measuring human experience requires human judgment.
The measurement disagreement this article describes, spanning 13%, 31%, 57%, and published page-level rates of 0% to 40%, is not a debate to resolve. It is a framework to understand. Each percentage answers a different question, and the answer a buyer needs depends on the question they are asking. "What fraction of the standard can automation enforce?" is a different question from "How many of my site's current bugs will a scanner find?" which is a different question from "What share of a typical audit's findings are machine-detectable?" Distinguishing these questions produces clarity; conflating them produces false confidence.
Finding
The most consequential result in the published literature is not any single detection rate. It is the keyboard gap: automated tools detect between 0% and 2.49% of keyboard-navigation barriers, according to Deque's own data and confirmed by the GDS experiment. The tools most organizations rely on to find accessibility problems are nearly blind to some of the most impactful barriers.
Keyboard access is not a niche concern. It is the primary interaction method for users with motor disabilities, many screen reader users, and anyone temporarily unable to use a pointing device.
Run the scanners on every build, every page, and every release. Treat their output as the minimum, not the maximum, of what you need to know about your site's accessibility.