Ask ten site owners whether their site is accessible, and most will point to a scanner dashboard showing a passing score. That score answers a narrower question than the one being asked: it confirms that a machine found no automatically detectable violations, not that a person using a keyboard, a screen reader, or an assistive device can actually complete a task. A trustworthy answer to "is my site accessible?" comes from climbing a hierarchy of evidence, where each layer catches problems the layer beneath it cannot see, rather than from any single tool's verdict.
What an automated scan actually proves
An automated scanner compares a page's markup against a fixed rule set built from the parts of the Web Content Accessibility Guidelines (WCAG) that can be mechanically verified: color contrast ratios, missing form labels, empty links, duplicate identifiers, a missing document language attribute, and similar structural defects. These checks are genuinely useful, and running them across an entire site catches errors that would take a human reviewer weeks to find by hand, particularly on large sites with thousands of templated pages.
The limitation is not that the tool is bad; it is that most of what determines whether a page is actually usable cannot be reduced to a rule a machine can apply. Whether alternative text describes an image usefully, whether a focus order matches the logical flow of a task, and whether an error message gives someone enough information to recover are judgment calls a scanner cannot make, and neither can it reliably test forms behind a login, multi-step checkouts, or PDF documents attached to a page. The W3C's guidance on selecting evaluation tools makes this point directly: tools support evaluation, they do not replace it. Treating a clean scan as proof of accessibility mistakes the floor for the ceiling.
A manual keyboard pass reveals what a scanner cannot see
Unplugging the mouse and navigating a page using only the Tab, Shift+Tab, Enter, and arrow keys exposes a category of failure that automated tools routinely miss. A scanner can confirm that a button exists in the markup; it cannot confirm that a keyboard-only visitor can actually reach that button, see where focus currently sits, or escape a dropdown menu once it opens.
The check itself takes minutes once someone knows what to look for. Every interactive element, links, buttons, form fields, and custom widgets, should be reachable in a sensible order, should show a visible focus indicator at each stop, and should never trap the keyboard inside a modal or menu with no way out. WebAIM's keyboard accessibility techniques walk through the specific patterns worth testing, including the traps that custom-built dropdowns, carousels, and date pickers introduce most often. A site can pass every automated check and still be unusable for someone who cannot operate a mouse, which is exactly the gap this second layer is designed to close.
A screen reader spot check surfaces a different category of failure
Turning on a screen reader, such as Apple's built-in VoiceOver or the free NonVisual Desktop Access (NVDA) tool for Windows, and navigating a page by headings, links, and form fields reveals problems that neither a scanner nor a keyboard pass will catch. A heading structure that looks fine visually might skip levels or repeat the same text across a page, making it impossible to build a mental map of the layout by ear alone. An image might carry alternative text that is technically present but functionally useless, describing a decorative flourish while staying silent about the chart beside it that actually carries information.
This layer also tests the plumbing that the Web Accessibility Initiative's Accessible Rich Internet Applications (WAI-ARIA) specification exists to support: whether a custom dropdown announces its expanded state, whether a form error is read aloud the moment it appears, and whether a loading spinner tells anyone it is there at all. WebAIM's guidance on testing with screen readers is a reasonable place to learn the handful of commands a sighted tester needs to run a competent spot check without years of screen reader fluency.
Real user testing closes a gap the earlier layers cannot
Here is the honest limitation of everything above: a sighted staff member running a screen reader for the first time is testing whether a page is theoretically operable, not whether it matches how an experienced daily user actually navigates it. Years of practice build shortcuts, expectations, and workarounds that no amount of internal testing reproduces, and the same gap exists for keyboard-only users, switch device users, and people with low vision relying on extreme zoom or high contrast display modes.
Testing with people who use these tools as their daily interface, whether through a structured usability study or a smaller informal session, is the layer that catches what every earlier check was structurally unable to see. The W3C's guidance on involving users with disabilities in evaluation lays out how to recruit participants and structure sessions without an enterprise research budget. I think of the first three layers as increasingly rigorous simulations of a real visitor, and this fourth layer as the point where the simulation ends and the actual answer begins.
What each layer proves, and what it leaves unanswered
| Layer | What it proves | What it cannot answer |
|---|---|---|
| Automated scan | Machine-checkable WCAG criteria pass across the whole site | Whether content, labels, and flows actually make sense to a person |
| Manual keyboard pass | Every interactive element is reachable and escapable without a mouse | Whether a screen reader user can interpret what they land on |
| Screen reader spot check | Headings, labels, and interface states are announced correctly | Whether an experienced daily user finds the experience efficient |
| Real user testing | Whether the people the site is built to serve can complete real tasks | Whether every other page and template on the site behaves the same way |
A realistic order to run these checks in
None of this requires hiring an auditor before taking a first look. Running an automated scan across the site's core templates, then tabbing through the two or three highest-traffic pages without a mouse, then spending twenty minutes with a screen reader on a form or checkout flow, surfaces most of the obvious problems before anyone commits a procurement budget. Each layer is cheap enough to repeat regularly, and repeating them after a redesign or a new template ships is often more valuable than running any single layer once and considering the question closed.
The fourth layer, testing with people who actually use assistive technology daily, is the one most owners defer, usually because it sounds like it requires a formal study. A handful of short sessions with real users tends to reveal problems that internal review methods cannot find, because it tests what the first three layers can only approximate.
Key takeaways
- An automated scan only confirms that a page passes the subset of WCAG criteria a machine can check mechanically, not that a person can actually complete a task on it.
- A manual keyboard pass, tabbing through a page without a mouse, catches focus order and keyboard trap problems that no scanner can detect.
- A screen reader spot check with a tool like VoiceOver or NVDA reveals heading structure and labeling failures that remain invisible to sighted testers using a mouse.
- Real user testing with people who rely on assistive technology daily is the only layer that tests efficiency and real-world usability rather than theoretical operability.
- The honest answer to whether a site is accessible names which of these four layers have actually been checked rather than offering a single yes or no verdict.
Questions readers ask
Is a passing automated accessibility scan enough to call a site accessible?
No. An automated scan only verifies the subset of WCAG success criteria that can be mechanically checked, such as color contrast and missing form labels, and it cannot judge whether alternative text is meaningful, whether a focus order makes sense, or whether a form behind a login actually works for a keyboard-only visitor.
What does a manual keyboard test check that a scanner cannot?
Navigating a page using only Tab, Shift+Tab, Enter, and arrow keys reveals whether every interactive element is reachable in a sensible order, whether focus is visible at each stop, and whether a modal or menu traps the keyboard with no way to escape, none of which a scanner can confirm from markup alone.
Do I need a screen reader user to run a screen reader spot check?
A sighted staff member can run a useful spot check with a free tool like NVDA or the built-in VoiceOver to catch obvious heading and labeling failures, but this only tests theoretical operability, not the efficiency an experienced daily screen reader user would expect.
Why does real user testing matter if the earlier layers already passed?
Years of practice give experienced assistive technology users shortcuts, expectations, and workarounds that no amount of internal testing reproduces, so real user testing is the only layer that reveals whether the site actually works for the people it needs to serve.
What order should these four checks be run in?
Running an automated scan across core templates first, then a manual keyboard pass on the highest-traffic pages, then a screen reader spot check on a key flow like checkout, surfaces most obvious problems cheaply before a team commits to a formal usability study with real users.