Automated accessibility scans need a false-positive budget: protected human review time for signals that a rules engine cannot resolve from code alone. The budget is not permission to tolerate defects; it is an operating rule that separates confirmed, reproducible failures from findings whose truth depends on context, behavior, or meaning. Deque’s axe-core results-object documentation distinguishes violations from incomplete results, and that distinction should determine how a team spends specialist attention.

This is an evidence-led analysis of published accessibility standards and scanner documentation, not a benchmark, a measured false-positive rate, or a claim about any organization’s backlog. The central conclusion is narrower: a scan is evidence of what one rules engine could determine in one inspected state, whereas accessibility evaluation also requires human tasks. The World Wide Web Consortium (W3C) Easy Checks guidance combines tool-assisted checks with reviewer assessment for precisely that reason.

The operational error is treating one queue as one kind of work

Accessibility quality assurance (QA) becomes difficult to manage when a dashboard’s raw finding count is treated as a risk measure. A missing programmatic association can often be reproduced, assigned, and fixed through a normal engineering path. An uncertain result about a changing visual treatment or a keyboard interaction can require several states, content variants, and task paths before a reviewer can decide whether a defect exists. The W3C guidance on evaluating web accessibility recommends using multiple evaluation approaches, rather than treating automated output as a final verdict.

A false-positive budget therefore allocates review capacity according to uncertainty, not alert volume. It makes a previously invisible cost visible: the time needed to establish whether a signal represents a defect, an exception, or an untested state. That is a governance choice, not a requirement supplied by WCAG, and its size should reflect local release cadence, queue age, recurrence, and the importance of affected user journeys.

The useful question is not whether scanners are reliable in the abstract. It is what evidence a particular rule has produced, what it cannot know, and who is responsible for the remaining judgment. Once those questions are answered before findings enter a backlog, routine repairs can move quickly without converting unresolved uncertainty into a closed ticket.

Exhibit: a scan finding carries different evidence at each layer

Accessibility scan findings compared by the evidence available to an automated rule
Finding categoryEvidence a rule can inspectJudgment still requiredAppropriate next step
Structural presenceAttributes, associations, identifiersApplicability and user impactReproduce and route routine fixes
Static text contrastSpecified foreground and background valuesRendered state and content scopeValidate against W3C Contrast (Minimum) guidance
Variable visual backgroundA sampled rendered conditionLegibility across relevant statesReview the affected experience
Accessible Rich Internet Applications (ARIA)Roles, states, and propertiesInteraction and announced behaviorReview against the W3C ARIA Authoring Practices Guide
Headings and focus orderElement order and heading sequenceWhether the task flow makes senseReview structure and keyboard path
The exhibit separates inspectable code evidence from the contextual assessment that follows a flag or incomplete result.

Presence checks establish conditions, not equivalence

A structural presence check asks a narrow question about the Document Object Model (DOM), the browser representation of page elements. A rule can identify an image without an alt attribute, an input without an associated label, or duplicated identifiers. These are inspectable conditions; however, the W3C explanation of WCAG Success Criterion 1.1.1, Non-text Content frames text alternatives in terms of equivalent purpose, which is a separate semantic judgment.

Alt text exposes the boundary cleanly. Software can determine whether an alt attribute is present, yet it cannot settle whether the text serves the image’s function without knowing why that image appears in that location. An empty attribute can be correct for a decorative image and inadequate for an informative one, so attribute presence alone should not be confused with an accessibility verdict.

Heading signals require the same restraint. Tools can identify heading elements and flag a sequence such as h2 followed by h4, but the sequence by itself does not establish that a page is incoherent. The W3C headings tutorial describes headings as a means of communicating organization, and only a reviewer can assess that organization against the page’s actual sections and labels.

Static color contrast is often more amenable to automation when foreground and background values remain stable. The W3C explanation of WCAG Success Criterion 1.4.3, Contrast (Minimum) defines the text-contrast requirement, while axe-core’s documented result model includes incomplete outcomes for checks it cannot decide automatically. Text placed over images, video, or changing treatments needs examination across relevant rendered states rather than a universal decision based on one sample.

Incomplete results need an explicit review service

Axe-core defines incomplete findings as checks that require further testing, rather than confirmed violations, in its official results documentation. That label should alter the workflow. A lightweight closure process suited to a routine code repair may record a status for an incomplete result while leaving the original uncertainty unresolved.

This is an inference about queue design, not an assertion that every team mishandles scanner output. Still, combining confirmed failures and unresolved questions in one operational measure combines different kinds of labor. The backlog may appear precise while concealing the review time needed to establish the status and significance of each uncertain signal.

Suppression is a useful test of whether a team understands its coverage. Disabling a noisy rule category removes that category’s future signals; it does not make the underlying question disappear. The decision should document the condition no longer surfaced automatically, assign an owner, and identify another evaluation method, consistent with the W3C multiple-method evaluation guidance.

Build the budget into release governance

Start with each scanner’s published result definitions. Axe-core defines violations, incomplete findings, and related categories in its results documentation; labels from another product should not be assumed to carry identical meanings. Classify results by the question they answer: markup presence, computed presentation, interaction pattern, or semantic appropriateness.

Route routine, reproducible code findings into the delivery path that can correct them. Reserve reviewer capacity for results requiring context, expressed as a weekly allocation, a service-level target for incomplete findings, or a release gate for high-risk interactions. ARIA warrants that path because it is a specification for exposing roles, states, and properties to assistive technologies, not a guarantee of usable interaction. The W3C ARIA specification defines that semantic layer, while the ARIA Authoring Practices Guide documents interaction and keyboard patterns.

Method note: This article synthesizes published standards and scanner documentation into a triage framework. It does not report a comparative scanner test, measured false-positive rates, or evidence about any organization’s backlog.

Key takeaways

  • Automated accessibility scans provide different kinds of evidence, so their findings should not all receive the same workflow.
  • Axe-core distinguishes confirmed violations from incomplete findings requiring further testing in its results documentation.
  • Presence checks can establish narrow markup conditions, while equivalence, structure, and interaction behavior require contextual assessment.
  • A false-positive budget makes human judgment a planned accessibility QA cost rather than an undefined remainder after routine fixes.
  • Disabling a noisy rule category is a coverage decision that needs a documented alternative evaluation method.

Questions readers ask

What is the difference between a scan violation and an incomplete result?

A violation is a result the rules engine reports as a failure, while an incomplete result is one it reports as needing further testing. Deque defines both categories in its axe-core results documentation.

Can an automated scan decide whether alt text is good?

It can inspect markup conditions, but whether a text alternative serves an equivalent purpose depends on context. The W3C explanation of Non-text Content describes that purpose-based requirement.

Do skipped heading levels prove a page is badly structured?

No. A tool can identify the sequence, but the meaning of the page structure still needs review. The W3C headings tutorial explains the organizational role headings are meant to communicate.

Does valid ARIA remove the need for manual review?

No. ARIA supplies semantics, while the interaction still needs to work in the experience. The W3C ARIA Authoring Practices Guide is a reference for those patterns.