A WCAG audit report answers the question of whether a button has an accessible name. It does not answer the question of whether anyone finds it. Between those two sentences sit findings that can become expensive later: the search that is marked up correctly and is still abandoned; the checkout that meets every success criterion and breaks at a point nobody considered during design. Places like that only become visible when someone uses the site who has to use it - with a screen reader, with heavy magnification, with the keyboard alone. A usability test with disabled participants does not replace the WCAG audit; it answers the question the audit leaves open. This article describes how such a round is organised: recruitment, consent, task script, logging, interpretation - and the four mistakes that make a round worthless.
Key takeaways
- Automated tools cover part of the criteria, everything else requires human judgement (W3C). A usability test closes exactly that gap instead of repeating the automated scan.
- Evaluating with disabled users finds usability issues that conformance evaluation alone does not find (W3C) - and the reverse holds too. Only both routes together produce a complete picture.
- Small rounds deliver solid observations, but no statistical significance (W3C). The result is a list of failures with evidence, not a score or a pass rate.
- Participant technology is part of the setup: 71.6 percent of respondents use more than one desktop screen reader (WebAIM). Prescribing the choice measures something other than real use.
- Recording a session touches health data. It is only permissible with explicit consent for specified purposes (Article 9(2)(a) GDPR), and participants may stop at any time (W3C).
- Test findings belong in the same remediation plan as audit findings. With two separate lists, one of them is easily left untouched.
What an automated check cannot see
The limit is not controversial; it is written into the W3C guidance on selecting evaluation tools: tools cannot check all accessibility aspects automatically, and human judgement is required for the rest (W3C). That is not an objection to the tools. A scan finds in minutes what takes a person hours - missing alternative text, insufficient contrast, unlabelled form fields, duplicate identifiers. What it cannot do is judge: whether an existing alternative text matches the purpose of the image, whether the heading order follows a comprehensible structure, whether an error message is understood, and whether a path through a checkout with a screen reader ends within a reasonable time. Those are exactly the questions that decide whether a site is usable.
The second limit lies in the question being asked. A conformance evaluation asks about rules; a usability test asks about tasks. The W3C describes the difference in its guidance on involving users: evaluating with users with disabilities and with older users identifies usability issues that conformance evaluation alone does not discover (W3C). The reverse applies as well: a test round will not find a missing language attribute in the source or a contrast ratio just below the threshold. Treating the two routes as alternatives means choosing between two incomplete pictures. Which method delivers what is covered in our overview of methods and evaluation approaches; this article describes the part that cannot be automated.
Three questions no scan answers
Who to invite - and who not to
At the end of 2023, around 7.9 million people with a severe disability lived in Germany (Federal Statistical Office). For a test round that figure is only background; what matters is the composition. A round in which every participant is blind and uses the same software describes one mode of use, not use in general. So the seats are filled along access routes: screen reader on the desktop, screen reader on mobile, heavy magnification, keyboard-only operation, voice or switch control, plus participants with reading or concentration difficulties. We assemble rounds like this as part of a moderated screen reader test. Invitations are best sent through associations, self-advocacy organisations and university groups rather than through a notice on your own website.
The second dimension is experience level. For early stages the W3C recommends participants with a fairly high experience level, because they operate their own technology confidently and the observation is not obscured by difficulties with the assistive technology itself (W3C). That recommendation has a flip side that belongs in the analysis: someone who has used a screen reader for years knows workarounds an inexperienced person does not. The tenth WebAIM screen reader survey, with 1539 valid responses, shows the difference clearly: 78 percent of respondents with advanced proficiency turn to the headings first when looking for information on a lengthy page, compared with 47 percent of beginners (WebAIM). A round made up only of experienced users therefore measures the best possible use, not everyday use.
| Access route | What this seat makes visible | What to watch for when selecting |
|---|---|---|
| Screen reader on desktop | structure, names of controls, reading order | let participants bring their own software and settings |
| Screen reader on mobile | swipe gestures, on-screen keyboard, target sizes | own device instead of an agency test device |
| Heavy magnification | viewport changes, reflow, sticky bars | ask about zoom factor and colour inversion in advance |
| Keyboard only | focus order, visible focus, keyboard traps | invite participants without visual impairment as well |
| Voice or switch control | labels that match the visible text | plan longer sessions and more breaks |
| Reading and comprehension | wording, error messages, technical terms | read tasks aloud instead of handing them over in writing |
Three groups do not belong in a test round, however readily available they are. First, your own staff: they know how the offering is built and find routes a stranger will not. Second, colleagues wearing a blindfold. Simulation produces empathy but no findings - someone hearing a screen reader for the first time fails at the software, not at the site. Third, people who are only asked for an opinion without using the site. For areas behind a login, what we described about customer portals behind authentication applies as well: test accounts with real permissions and suitable test data must exist before the appointment, otherwise half the session is spent waiting.
Consent, recording, compensation
A test session produces personal data, and it falls into a specially protected category. A recording reveals that the person has a disability and which one; that is health data. Processing it is prohibited in principle and only becomes permissible once the data subject has given explicit consent for one or more specified purposes (Article 9(2)(a) GDPR). In practice that means the consent form names the purpose - the analysis of this test round - the scope of the recording, that is screen and audio without the face, the retention period, and the group of people allowed to view the recording. A blanket signature under "recording for research purposes" does not meet that standard, because the purpose is not specified.
Added to this are the points the W3C lists under research ethics: participants are told they are free to stop at any time, and their time is compensated appropriately (W3C). Compensation is not a tip; it is the condition for a round happening at all and not consisting solely of people who already work with accessibility professionally. It is set before the invitation goes out rather than negotiated afterwards, it is the same for every participant, and it is paid even when a session is abandoned after a few minutes. Tying compensation to a completed script creates exactly the pressure that prevents people from stopping - and stopping is a finding.
Own device
Participants work on their own technology and with their own settings: speech rate, voice, zoom factor, colour inversion. A supplied test device mostly measures how quickly someone adapts to an unfamiliar system.
Own screen reader
71.6 percent of respondents use more than one program on the desktop (WebAIM). Which one runs during the test is the participant's decision, not the script's, otherwise the log records an adjustment period instead of a barrier.
Plan for braille
A braille display is used by 38 percent of respondents (WebAIM). People who use one read differently: abbreviations, tables and long link texts behave differently on a display than in speech output.
Test mobile separately
91.3 percent of respondents also use a screen reader on a mobile device (WebAIM). The mobile session is its own appointment with its own tasks, not an appendix to the desktop session.
Limit the recording
Record screen and audio, not the face. Whatever is not needed is not captured; that shortens the consent form, the retention period and the later discussion about approvals.
Prepare remote sessions
A session from the participant's home saves travel and keeps their familiar technology in play. Screen sharing has to work alongside the assistive technology, and that is verified beforehand, not during the appointment.
Before the first real appointment comes a pilot run. It does not test the site, it tests the script: is the allotted time enough? Are the tasks understandable when read aloud? Does screen sharing work together with the screen reader? A pilot run with someone from your own organisation who does not know the offering exposes the organisational mishaps before they cost a real session. The pilot does not replace a participant and produces no findings about the site; it makes sure the first real session is not the dress rehearsal.
The most expensive mistake happens before the session
The task script: assignment instead of opinion
A usability test stands or falls with the wording of its tasks. "How do you like the navigation?" produces an opinion. "Please find out by when an order can be cancelled" produces an observation. A task describes a goal, never a route. As soon as the wording gives away the route - "open the Service section and choose Returns" - the round only checks whether participants can listen. Tasks around forms are especially vulnerable, because naming a field in the task text is already half the solution; how labels and error messages behave is covered in our article on accessible forms.
The order follows use, not site structure. A round starts with a simple task that is sure to succeed, so that participants settle in. Then come the paths that stood out in the audit, and only at the end the tasks likely to fail - an early abandonment otherwise spoils the rest of the session. For paths that have to work without a mouse, it is worth reviewing keyboard operation before the script is written: anything that is already visible at the desk as a focus-order problem needs a fix, not a test round.
- Every task names a result the person can recognise - a number, a date, a completed order - and never "have a look around"
- No term from the interface labels appears in the task text, otherwise searching turns into word matching
- One task per card, read aloud and also supplied as a text file so it can be re-read with the participant's own output
- A time limit per task is set in advance but not announced; it ends the task, not the session
- No task requires a real payment, a login with private credentials or a real cancellation
- Login and authentication are a task of their own, because that is where most abandonments happen - see accessible authentication
- At least one task deliberately routes through an error, such as a wrong format or an empty required field, because error paths are rarely evaluated
- The last task stays open: "What would you do next?" - that is where the hints no script anticipates appear
The script also contains what the moderator says when someone gets stuck. Without that part, the most common mistake in a round appears: the moderator helps. A hint like "the filter is at the top right" ends the observation and replaces it with a demonstration. So a fixed escalation is agreed - first silence, then "what are you hearing or seeing right now?", then "what would you try next?", and only after that the task is ended. Anyone who helps before the escalation is exhausted has not measured the task, they have rescued it.
Logging without interpreting
The log records what happened, not why. "Participant moves through the page with the heading key, stops at the fourth heading, goes back, opens the search" is an observation. "Participant finds the navigation confusing" is already an interpretation and is worthless in analysis, because there is no way to establish what it rests on. This separation is why a round needs two roles: the moderator speaks, the note-taker writes. One person does both badly. Which output is running belongs in the log as well - program, version, browser; the technical side of that is covered in screen reader optimisation.
Recording the technology matters because it shifts the result. 71.6 percent of respondents use more than one program on the desktop, and the two most used are close together: 65.6 percent of respondents use NVDA and 60.5 percent use JAWS (WebAIM). A finding that appears in only one program is still a finding - but it is classified differently from one every participant runs into. Without the note in the log that distinction cannot be made afterwards, and the discussion with the development team starts from zero.
A second point concerns the clock. Times are written down, but not reported as a metric. The W3C points out explicitly that accessibility testing focuses on understanding the failures rather than on task time or satisfaction, and that this usually involves a think-aloud technique with high facilitator interaction (W3C). Thinking aloud makes people slower. A duration from such a session is an indication of an obstacle, not a measurement, and certainly not a basis for comparing two participants.
Four mistakes that make a round worthless
Preference or barrier?
Not every observation is a defect. One participant who consistently navigates by headings while another uses landmarks shows two habits, not a fault. The survey data supports this: 31.7 percent of respondents always or often use landmarks when they are present, but only 3.7 percent use them as their primary way of finding information on a lengthy page (WebAIM). Deriving a rule from a single session means rebuilding the offering around one habit. The W3C warns against this in two places: feedback from one person with a disability does not apply to all people with disabilities, and results from a couple of sessions cannot be generalised (W3C).
- Does the task fail, or does it merely take longer? An abandonment is a barrier, a detour is an observation to begin with
- Does the issue appear with several participants, or only with one person using a particular setting?
- Can the issue be mapped to a success criterion? Then it is an audit finding and has to be fixed anyway
- Does the cause lie in the site, in the assistive technology, or in how it is operated? Only the first case is your finding
- Would the issue also be a problem without assistive technology? Then it affects every visitor and is a general usability defect
The fourth question is the hardest. The W3C describes accessibility as the interplay of several components - content, browser, assistive technology and knowledge of that technology - and an observed problem can sit in any of them. During the session a simple counter-check helps: the same kind of task on a comparable, working site. If it succeeds there, the cause is your offering. If it fails there too, a note goes in the log and no finding goes in the plan. That counter-check is quick to run during the session and saves a long debate later about whether an issue is even your responsibility.
| Observation in the session | Classification | Where it belongs |
|---|---|---|
| Screen reader announces "button" without a name | audit finding, success criterion affected | remediation plan, high priority |
| Error message is not announced | audit finding with immediate impact on use | remediation plan, high priority |
| Route to the goal takes several times as long | usability obstacle without a rule violation | remediation plan, medium priority |
| Participant never uses landmarks | habit, not a defect | note for the analysis |
| Output reads a date as a string of digits | behaviour of the assistive technology | note, not a finding |
| Task fails because of a technical term | comprehensibility of the text | editorial team rather than development |
Classification decides the later discussion. A finding that references a success criterion is rarely disputed. A usability obstacle without a rule violation needs the observation as evidence: task, participant profile, location, outcome. Both belong in the same report, but in separate columns. If a task fails because of jargon rather than technology, the recipient is the editorial team; how to set texts up for that is covered in plain language on the web. And if a round shows that a purchased overlay creates the obstacle, it is worth looking at overlays before another layer is added on top.
From the test log into the remediation plan
At the end there are two lists, and they must not stay two lists. The audit findings are sorted by success criteria, the test findings by task. They are merged through the location in the interface: which component, which page, which step in the path? We described the approach in our article on the remediation plan; for test findings one column is added, holding the observation the priority rests on. Without that column, a finding turns into an unsupported claim by next quarter.
The priority of a test finding follows from three inputs: how many participants failed at that point, how central the path is, and whether a reasonable detour exists. A point where everyone fails and which sits on the way to checkout goes to the top - regardless of whether a success criterion is violated. So that the round does not remain a one-off, the next appointment belongs in the same planning as the technical re-check; which events trigger a re-check is covered in the review cadence after go-live. Continuous measurement in accessibility monitoring shows in between when something has shifted.
A test log in which not a single task fails does not prove the accessibility of the offering; it proves the leniency of the tasks.
That leaves the question of effort. A round costs preparation, appointments, compensation and analysis - and it delivers findings no scan delivers. Anyone who wants to keep the effort small should test earlier: on a clickable prototype rather than the finished build, on one path rather than the whole offering. The W3C considers informal evaluations throughout development more effective than a single formal usability study at the end of a project (W3C). And anyone who wants to make the observation repeatable in their own team will find a starting point in our training.
Sources and studies
Related Articles
Screen Reader Optimization: Making Your Website Readable
Screen reader optimisation for websites: ARIA landmarks, roles, live regions, semantic HTML and screen reader testing on Windows, macOS and mobile devices.
Accessible Job Applications: Career Pages Without Barriers
Application flows fall outside the BFSG but inside the AGG and SGB IX. Where uploads, required fields and session timeouts cost you applications.
Accessible Customer Portals: Auditing the Login Area
Automated scans stop at the sign-in form. How a WCAG audit covers the protected area: scope, test accounts, session timeouts and screen reader testing.