Skip to content
Practice & implementation

Usability Testing With Disabled Users: What Tools Miss

How a test round with blind, low-vision and motor-impaired participants is organised: recruitment, consent, task script, logging and the route into the remediation plan.

14 min read NutzertestsScreenreaderUsability-TestRekrutierungBarrierefreiheit

A WCAG audit report answers the question of whether a button has an accessible name. It does not answer the question of whether anyone finds it. Between those two sentences sit findings that can become expensive later: the search that is marked up correctly and is still abandoned; the checkout that meets every success criterion and breaks at a point nobody considered during design. Places like that only become visible when someone uses the site who has to use it - with a screen reader, with heavy magnification, with the keyboard alone. A usability test with disabled participants does not replace the WCAG audit; it answers the question the audit leaves open. This article describes how such a round is organised: recruitment, consent, task script, logging, interpretation - and the four mistakes that make a round worthless.

Key takeaways

  • Automated tools cover part of the criteria, everything else requires human judgement (W3C). A usability test closes exactly that gap instead of repeating the automated scan.
  • Evaluating with disabled users finds usability issues that conformance evaluation alone does not find (W3C) - and the reverse holds too. Only both routes together produce a complete picture.
  • Small rounds deliver solid observations, but no statistical significance (W3C). The result is a list of failures with evidence, not a score or a pass rate.
  • Participant technology is part of the setup: 71.6 percent of respondents use more than one desktop screen reader (WebAIM). Prescribing the choice measures something other than real use.
  • Recording a session touches health data. It is only permissible with explicit consent for specified purposes (Article 9(2)(a) GDPR), and participants may stop at any time (W3C).
  • Test findings belong in the same remediation plan as audit findings. With two separate lists, one of them is easily left untouched.

What an automated check cannot see

The limit is not controversial; it is written into the W3C guidance on selecting evaluation tools: tools cannot check all accessibility aspects automatically, and human judgement is required for the rest (W3C). That is not an objection to the tools. A scan finds in minutes what takes a person hours - missing alternative text, insufficient contrast, unlabelled form fields, duplicate identifiers. What it cannot do is judge: whether an existing alternative text matches the purpose of the image, whether the heading order follows a comprehensible structure, whether an error message is understood, and whether a path through a checkout with a screen reader ends within a reasonable time. Those are exactly the questions that decide whether a site is usable.

The second limit lies in the question being asked. A conformance evaluation asks about rules; a usability test asks about tasks. The W3C describes the difference in its guidance on involving users: evaluating with users with disabilities and with older users identifies usability issues that conformance evaluation alone does not discover (W3C). The reverse applies as well: a test round will not find a missing language attribute in the source or a contrast ratio just below the threshold. Treating the two routes as alternatives means choosing between two incomplete pictures. Which method delivers what is covered in our overview of methods and evaluation approaches; this article describes the part that cannot be automated.

Three questions no scan answers

Does a blind user find the filter without knowing the site beforehand? Does someone using heavy magnification abandon the login because the error message appears outside the visible viewport? Does a participant using only the keyboard need several times the planned duration to book an appointment? All three questions have a clear answer. It simply does not appear in any audit report, because no tool asks them.

Who to invite - and who not to

At the end of 2023, around 7.9 million people with a severe disability lived in Germany (Federal Statistical Office). For a test round that figure is only background; what matters is the composition. A round in which every participant is blind and uses the same software describes one mode of use, not use in general. So the seats are filled along access routes: screen reader on the desktop, screen reader on mobile, heavy magnification, keyboard-only operation, voice or switch control, plus participants with reading or concentration difficulties. We assemble rounds like this as part of a moderated screen reader test. Invitations are best sent through associations, self-advocacy organisations and university groups rather than through a notice on your own website.

The second dimension is experience level. For early stages the W3C recommends participants with a fairly high experience level, because they operate their own technology confidently and the observation is not obscured by difficulties with the assistive technology itself (W3C). That recommendation has a flip side that belongs in the analysis: someone who has used a screen reader for years knows workarounds an inexperienced person does not. The tenth WebAIM screen reader survey, with 1539 valid responses, shows the difference clearly: 78 percent of respondents with advanced proficiency turn to the headings first when looking for information on a lengthy page, compared with 47 percent of beginners (WebAIM). A round made up only of experienced users therefore measures the best possible use, not everyday use.

Access routeWhat this seat makes visibleWhat to watch for when selecting
Screen reader on desktopstructure, names of controls, reading orderlet participants bring their own software and settings
Screen reader on mobileswipe gestures, on-screen keyboard, target sizesown device instead of an agency test device
Heavy magnificationviewport changes, reflow, sticky barsask about zoom factor and colour inversion in advance
Keyboard onlyfocus order, visible focus, keyboard trapsinvite participants without visual impairment as well
Voice or switch controllabels that match the visible textplan longer sessions and more breaks
Reading and comprehensionwording, error messages, technical termsread tasks aloud instead of handing them over in writing

Three groups do not belong in a test round, however readily available they are. First, your own staff: they know how the offering is built and find routes a stranger will not. Second, colleagues wearing a blindfold. Simulation produces empathy but no findings - someone hearing a screen reader for the first time fails at the software, not at the site. Third, people who are only asked for an opinion without using the site. For areas behind a login, what we described about customer portals behind authentication applies as well: test accounts with real permissions and suitable test data must exist before the appointment, otherwise half the session is spent waiting.

A test session produces personal data, and it falls into a specially protected category. A recording reveals that the person has a disability and which one; that is health data. Processing it is prohibited in principle and only becomes permissible once the data subject has given explicit consent for one or more specified purposes (Article 9(2)(a) GDPR). In practice that means the consent form names the purpose - the analysis of this test round - the scope of the recording, that is screen and audio without the face, the retention period, and the group of people allowed to view the recording. A blanket signature under "recording for research purposes" does not meet that standard, because the purpose is not specified.

Added to this are the points the W3C lists under research ethics: participants are told they are free to stop at any time, and their time is compensated appropriately (W3C). Compensation is not a tip; it is the condition for a round happening at all and not consisting solely of people who already work with accessibility professionally. It is set before the invitation goes out rather than negotiated afterwards, it is the same for every participant, and it is paid even when a session is abandoned after a few minutes. Tying compensation to a completed script creates exactly the pressure that prevents people from stopping - and stopping is a finding.

Own device

Participants work on their own technology and with their own settings: speech rate, voice, zoom factor, colour inversion. A supplied test device mostly measures how quickly someone adapts to an unfamiliar system.

Own screen reader

71.6 percent of respondents use more than one program on the desktop (WebAIM). Which one runs during the test is the participant's decision, not the script's, otherwise the log records an adjustment period instead of a barrier.

Plan for braille

A braille display is used by 38 percent of respondents (WebAIM). People who use one read differently: abbreviations, tables and long link texts behave differently on a display than in speech output.

Test mobile separately

91.3 percent of respondents also use a screen reader on a mobile device (WebAIM). The mobile session is its own appointment with its own tasks, not an appendix to the desktop session.

Limit the recording

Record screen and audio, not the face. Whatever is not needed is not captured; that shortens the consent form, the retention period and the later discussion about approvals.

Prepare remote sessions

A session from the participant's home saves travel and keeps their familiar technology in play. Screen sharing has to work alongside the assistive technology, and that is verified beforehand, not during the appointment.

Before the first real appointment comes a pilot run. It does not test the site, it tests the script: is the allotted time enough? Are the tasks understandable when read aloud? Does screen sharing work together with the screen reader? A pilot run with someone from your own organisation who does not know the offering exposes the organisational mishaps before they cost a real session. The pilot does not replace a participant and produces no findings about the site; it makes sure the first real session is not the dress rehearsal.

The most expensive mistake happens before the session

A round in which the technology is set up during the appointment loses session time and the participant's attention with it. Technology checks, test accounts and test data belong in a separate preparation step - even when everyone involved promises it will be quick.

The task script: assignment instead of opinion

A usability test stands or falls with the wording of its tasks. "How do you like the navigation?" produces an opinion. "Please find out by when an order can be cancelled" produces an observation. A task describes a goal, never a route. As soon as the wording gives away the route - "open the Service section and choose Returns" - the round only checks whether participants can listen. Tasks around forms are especially vulnerable, because naming a field in the task text is already half the solution; how labels and error messages behave is covered in our article on accessible forms.

The order follows use, not site structure. A round starts with a simple task that is sure to succeed, so that participants settle in. Then come the paths that stood out in the audit, and only at the end the tasks likely to fail - an early abandonment otherwise spoils the rest of the session. For paths that have to work without a mouse, it is worth reviewing keyboard operation before the script is written: anything that is already visible at the desk as a focus-order problem needs a fix, not a test round.

  • Every task names a result the person can recognise - a number, a date, a completed order - and never "have a look around"
  • No term from the interface labels appears in the task text, otherwise searching turns into word matching
  • One task per card, read aloud and also supplied as a text file so it can be re-read with the participant's own output
  • A time limit per task is set in advance but not announced; it ends the task, not the session
  • No task requires a real payment, a login with private credentials or a real cancellation
  • Login and authentication are a task of their own, because that is where most abandonments happen - see accessible authentication
  • At least one task deliberately routes through an error, such as a wrong format or an empty required field, because error paths are rarely evaluated
  • The last task stays open: "What would you do next?" - that is where the hints no script anticipates appear

The script also contains what the moderator says when someone gets stuck. Without that part, the most common mistake in a round appears: the moderator helps. A hint like "the filter is at the top right" ends the observation and replaces it with a demonstration. So a fixed escalation is agreed - first silence, then "what are you hearing or seeing right now?", then "what would you try next?", and only after that the task is ended. Anyone who helps before the escalation is exhausted has not measured the task, they have rescued it.

Logging without interpreting

The log records what happened, not why. "Participant moves through the page with the heading key, stops at the fourth heading, goes back, opens the search" is an observation. "Participant finds the navigation confusing" is already an interpretation and is worthless in analysis, because there is no way to establish what it rests on. This separation is why a round needs two roles: the moderator speaks, the note-taker writes. One person does both badly. Which output is running belongs in the log as well - program, version, browser; the technical side of that is covered in screen reader optimisation.

Recording the technology matters because it shifts the result. 71.6 percent of respondents use more than one program on the desktop, and the two most used are close together: 65.6 percent of respondents use NVDA and 60.5 percent use JAWS (WebAIM). A finding that appears in only one program is still a finding - but it is classified differently from one every participant runs into. Without the note in the log that distinction cannot be made afterwards, and the discussion with the development team starts from zero.

A second point concerns the clock. Times are written down, but not reported as a metric. The W3C points out explicitly that accessibility testing focuses on understanding the failures rather than on task time or satisfaction, and that this usually involves a think-aloud technique with high facilitator interaction (W3C). Thinking aloud makes people slower. A duration from such a session is an indication of an obstacle, not a measurement, and certainly not a basis for comparing two participants.

Four mistakes that make a round worthless

The moderator helps; participants work on unfamiliar technology; the task gives away the route; the log contains interpretations instead of observations. Each of these arises from the wish to let the session run smoothly - and each erases precisely the finding the round exists to produce.

Preference or barrier?

Not every observation is a defect. One participant who consistently navigates by headings while another uses landmarks shows two habits, not a fault. The survey data supports this: 31.7 percent of respondents always or often use landmarks when they are present, but only 3.7 percent use them as their primary way of finding information on a lengthy page (WebAIM). Deriving a rule from a single session means rebuilding the offering around one habit. The W3C warns against this in two places: feedback from one person with a disability does not apply to all people with disabilities, and results from a couple of sessions cannot be generalised (W3C).

  1. Does the task fail, or does it merely take longer? An abandonment is a barrier, a detour is an observation to begin with
  2. Does the issue appear with several participants, or only with one person using a particular setting?
  3. Can the issue be mapped to a success criterion? Then it is an audit finding and has to be fixed anyway
  4. Does the cause lie in the site, in the assistive technology, or in how it is operated? Only the first case is your finding
  5. Would the issue also be a problem without assistive technology? Then it affects every visitor and is a general usability defect

The fourth question is the hardest. The W3C describes accessibility as the interplay of several components - content, browser, assistive technology and knowledge of that technology - and an observed problem can sit in any of them. During the session a simple counter-check helps: the same kind of task on a comparable, working site. If it succeeds there, the cause is your offering. If it fails there too, a note goes in the log and no finding goes in the plan. That counter-check is quick to run during the session and saves a long debate later about whether an issue is even your responsibility.

Observation in the sessionClassificationWhere it belongs
Screen reader announces "button" without a nameaudit finding, success criterion affectedremediation plan, high priority
Error message is not announcedaudit finding with immediate impact on useremediation plan, high priority
Route to the goal takes several times as longusability obstacle without a rule violationremediation plan, medium priority
Participant never uses landmarkshabit, not a defectnote for the analysis
Output reads a date as a string of digitsbehaviour of the assistive technologynote, not a finding
Task fails because of a technical termcomprehensibility of the texteditorial team rather than development

Classification decides the later discussion. A finding that references a success criterion is rarely disputed. A usability obstacle without a rule violation needs the observation as evidence: task, participant profile, location, outcome. Both belong in the same report, but in separate columns. If a task fails because of jargon rather than technology, the recipient is the editorial team; how to set texts up for that is covered in plain language on the web. And if a round shows that a purchased overlay creates the obstacle, it is worth looking at overlays before another layer is added on top.

From the test log into the remediation plan

At the end there are two lists, and they must not stay two lists. The audit findings are sorted by success criteria, the test findings by task. They are merged through the location in the interface: which component, which page, which step in the path? We described the approach in our article on the remediation plan; for test findings one column is added, holding the observation the priority rests on. Without that column, a finding turns into an unsupported claim by next quarter.

The priority of a test finding follows from three inputs: how many participants failed at that point, how central the path is, and whether a reasonable detour exists. A point where everyone fails and which sits on the way to checkout goes to the top - regardless of whether a success criterion is violated. So that the round does not remain a one-off, the next appointment belongs in the same planning as the technical re-check; which events trigger a re-check is covered in the review cadence after go-live. Continuous measurement in accessibility monitoring shows in between when something has shifted.

A test log in which not a single task fails does not prove the accessibility of the offering; it proves the leniency of the tasks.

Guiding principle for test moderation

That leaves the question of effort. A round costs preparation, appointments, compensation and analysis - and it delivers findings no scan delivers. Anyone who wants to keep the effort small should test earlier: on a clickable prototype rather than the finished build, on one path rather than the whole offering. The W3C considers informal evaluations throughout development more effective than a single formal usability study at the end of a project (W3C). And anyone who wants to make the observation repeatable in their own team will find a starting point in our training.

Sources and studies

This article draws on the tenth WebAIM screen reader survey, on the W3C guidance for selecting evaluation tools and for involving users, on a press release from the German Federal Statistical Office and on Article 9 of the General Data Protection Regulation. The figures cited refer to the state of the respective publication.

Related Articles

Practice & implementation

Screen Reader Optimization: Making Your Website Readable

Screen reader optimisation for websites: ARIA landmarks, roles, live regions, semantic HTML and screen reader testing on Windows, macOS and mobile devices.

14 min read
Branchen & Anwendungsfälle

Accessible Job Applications: Career Pages Without Barriers

Application flows fall outside the BFSG but inside the AGG and SGB IX. Where uploads, required fields and session timeouts cost you applications.

13 min read
Practice & implementation

Accessible Customer Portals: Auditing the Login Area

Automated scans stop at the sign-in form. How a WCAG audit covers the protected area: scope, test accounts, session timeouts and screen reader testing.

13 min read