Skip to content
UI & UX Design

Usability Testing: What Actually Works

A phase-by-phase usability testing playbook: framing, recruiting, realistic tasks, neutral sessions, severity ratings, and small iterative rounds that work.

Priya Raghunathan Head of Design 26 min read 25 views
Usability Testing: What Actually Works

Usability testing is the practice of watching representative people attempt real tasks with a design and noting exactly where they hesitate, misread, backtrack or give up. It is not a survey, a focus group or a demo followed by a round of opinions. The value comes from observation: you see behavior that nobody on the team can predict, because everyone on the team already knows how the product is supposed to work.

It matters to anyone who is about to spend money on the assumption that people will understand an interface. That includes founders shipping a first checkout, marketers relaunching a lead form, product managers redesigning an onboarding flow and in-house teams that have been staring at the same screens for two years. Done well, usability testing is the cheapest way to find serious problems before they reach customers, and it consistently surfaces issues that no internal review, heuristic audit or stakeholder walkthrough catches.

This playbook walks through the method as a sequence of phases you follow in order. Each phase has a goal, the actions to take, what you should have in hand at the end, and the checks that tell you it is safe to move on. It closes with an illustrative worked example, a realistic schedule, and guidance on when to bring in outside help.

What usability testing measures, and the full procedure at a glance

At its core, a usability test answers one question: can the intended people complete the tasks that matter, and if not, what specifically stops them? The standard format is simple. A participant receives a realistic task, attempts it with the design while thinking aloud, and a facilitator observes without helping. Notes capture where the participant struggled, what they expected to happen, and what they did instead. Those observations become a list of problems ranked by severity, which the team fixes before testing again.

What the method does not measure is just as important. A handful of sessions will not tell you what percentage of your traffic will convert, which of two headlines performs better, or whether people like your brand colors. Those are questions for analytics, experiments such as A/B testing for design decisions, or larger quantitative studies. Usability testing is a diagnostic tool. It tells you why something fails and what to change, and it does that with far fewer people than a statistical study needs.

The Nielsen Norman Group has long argued that small samples find most of the issues in a given round, and that the smarter investment is several small rounds rather than one large one. The reasoning is practical: the first few participants expose the biggest problems, later participants mostly repeat them, and once you fix those problems a new round reveals the next layer that was hidden underneath. That principle shapes the whole procedure below.

  1. Frame the decision Agree which design question the test must answer and which flows are in scope, so the sessions produce something the team can act on.
  2. Recruit the right people Screen for participants who match the real audience, including people who have never used the product.
  3. Write realistic tasks Turn each goal into a scenario with a clear end state, avoiding interface vocabulary that gives away the answer.
  4. Choose the setup Decide between in-person, remote moderated and remote unmoderated sessions based on what you need to observe.
  5. Run neutral sessions Facilitate without leading, prompt participants to narrate, and let them struggle long enough to see why.
  6. Analyze and rate severity Consolidate observations into distinct problems, rate each one, and tie it to evidence.
  7. Fix, then test again Change the design before the next round and check that the fixes worked without creating new problems.

Phase 1: Frame the decision the usability test must inform

Goal: make sure the test answers a question someone will act on. The most common reason usability testing fails to change anything is not bad facilitation. It is that nobody agreed beforehand what the results were for, so the findings arrive as an interesting list with no owner.

Actions

Start with a short conversation with whoever owns the design and whoever owns the business outcome. Ask three things: which flow or screen are you least confident in, what would you change if the test showed a problem, and when is the latest point a change can still be made cheaply? The answers define scope. A good scope is narrow: "Can first-time visitors find and book the right service tier?" is testable in a 45-minute session. "Is the new website usable?" is not.

Next, list the tasks that matter most to the business and to users. For an ecommerce site that might be finding a product by need rather than by name, comparing two options, and completing checkout as a guest. For a SaaS dashboard it might be setting up the first project, inviting a colleague and exporting a report. Rank these by risk: how likely is failure, and how costly would it be? Test the top three to five.

Finally, decide what you will test with. Teams routinely wait for polish, and that is a mistake. A clickable wireframe or a rough prototype tested in week two is worth more than a pixel-perfect build tested the week before launch, because problems found early are cheap to fix. The trade-offs between paper sketches, grayscale prototypes and near-final builds are covered in depth in our guide to prototype fidelity for testing, but the short version is that fidelity should match the question. Navigation and comprehension can be tested at low fidelity; visual hierarchy, microcopy and form behavior need something closer to real.

Outputs and checks

The output of this phase is a one-page test plan: the research questions, the flows in scope, the prototype or build to be used, the target participant profile, the number of rounds and the date by which findings are needed. The check is simple. If you cannot name a person who will decide what to change based on the findings, stop and fix that before recruiting anyone.

  • Written research questions, each tied to a specific flow or screen
  • Three to five priority tasks ranked by likelihood and cost of failure
  • A named decision owner who has agreed to review findings
  • A prototype or build that supports every task end to end
  • A date for findings that still leaves time to change the design
  • Agreement on how many rounds the project budget and schedule allow
What to do and what to avoid with usability testing, side by side
Good practice against the usual mistakes, from the sources listed below.

Phase 2: Recruit participants who match the real audience

Goal: put the design in front of people whose knowledge, motivation and context resemble the people who will actually use it. The wrong participants produce confident, misleading results.

Who to recruit

Write a screener based on behaviors and circumstances, not demographics alone. "Has booked a home service online in the past six months" predicts relevant behavior far better than an age bracket. Include questions that screen out people who work in design, marketing or user research, since they tend to critique rather than use. For B2B products, screen for role and responsibility: an office manager who buys supplies behaves differently from a procurement lead who approves contracts.

One of the most damaging recruiting habits is testing only with existing customers. They are easy to reach and happy to help, but they have already learned the product's quirks and vocabulary. They will sail through flows that new users find baffling. Unless the research question is specifically about existing users, recruit a mix, and make sure newcomers make up a meaningful share of each round.

How many people per round

For qualitative usability testing aimed at finding problems, five to eight participants per distinct user group per round is a common working range. If you have two audiences that use the product very differently, such as patients and clinic staff, treat them as separate groups and recruit for each. The logic, and the cases where you need more, such as benchmarking or statistical comparison, are explained in our article on usability test sample sizes. What matters most is that you plan several rounds rather than spending the whole budget on one.

Logistics that save sessions

Recruit one or two more people than you need, because no-shows are normal. Confirm the day before and again an hour before. For remote sessions, send a short technical check in advance: the link, the browser to use and whether screen sharing is needed. Offer an incentive that respects the participant's time; the right amount depends on the audience, with specialist professionals typically needing considerably more than general consumers. Collect consent for recording in plain language, and explain how recordings will be stored and when they will be deleted.

Shortcut: keep a lightweight panel of people who have opted in to future research, tagged by the screener answers that matter to you. Rebuilding a recruit from scratch every round is where most small teams lose a week.

  • Record when each person last took part so you can avoid over-testing the same people.
  • Add a new-user quota to every round so the panel never drifts toward experts.

Phase 3: Write realistic tasks, not tours of the interface

Goal: give participants something to accomplish, so you observe behavior instead of collecting opinions. The task script is the single biggest lever on the quality of your findings.

What a good task looks like

A good task describes a situation and a goal in the participant's language, with a clear end point. Compare these two versions:

  • Weak: "Use the Filters panel to find a size 10 running shoe under the Sale category."
  • Strong: "You run three times a week and your current shoes are worn out. Your budget is fairly tight. Find a pair you would consider buying and get to the point where you would pay."

The weak version names the feature, the category and the path, so it tests reading comprehension, not usability. The strong version leaves the participant to decide how to search, which is exactly what real visitors must do. It also has an observable finish line, which lets you judge success consistently across sessions.

Rules for task wording

  • Avoid words that appear in the interface's labels. If the button says "Schedule," do not ask people to "schedule" something; ask them to "arrange a visit."
  • Give participants the information a real person would have: an account to log into, a delivery address, a problem to solve.
  • Keep one goal per task. Compound tasks make it hard to tell which part caused a failure.
  • Order tasks the way a real journey would unfold, but be ready to reset state if a failure would block the next task.
  • Pilot the script with a colleague outside the project to catch ambiguous wording before real participants see it.

Deciding what counts as success

For each task, write down the success criteria before the first session: what end state counts as complete, and what counts as complete with difficulty. This prevents the drift where observers, having watched someone struggle for four minutes, generously call it a success because the person eventually got there. Many teams track three outcomes per task: completed unaided, completed with difficulty or after a wrong path, and failed or abandoned. If you also want a standardized satisfaction measure at the end of each session, the System Usability Scale is a well-established ten-item questionnaire, though with small samples it should be read as a directional signal rather than a benchmark.

If the research question is specifically about navigation or labeling, consider running a quick structural method first. Tree testing isolates whether people can find things in your information architecture without the visual design getting in the way, and it can sharpen the tasks you later give participants in a full usability session.

  • Every task is a scenario with a goal, not an instruction to use a feature
  • No task wording repeats a label from the interface
  • Each task has written success, partial and failure criteria
  • Test data, logins and example details are prepared for each participant
  • The script has been piloted once with someone outside the project
  • Session length fits the tasks, with time left for a short debrief

Phase 4: Choose between in-person, remote moderated and unmoderated setups

Goal: pick the setup that lets you observe what you need to observe at a cost and pace the project can sustain. There is no universally best option; each makes different trade-offs.

SetupBest forMain strengthsMain limitations
In-person moderatedPhysical products, kiosks, specialist users, complex or sensitive flowsRichest observation of body language and context; easiest to handle surprises and probe in depthTravel and room logistics; smaller geographic reach; slower to schedule
Remote moderatedMost web and app work, distributed audiences, B2B rolesReal-time follow-up questions; participants use their own devices and setting; wide reachDependent on connection and screen sharing; harder to see physical cues
Remote unmoderatedShort, well-defined tasks; quick checks between rounds; larger samplesFast turnaround; runs in parallel; lower cost per sessionNo chance to probe; tasks must be airtight; more low-quality sessions to discard

Remote moderated sessions are the default for most web projects because they combine real-time probing with broad reach. Unmoderated tools are excellent for speed, but they punish vague tasks: with no facilitator to clarify, a confusing instruction produces a session that tells you nothing. Our guide to remote unmoderated testing covers how to write tasks and screen sessions for that format.

Device matters as much as location. If most of your traffic arrives on phones, test on phones. Mobile sessions introduce their own practical problems, such as recording touch gestures, managing notifications and seeing the participant's thumb reach, and these are addressed step by step in usability testing on mobile devices. Desktop findings do not transfer cleanly to small screens, so avoid the temptation to test only on the device that is easiest to record.

Roles in a moderated session

Assign a facilitator who talks to the participant and one or two note-takers who say nothing. Observers from the wider team can watch live through a separate link with their cameras and microphones off. Having stakeholders watch even two sessions is one of the most effective ways to build agreement on what needs fixing, because it replaces debate about opinions with shared memory of what actually happened.

Phase 5: Run the sessions without leading the participant

Goal: observe genuine behavior. Leading questions and helpful hints invalidate a result faster than any other mistake, because they turn a failure into a success that will not happen in the real world.

A session structure that works

A typical moderated session of 45 to 60 minutes breaks down roughly as follows: five minutes of introduction and consent, five minutes of background questions about the participant's context, 30 to 40 minutes of tasks, and five to ten minutes of debrief. Tell participants up front that you are testing the design, not them, that there are no wrong answers, and that you did not build it, so they can be honest. Ask them to think aloud: to say what they are looking for, what they expect to happen and what surprises them.

What the facilitator says, and does not say

When a participant struggles, say nothing except to prompt them to narrate. Silence is uncomfortable, and the instinct to help is strong, but the struggle is the data. Useful neutral prompts include:

  • "What are you looking for right now?"
  • "What do you expect will happen if you click that?"
  • "You paused there. What were you thinking?"
  • "Is this what you expected to see?"

Prompts to avoid are the ones that contain the answer or signal approval: "Did you see the menu at the top?", "Would you click Continue now?" or "Great, that's right." Also avoid asking participants to predict other people's behavior or to redesign the interface. Their opinions about design are much less reliable than their actions.

Set a time limit for each task in advance. If a participant is clearly stuck and further time will not produce new information, thank them, mark the task as failed and move on. If a later task depends on the failed one, reset the prototype to the right state rather than walking them through the solution.

Note-taking that speeds up analysis

Ask note-takers to record observations, not interpretations, with a timestamp so the moment can be found on the recording. "Scrolled past the pricing table twice, then asked where the prices were" is useful. "Pricing is confusing" is a conclusion that belongs in analysis. A shared spreadsheet with columns for participant, task, timestamp, observation and a direct quote makes consolidation much faster later.

Shortcut: hold a 15-minute debrief with observers immediately after each session, while memory is fresh. Each person names the two or three most important things they saw. By the last session of the round, the major themes are usually already clear, and formal analysis becomes confirmation rather than discovery.

Phase 6: Analyze findings and rate each problem by severity

Goal: turn many observations into a short, prioritized list of distinct problems, each backed by evidence, that the decision owner can act on.

From observations to problems

Group related observations across participants. Five different notes about people missing the delivery options may all describe one problem: the options sit below a visual break that reads like the end of the page. Write each problem as a statement of what happened and why it matters, not as a solution. "Four of six participants did not find delivery options before trying to pay, because the section sits below a full-width banner" is a finding. "Move delivery options up" is a recommendation, and it should be listed separately so the team can consider alternatives.

With small samples, report counts plainly ("four of six") rather than converting them to percentages, which imply a precision the study does not have. Be clear that one participant hitting a serious problem is still a finding worth fixing if the consequences are severe, for example accidentally submitting a payment twice.

Rating severity

Rate each problem on a simple scale so the team can prioritize. A practical four-level scale looks like this:

SeverityDefinitionTypical response
CriticalPrevents task completion or causes a costly error, such as a lost order or a wrong bookingFix before the next round or before release
SeriousCauses significant delay, confusion or a wrong path, though most people recoverFix in the current iteration
MinorCauses brief hesitation or mild annoyance without affecting outcomeSchedule in the backlog; fix when touching that area
ObservationWorth noting, such as a positive reaction or a suggestion, but not a problemRecord for context; no immediate action

Severity combines three factors: how many participants were affected, how badly it affected them, and whether it would recur for the same person every time. Established frameworks such as Nielsen's usability heuristics are useful for explaining why a problem occurs, for example a lack of system status feedback or a mismatch between the interface and real-world language, which helps designers find the right fix rather than patching the symptom.

Reporting so people act

Keep the report short. A one-page summary of the top problems, a table of all findings with severity, and a handful of short video clips will be read and used; a 60-page deck will not. Lead with what worked, briefly, then the critical and serious problems, each with evidence and a recommended direction. The discipline of turning research into choices that stick, including how to run the review meeting and assign owners, is covered in our article on turning research into decisions.

Public research bodies offer good models for this. The GOV.UK Service Manual treats research as a continuous activity built into every phase of service delivery rather than a one-off gate, and that framing helps teams see each report as input to the next iteration rather than a final verdict.

  • Every observation is grouped into a distinct problem statement
  • Each problem has a severity rating and a count of affected participants
  • Findings and recommendations are listed separately
  • Critical problems include a short video clip or direct quote as evidence
  • The decision owner has reviewed the findings and assigned fixes
  • Things that worked well are recorded so they are not broken by later changes

Phase 7: Fix the problems, then run the next round

Goal: convert findings into design changes and verify that those changes work. A round of testing whose findings are not fixed before the next round is mostly wasted, because the next group of participants will simply trip over the same problems and hide everything behind them.

Why several small rounds beat one large one

Imagine a checkout where the first critical problem is that people cannot find the guest checkout option. Until that is fixed, participants never reach the payment step, so any problems there remain invisible. Test with 15 people at once and you will see the guest checkout problem 15 times and learn almost nothing about payment. Test with five, fix the entry point, then test with five more, and the second round reaches the payment step and exposes the next layer. Spending the same budget on three rounds of five produces far more distinct findings than a single round of 15, and it also checks whether each fix actually worked.

That is why one large test at the end of a project is among the costliest mistakes in the practice. By the time results arrive, the design is built, the launch date is fixed and only cosmetic changes are affordable. Moving testing earlier and repeating it is almost always a better use of the same money.

What to change between rounds

  • Fix critical and serious problems first. Resist the urge to redesign everything at once, which makes it harder to know which change solved what.
  • Keep the core tasks the same so you can compare performance across rounds, but add new tasks as the design grows.
  • Re-test anything you changed. A fix can introduce a new problem, for example a more visible guest checkout button that now draws people away from saved-account login.
  • Update the screener if a round revealed that you recruited the wrong mix of people.

Once a flow is stable and live, the questions shift. You may want quantitative evidence of how it performs over time or against a competitor, which is the territory of UX benchmarking, or a precise test of whether people's first click goes to the right place, which is where first-click testing earns its keep. Large-scale research organizations such as the Baymard Institute publish usability research on common patterns like checkout and product listings, which is useful for spotting likely issues before you test, though it never replaces watching your own users on your own design.

A common objection is that a small usability test is not statistically valid, so its findings can be ignored. In practice, qualitative usability testing is not trying to estimate a population rate. It is trying to find problems and understand their causes. If several people in a small round cannot find the checkout button, you do not need a larger sample to know that the button needs work. Statistical rigor matters when you are measuring and comparing, and those are different studies with different sample sizes.

A worked example: three rounds on an appointment booking flow

The following example is illustrative, not a client story. It shows how the phases fit together on a realistic small project, with plausible numbers for a team of one researcher, one designer and a product owner.

A regional physiotherapy clinic group is redesigning its online booking. The research question is whether first-time patients can choose the right appointment type, pick a location and time, and complete booking without phoning the clinic. The team plans three rounds of five remote moderated sessions over five weeks, each session 45 minutes, with a mix of first-time patients and returning patients at roughly three to two. Considerations particular to this kind of flow, from appointment type naming to insurance details, are explored in our piece on healthcare appointment booking UX.

Round 1: clickable wireframes

Four of five participants hesitate at the appointment type step, because options are named with internal terms such as "IA" and "FU." Two of the four pick the wrong type. Three of five do not notice that the location selector filters available times, and assume the clinic has no availability. One participant completes the flow unaided. Critical problems: appointment naming and the hidden location filter. The designer renames types to plain-language descriptions with a short line on who each is for, and moves location selection before the calendar.

Round 2: revised mid-fidelity prototype

All five participants now choose a type without hesitation, and four of five choose correctly. With the earlier blockers gone, a new serious problem appears: three participants reach the details form and are unsure whether insurance information is required, and two abandon to "check with the clinic first." The team adds a clear optional label and an explanation of when insurance details are needed. Four of five complete the flow, compared with one of five in round one.

Round 3: near-final build on phones

Because most bookings are expected on mobile, the third round runs on participants' own phones. Five of five complete booking. Two participants struggle to tap the time slots, which are small and tightly spaced, and one misses the confirmation because it appears below the fold. Both are fixed before release: larger tap targets and a full-screen confirmation state. Remaining minor issues go into the backlog.

Across 15 sessions, the team found and fixed two critical problems and three serious ones, most of which could not have been seen until an earlier problem was removed. A single round of 15 at the end would have spent the entire budget watching people fail at the appointment type step.

  1. Week 1 Agree research questions and scope, write the screener, start recruiting, draft and pilot the task script.
  2. Week 2 Run round 1 with five participants on wireframes; hold daily debriefs; deliver a short findings summary by the end of the week.
  3. Week 3 Designer fixes critical and serious problems; researcher recruits round 2 and updates tasks; run round 2 late in the week.
  4. Week 4 Consolidate round 2 findings, fix the new serious problem, prepare the near-final build and mobile setup, recruit round 3.
  5. Week 5 Run round 3 on phones, fix remaining critical items, log minor issues and hand a final findings table to the product owner.

When to bring in specialist usability testing help

Many teams can run competent usability tests themselves once they have a template, a recruiting channel and a facilitator willing to stay quiet. There are three situations where outside help usually pays for itself.

A flow is underperforming and nobody knows why

When analytics show a sharp drop-off at one step, but internal reviews keep producing competing theories, an independent researcher can design a focused study, facilitate neutrally and report findings without the politics of an internal team defending its own work. Independence matters most when the problem touches decisions that senior people have already made.

Testing must be moderated with specialist users

Clinicians, engineers, financial advisers and other specialist audiences are hard to recruit and expensive to waste. Sessions with them need a facilitator who understands the domain well enough to write credible scenarios and follow the conversation, but who can still resist explaining the interface. An experienced external team typically brings both the recruiting relationships and the facilitation discipline.

Accessibility testing with assistive technology

Testing with people who use screen readers, switch devices, magnification or voice control is a distinct skill. Sessions need different setups, longer timings and facilitators who understand how each technology behaves. Automated checkers catch only part of the picture. Our guide to accessibility testing with assistive technology explains how to plan these sessions properly.

If any of these apply, our usability testing service can plan, recruit, facilitate and report on rounds as a standalone engagement or as part of a wider design project.

The value of usability testing does not come from the size of any one study. It comes from watching real people attempt real tasks, fixing what they struggle with, and doing it again before the design is too expensive to change.

Common failure points in usability testing and how to avoid them

Most failed usability programs do not fail because of a single dramatic error. They fail through a handful of habits that quietly drain the value from every session. Knowing them in advance is the easiest way to protect your investment.

Demonstrating the interface and asking for opinions
Walking participants through the design and asking what they think produces polite approval and design suggestions, not evidence. Give them a task and watch instead.
Leading participants toward the intended path
Hints, approving noises and questions that name interface elements all contaminate results. Script your prompts and review a recording of your own facilitation after the first session of each round.
Testing too late
Waiting for a finished build guarantees that only cosmetic findings can be acted on. Test rough prototypes early, when a structural change costs a few hours of design time.
Testing only with current users
Experienced customers have learned the workarounds. Newcomers reveal the problems that stop people from ever becoming customers.
Reporting without owners
Findings that are not assigned to someone with authority to change the design rarely get fixed. Name the decision owner in the test plan and hold a review meeting within a few days of each round.
Treating one round as the finish line
Each round hides the next layer of problems behind the current blockers. Budget for iteration from the start, even if each round is small.

If you run usability testing this way, as a sequence of short, well-scoped rounds with realistic tasks, neutral facilitation and fixes applied before each new round, it becomes a routine part of design rather than a special event. For more on the wider discipline, browse our UI and UX design articles.

Where this comes from

The figures and practices above come from the sources listed.

Working on something like this?

We take on UI & UX Design work for teams who want it done once, properly. Tell us what you are building and we will tell you honestly whether we are the right studio for it. Start a project.

Where to go next

Spotted something wrong? Report an error on this page. We correct on the page and say what changed.

Frequently asked questions

For qualitative testing aimed at finding problems, around five to eight people per distinct user group per round is a common working range. It is usually better to run several small rounds with fixes in between than to spend the whole budget on one large session. Quantitative benchmarking or statistical comparison needs larger samples.
In moderated testing a facilitator observes live and can ask follow-up questions, which gives richer insight into why people struggle. Unmoderated testing lets participants complete tasks on their own with recording software, which is faster and cheaper per session but offers no chance to probe. Unmoderated sessions demand very clearly written tasks.
Yes, and you usually should. Testing clickable wireframes or rough prototypes early finds structural problems while they are still cheap to fix. Match fidelity to the question: navigation and comprehension work at low fidelity, while microcopy and form behavior need something closer to final.
Most moderated sessions run 45 to 60 minutes, covering an introduction, a few background questions, 30 to 40 minutes of tasks and a short debrief. Unmoderated sessions are typically shorter because there is no facilitator to keep participants engaged. Set a time limit for each task in advance.
Existing customers can be useful, but testing only with them hides problems because they have already learned how the product works. Recruit a mix that includes people who have never used the product. Newcomers reveal the barriers that stop people from becoming customers in the first place.
Very little. Ask neutral prompts such as what they are looking for or what they expect to happen, and avoid any hint that names an interface element or signals the right answer. If they are clearly stuck and no new information is coming, mark the task as failed and move on.
Outside help pays off when a flow is underperforming and nobody can agree why, when sessions must be moderated with hard-to-recruit specialist users, or when you need accessibility testing with people who use assistive technology. An independent researcher also avoids the bias of a team evaluating its own work.
All services

The work behind this article, and what it costs.

Priya Raghunathan

Interface and identity. Writes about design systems, research, and why most navigation problems are structure problems.

Keep reading