2026-09-09 · 7 min · By Alcott Dube
How to run a usability test with five users
I use five-person usability tests to find friction in a specific task, with a neutral script, consistent notes, and clear limits on the conclusions.

I run usability testing with five users by recruiting people from one relevant audience, giving them realistic tasks, and observing without coaching. I record outcomes and points of friction, prioritise problems by consequence, then test revisions rather than treating five sessions as a reliable estimate of how everyone behaves.
When are five participants enough for usability testing?
Five participants is a useful starting point for a small formative study: testing a design to find problems worth fixing. It's not a certificate that a product is usable. I keep the scope narrow, usually one flow, one audience, and a few decisions the team can act on.
Nielsen Norman Group's five-user recommendation rests on diminishing returns: additional participants often reveal problems already observed. Its argument favours repeated small studies over spending the entire budget on one large round. That reasoning doesn't mean every fifth session completes the picture, particularly when behaviours or user groups differ.
If I'm testing both experienced administrators and occasional account holders, I don't combine them into a single five-person sample and call both groups covered. I either choose one group for this round or plan separate rounds. For precise completion rates, comparisons between variants, or rare behaviours, I'd use a larger study designed around that question.
How to recruit five relevant test participants
I recruit for behaviour before demographics. For an invoice tool, I'd ask whether someone has created, checked, or chased an invoice recently, which tools they used, and how often. A matching job title doesn't guarantee relevant experience.
My screener uses concrete questions that don't advertise the preferred answer. Instead of asking whether someone is comfortable managing invoices, I'd ask them to describe the last invoice-related task they completed. I exclude people who helped design the flow because their knowledge changes what the test can reveal.
I also check device requirements and access needs before scheduling. If assistive technology is part of the intended use, it belongs in recruitment and session planning, not a footnote afterwards. I arrange consent, recording permission, compensation, and a cancellation plan before the first session.

How to write usability test tasks without giving clues
I write tasks around an intention, not an interface action. Nielsen Norman Group's guidance on task scenarios makes this distinction useful: a task should establish what someone wants to accomplish without naming the controls that get them there. I avoid borrowing navigation labels unless users would naturally arrive knowing them.
For a hypothetical invoicing prototype, I'd use three tasks like these. Each needs prepared account data and a reachable end state:
- You finished a £480 project for an existing customer, Rowan Studio. Send them an invoice that is due in 14 days.
- Rowan Studio says its billing address has changed. Correct the address on the invoice you just prepared.
- The payment has arrived by bank transfer. Update your records so you can distinguish this invoice from unpaid work.
A practical script for a 30-minute usability test
I plan roughly four minutes for the introduction, three for context, eighteen for tasks, and five for follow-up. Before the study, I pilot the session to check links, permissions, task wording, and timing. I also define success for each task. For sending an invoice, reaching a preview isn't the same as confirming delivery.
I open with: "I'm checking how this design works, not testing your ability. Please work as you normally would and say what you're looking for or expecting. You can stop at any time." I then confirm recording consent and ask about their recent experience with the activity.
I give one task at a time. If someone goes quiet, I ask, "What are you thinking about now?" If they ask where to click, I reply, "What would you try if I weren't here?" I leave room for silence. Constant prompting can turn ordinary exploration into a performance for the moderator.
When someone is stuck, I note the point before helping. If later tasks depend on that step, I move them to the necessary state and mark the assistance. I finish by asking which part was hardest and whether anything behaved differently from their expectations.
What to record during each usability test
I use one note-taking sheet across all five sessions. Each task gets an outcome of independent completion, assisted completion, or non-completion. I define these before testing so I don't quietly become more generous with a design I like. A second person can take notes while I moderate, provided their presence is explained.
Alongside the outcome, I record the observed action, its consequence, and a timestamp if recording is permitted. "Opened three invoices to find the overdue one" is evidence. "Needs a dashboard" is a proposed solution. Keeping those apart prevents an early design idea from becoming the finding.
I note wrong turns, repeated attempts, requests for help, and whether the person notices an error. I treat completion time cautiously when participants are thinking aloud or using an incomplete prototype. Comments about confidence can add context, but I don't let a positive rating erase an observed failure to finish.
How to analyse five-user usability test results
I make a participant-by-task table, then group observations by the problem that appears to explain them. Three different wrong clicks may point to one unclear label. The same wrong click may also have different causes. I check the session notes before merging observations into a single finding.
I report counts with context: "Three of five participants needed help finding the payment status." I don't translate that into "60% of customers will struggle." Five qualitative sessions don't provide a dependable population estimate. The Interaction Design Foundation describes usability testing as observing representative users attempting tasks; I keep my claims anchored to those observed attempts.
Frequency is only one input. One person sending an invoice to the wrong recipient can matter more than four people pausing briefly over a label. I assess the consequence, whether recovery was possible, and how strong the evidence is. If nobody encounters a problem, I write "not observed in this round", not "validated".
How to prioritise fixes and plan the next test
I turn each finding into a small decision record: the task, observed behaviour, consequence, supporting sessions, and proposed next action. For example, a participant marking an unpaid invoice as paid without noticing would justify investigating the status control and recovery options. It wouldn't, by itself, justify rebuilding the whole invoicing area.
I prioritise blockers and consequential errors, then recurring confusion that people recover from. Cosmetic preferences usually wait unless they affect understanding or access. I also check alternative explanations: did the task wording confuse people, was necessary information missing from the prototype, or did the moderator accidentally steer the attempt?
For the next round, I keep the task intent and success criteria stable, recruit fresh relevant participants where possible, and test the changes with another small group. Reusing participants risks measuring what they learnt last time. If a suspected cause remains uncertain, I test that assumption before commissioning an expensive redesign.
Questions people ask
Is five users enough for usability testing?
I use five participants to find problems in a focused flow for a reasonably consistent audience. I don't use that sample to estimate population success rates or claim coverage of several distinct user groups.
How many tasks should a usability test include?
For a 30-minute session, I start with three short tasks and pilot the timing. One complex task may fill the available time, so I prioritise the decisions I need to make rather than a fixed task count.
Can I run usability tests with colleagues?
I use colleagues to check the script and setup. I only count them as study participants if they genuinely match the intended audience and lack inside knowledge that would distort their behaviour.
Should I help participants when they get stuck?
I first give them space to try and record where progress stops. If I help so the session can continue, I mark that task as assisted and avoid reporting it as independent completion.