2026-09-30 · 7 min · By Alcott Dube
How to run a heuristic evaluation: checklist and severity
I use a task-based heuristic evaluation checklist to find usability problems, score their severity, and turn observations into fixes without mistaking inspection for user research.

I run a heuristic evaluation by walking through defined product tasks, checking each interaction against Nielsen Norman Group’s ten usability heuristics, and recording specific failures. I score each finding from 0 to 4 using frequency, impact and persistence, then prioritise fixes while keeping observed behaviour separate from assumptions about users.
Choose the tasks and boundaries for your evaluation
I start with a task, not a collection of screens. ‘Change the billing contact without changing the subscription’ gives me a route and a success condition. ‘Review account settings’ does neither. For a first pass, I choose three tasks: the main reason someone uses the product, a frequent maintenance task, and a recovery task such as correcting a failed payment.
I record the starting state, account permissions, device, browser and intended outcome. I also specify what is outside the review. A desktop account-management audit cannot support conclusions about mobile onboarding. For a small product, I would budget an initial two-hour session and narrow the scope if necessary. That is a planning choice, not a research-backed duration. I keep setup time separate, because creating test accounts and realistic data can consume much of the session.
How to evaluate your own product without excusing it
Knowing the product makes inspection faster and judgement less reliable. I already know where settings live and what internal terminology means. I counter that by writing the task before opening the product and noting every point where my existing knowledge supplies information the interface does not. Familiarity is not evidence that an interaction is clear.
Nielsen Norman Group recommends multiple evaluators, commonly three to five, because different people find different problems. A solo review is still useful, but I treat it as one inspection rather than comprehensive coverage. I make one pass to understand the flow, then another to inspect individual interactions. I avoid designing fixes during that second pass. Moving a button is tempting; first I need to establish what makes its current position a problem.

A heuristic evaluation checklist for each task
I use Nielsen Norman Group’s ten heuristics as prompts, not as ten boxes to tick once per screen. I check them across the whole task, including transitions, waiting states and recovery. A screen can look reasonable while the sequence around it fails.
For each prompt, I capture a concrete observation or leave it unflagged. I do not manufacture a finding to complete the checklist. These checks identify suspected usability problems; they do not establish how many customers experience them.
- System status: After an action, can I tell what happened, whether work is continuing, and whether the result was saved?
- Real-world language: Do labels, units and sequences match the task, rather than expose the organisation’s internal terminology?
- Control and escape: Can I cancel, go back or undo an action without losing unrelated work?
- Consistency: Do equivalent controls behave alike, and do familiar platform conventions still apply?
- Error prevention: Are invalid choices constrained, and are consequential actions checked before damage occurs?
- Recognition: Are necessary options and instructions visible where needed, rather than recalled from another screen?
- Efficiency: Can experienced users repeat common tasks with less effort without making the basic route harder to understand?
- Relevant information: Does each element support the current decision, or compete with something more useful?
- Error recovery: Does an error explain the problem in plain language and provide a workable next step?
- Help: When explanation is necessary, is task-specific guidance easy to find and act on?
Record findings that someone else can reproduce
My finding log contains a task, location, starting state, reproduction steps, observed behaviour, likely consequence, heuristic, severity and supporting capture. I add a confidence note when the consequence is inferred. One underlying problem gets one entry, even if it appears on several screens. Otherwise, a repeated component defect can distort the apparent size of the backlog.
A useful hypothetical entry is: ‘After changing the billing contact, selecting Save gives no visible confirmation. Leaving and returning shows the new contact. This creates uncertainty about whether the change succeeded.’ I would tag system status and attach a short recording. ‘The form is confusing’ is not actionable evidence. I keep the proposed fix in a separate field so a plausible solution does not disguise a weak diagnosis. I also note any workaround and the effort it requires.
How to score usability severity from 0 to 4
I use Nielsen Norman Group’s severity scale and its three considerations: how often the problem occurs, how much it disrupts the task, and whether it keeps causing trouble after someone encounters it. Frequency needs care. I can observe that a defect appears on every save without knowing how often customers save.
I assign one severity rating with a short justification rather than inventing a formula. The missing confirmation might deserve a 2 if the result is easily checked. It could warrant a 3 if uncertainty encourages repeated submissions with meaningful consequences. I label that consequence as a hypothesis until supported. A blocked critical task with no recovery route may warrant a 4; an unattractive button does not. When evidence is incomplete, I mark the rating provisional.
- 0: Not a usability problem. I retain the entry only if documenting a rejected concern is useful.
- 1: Cosmetic issue. I would fix it when practical, not ahead of functional usability problems.
- 2: Minor problem. It creates friction but has a manageable workaround.
- 3: Major problem. It materially disrupts a task and deserves high priority.
- 4: Usability catastrophe. It must be resolved before release.
Turn severity ratings into a fix list
Severity describes the usability problem. Priority also considers exposure, business consequences, dependencies and implementation effort. I keep those fields separate. A low-effort cosmetic fix can fit into existing work without pretending it is more severe than a difficult recovery failure. Equally, a rare problem can remain serious if it causes irreversible loss.
I group findings by underlying cause, then choose a small batch with clear acceptance criteria. ‘Improve feedback’ is too loose. ‘After a successful save, show confirmation; after failure, retain the entered values and offer a retry’ is testable. If a fix depends on an unverified assumption about customer behaviour, I specify the question to investigate before committing to a large redesign. I do not use development effort to quietly downgrade the severity rating.
Validate fixes and repeat the same audit
I rerun the original task from the recorded starting state after a fix, including the relevant error path. I check that the problem has disappeared and that the change has not moved it elsewhere. A confirmation message is not enough if it reports success before the underlying operation finishes. For consequential changes, I also seek evidence from representative users.
I track findings closed, findings reopened and unresolved issues by severity. Where product data exists, I compare relevant measures such as completion, repeated submissions or recovery success, while checking for other changes that could explain a difference. My own completion time is not a customer benchmark. I save the scope, checklist and task scripts with the findings so the next evaluation can repeat the conditions. I review changed flows after meaningful releases rather than rerunning every screen on an arbitrary monthly schedule.
Questions people ask
Can I do a heuristic evaluation by myself?
Yes. I can identify useful findings alone, especially with defined tasks and recorded evidence. I would not treat a solo inspection as comprehensive, because familiarity hides problems and different evaluators notice different things.
What is the difference between heuristic evaluation and usability testing?
In a heuristic evaluation, I inspect interactions against usability principles. In usability testing, I observe representative participants attempting tasks. I use inspection to identify suspected problems, not to claim that users have experienced them.
Should I give every screen a usability score?
I score individual findings rather than average a screen’s checklist results. An average can hide a critical failure among several acceptable interactions. I report the affected task and severity distribution instead.
How often should I run a heuristic evaluation?
I evaluate when a significant flow changes, before a consequential release, or when evidence suggests a recurring problem. For an unchanged area, I repeat the review when new evidence or a planned change gives it a clear purpose.