A clickable prototype, staging system, or live product with the test paths is available and stable.
Usability Testing
Prerequisite
What needs to be finished first
3-5 realistic tasks with clear success criterion are formulated, derived from actual usage scenarios.
Preparation
What needs to be ready before start
Testable artifact (prototype, staging, live); task script with exact wording; recording tool (Lookback, Maze, Zoom with screen share); observation template (task, observed behavior, problems, severity 1-4); participant incentive.
One moderator for sessions; one to three silent observers from the team for notes; optional note-taker for synthesis; recruiter for 5-8 participants from target group.
Tasks in order with wording; success criteria per task; known hypotheses (what we expect to fail); participant demographics; test environment with data setup.
1-2 h per session, plus 4-8 h preparation and synthesis
Plan 5-8 sessions of 45-60 min each (Nielsen threshold: 5 testers uncover about 85% of usability problems). Pilot with 1 person, then main runs. Silent observers in room or remote, moderator leads alone. Define severity scale (1=show-stopper, 4=cosmetic).
Core question
The one question this method answers
Can typical users complete the central tasks successfully, efficiently, and satisfactorily, and where do they fail on what?
Flow
Marker: Phase
| Step | Duration | Action | Hint |
|---|---|---|---|
1Phase 1: Briefing and warm-up | 5-10 min | Welcome participant, explain purpose ("We are testing the product, not you"). Introduce think-aloud method, confirm consent. Demographic warm-up questions. | Putting testers under pressure distorts results. Explicitly say "There is no wrong here". If tester goes silent, gently remind think-aloud. |
2Phase 2: Run tasks | 30-45 min | Read tasks in order. Testers work independently, moderator observes silently. If stuck, ask carefully ("What are you thinking right now?"). Do not help, do not suggest solution path. | If tester is stuck after 5 min and success criterion is not reached, stop task. That is a data point, not a failure. |
3Phase 3: Post-task questions | 5-10 min | After each task: difficulty rating (1-7), short "What was frustrating?" and "What was helpful?". At end, SEQ or UMUX-Lite questionnaire. | Quantitative mini-scale plus qualitative rationale. Only quantitative is thin, only qualitative hard to compare. |
4Phase 4: Debrief per session | 10 min after each session | Moderator and observers note top-3 problems each. Assign severity. Secure and store recording. | Debrief immediately after session. Otherwise sessions blur in memory and detail is lost. |
5Phase 5: Synthesis and ranking | 3-6 h after all sessions | Aggregate all problems, consolidate duplicates. Per problem: frequency (how many testers affected) and severity. Sort top problems by frequency x severity. Derive recommendations. | Frequency alone misleads. Show-stopper in 1/5 matters more than cosmetic note in 5/5. Take severity seriously. |
Artifact
What comes out at the end
Usability report with sample description, task success rates, problem list (sorted by frequency x severity), quote evidence, recommendations per problem, and recording links (internal).
Own reports per test round with date and artifact version. Retests after fixes with explicit "tested against version Y on ..." status. Visualize trajectory of the most important problems across test rounds.
- Lookback for remote recording and highlights
- Maze for asynchronous tests and automated recording
- Zoom with screen share and manual notes
- UserTesting for external tester pools
- Dovetail for synthesis and recording tagging
usability-testing-working-template.md
Compact working template for Usability Testing with context, input, output artifacts, and next step.
Usability Testing Working Template
Goal
Observes people using a design to uncover usability problems.
Context
When and for what do we use this method?
Input
Which data, observations, decisions, or materials are available?
Execution
Short notes along the runsheet.
Output artifacts
- Usability Issues:
- Severity List:
- Recommendations:
Assumptions and open questions
- ...
Decision / Next step
Owner, date, and success signal.
Example output
Concrete filled scenario, fictional example
usability-testing-beispiel.md
Concrete filled scenario, fictional example
Usability Test - Checkout prototype v2 (CW 19/2026, n=5)
Sample: 5 online shoppers from target group (3 mobile, 2 desktop). Incentive EUR 30.
Task success rates
| Task | Success | Avg time | Avg SEQ |
|---|---|---|---|
| Add product to cart | 5/5 | 0:42 | 6.4 |
| Start checkout | 5/5 | 0:18 | 6.8 |
| Enter delivery address | 3/5 | 2:24 | 4.2 |
| Choose payment method + pay | 4/5 | 1:48 | 5.0 |
| Understand order confirmation | 5/5 | 0:25 | 6.2 |
Top problems
P1 (Show-stopper, 2/5 affected, Severity 1): ZIP field validates on blur, error text appears BELOW the next field. Testers did not see error and failed on submit.
- Quote (T2): "I pressed three times and nothing happened, then I gave up."
- Recommendation: inline error directly below ZIP field, plus focus on affected field on submit error.
P2 (Major, 3/5 affected, Severity 2): "Delivery address same as billing address" checkbox was off, testers filled address twice.
- Quote (T4): "Why do I have to enter this twice?"
- Recommendation: checkbox on by default.
P3 (Moderate, 2/5 affected, Severity 3): Apple Pay button was not recognizable as clickable (too small, missing affordance).
- Recommendation: larger touch area, clear CTA label.
Recommendation priority
- Fix P1 before launch (show-stopper)
- Fix P2 before launch (cheap, high effect)
- P3 in next iteration (Apple Pay adoption lever)
Pitfalls
Recognize symptoms and steer against them
Moderator helps
When difficulties arise, moderator gives hints or defends the design.
Practice standard phrases: "What are you thinking?", "What would you do next?". Tolerate silence. Help completely distorts the result.
Artificial usage situation
Tester sits in a room before unknown device with observers; behavior differs from real use.
Test on tester's own device where possible (remote test, own phone). Observers invisible or minimal. Realistic task scenarios.
Too many tasks
10+ tasks in one session, testers get tired, last tasks produce garbage data.
Maximum 5 main tasks in 60 min. If more is needed, run several sessions or shorter tasks. Fatigue creates pseudo-data.
Confirmation bias in synthesis
Problems the team dislikes are categorized as "user error" and ignored.
Rule: anything affecting two or more testers is always a design problem. Discuss severity, not existence.
No repetition after fix
Problems are fixed but not tested again. It is unknown whether fix works or creates new problems.
Retest after major fixes with 3-5 new testers. Fix verification is part of Definition of Done.
Stop criteria
Done signals checkable in under a minute
Finished the runsheet?
Go to the profile for purpose, similar methods, and sources or continue to the next method in the catalog.