A concrete IA structure (hierarchy from card sorting or existing navigation) exists and should be tested.
Tree Testing
Prerequisite
What needs to be finished first
5-10 typical search tasks from real user scenarios are formulated without revealing the target category in the task wording.
Preparation
What needs to be ready before start
Tree-testing tool (Treejack by Optimal Workshop, UserZoom, Maze); tree structure as JSON or CSV; tasks with target path marked; survey setup; recruiting link.
One UX researcher as owner; recruiter for 15-30 participants per variant; optional stakeholder for task validation; notes recipient in the team.
IA structure as tree with all nodes and leaf nodes; task list with target paths; hypotheses (which tasks likely fail); target-group demographics.
1-2 days setup, 3-7 days data collection, 1-2 days analysis
Build tree in the tool (all nodes exactly as planned). Formulate tasks: concrete search task, do NOT name category. Pilot with 2-3 people. Main run with 15-30 participants (more for variant comparison).
Core question
The one question this method answers
Can participants find the correct answers in the proposed navigation structure, and which paths lead them astray?
Flow
Marker: Phase
| Step | Duration | Action | Hint |
|---|---|---|---|
1Phase 1: Build tree | 2-4 h | Import or manually create IA structure in the tool. Label every node exactly as planned. Omit visual aids, only text hierarchy. | One wrong label in the tree distorts all downstream tasks. Before pilot, check tree against IA document node by node. |
2Phase 2: Formulate tasks | 2-3 h | Formulate realistic search scenario per task. Mark target path in the tool (can have several correct paths). | Avoid category words in task text. If category is "Account & Profile", do not write "in your account" in the task, otherwise the text gives away the answer. |
3Phase 3: Pilot | 1-2 h | Let 2-3 people run through the test. Watch for confusion: unclear tasks, ambiguous labels, tree too deep. Adjust tasks or tree after pilot. | Pilot reveals format errors early. Running the main test directly after a faulty pilot wastes data. |
4Phase 4: Main run | 3-7 days runtime | 15-30 participants per tree variant. For A/B comparison (old vs new), at least 30 per group. Anonymous asynchronous work, tool records paths. | Below 10 participants, statistical signal is thin. Multiple variants require exponentially more participants. |
5Phase 5: Analysis and iteration | 1-2 days | Per task: success rate (correct leaf nodes), directness (direct path without detours), most frequent wrong paths. Tasks below 60% success rate are critical. Adjust label or hierarchy, retest if needed. | Success rate alone misleads. Low directness plus high success means trial and error. Path analysis shows where users fail. |
Artifact
What comes out at the end
Tree-testing report with tree visualization, task list, success-rate table per task, directness values, top-3 wrong paths per task, sample description, and prioritized recommendations for IA adjustment.
Own test run per tree version with date. Explicitly document comparison between versions (improvement per task in percentage points). Store tree snapshots as JSON export in repo so IA history remains traceable.
- Treejack by Optimal Workshop
- UserZoom with tree-test module
- Maze for structural tests
- Lyssna (formerly UsabilityHub)
- Custom setup with survey tool plus manual analysis
tree-testing-working-template.md
Compact working template for Tree Testing with context, input, output artifacts, and next step.
Tree Testing Canvas
Context
What is this method used for?
Core question
Which question should be answered at the end?
Input
Which data, observations, or materials are available?
Working area
- Area 1:
- Area 2:
- Area 3:
- Relationships / patterns:
Output artifacts
- Findability Metrics:
- Path Analysis:
- Revised IA:
Open questions
- ...
Next step
Owner, date, success signal.
Example output
Concrete filled scenario, fictional example
tree-testing-beispiel.md
Concrete filled scenario, fictional example
Tree Testing - Help Center IA v2 (CW 20/2026, n=24)
Tree variant: New IA with 6 top-level categories (see Card Sorting CW 18) Sample: 24 existing users, asynchronous via Treejack, EUR 10 incentive
Task results
| Task | Success Rate | Directness | Most frequent wrong path |
|---|---|---|---|
| Reset password | 96% (23/24) | 87% | Account & Security -> Account Settings -> correct |
| Download invoice | 79% (19/24) | 63% | Billing -> Invoices correct (error: 5 first went to Features) |
| Create API token | 42% (10/24) | 28% | CRITICAL: 8 went to Account, 6 to Features, only 10 to Integrations & API |
| Invite employee | 71% (17/24) | 54% | Account -> Employees wrong (correct: Getting Started) |
| Cancel subscription | 88% (21/24) | 79% | Billing -> Cancellation correct |
Insights
- API-token task is show-stopper: 58% fail, almost no direct path usage. Label "Integrations & API" is not recognized as API management.
- Invite employee is searched in account area (mental model), not Getting Started. Consider moving or cross-linking.
- Invoice download works, but 21% first go to Features. Check clarity of "Features" label.
Recommendations
- Rename "Integrations & API" to "Developers & API" and increase top-level visibility
- Cross-link "Invite employee" additionally from Account
- Rename or split Features area more clearly
Pitfalls
Recognize symptoms and steer against them
Task text gives away answer
Task contains category keyword, testers find path too easily.
Formulate task without category language. Instead of "Where do you find billing?", use "You want to download the last invoice, where do you click?" Pilot checks wording.
Too few participants
Test with 5-8 participants, success rates fluctuate extremely.
At least 15 for pattern insight, 30+ for statistically more reliable comparisons. A/B needs separate samples.
Tree too deep
Hierarchy has 5+ levels, testers lose orientation and jump back.
Limit depth to max 3-4 levels. If more is needed, restructure IA instead of tree-testing.
Findability confused with visuals
Stakeholders expect insight about design, tree test only provides text-structure insight.
Set expectations upfront: tree test measures IA and labels, not visual design. For design feedback, run usability test or first-click test.
No iteration after test
Test shows problems, IA is rolled out unchanged.
Use test findings as gate for IA launch. Below 60% success rate, iterate and retest instead of launching.
Stop criteria
Done signals checkable in under a minute
Finished the runsheet?
Go to the profile for purpose, similar methods, and sources or continue to the next method in the catalog.