YuzuTrace
Home Markets Report Pricing Blog Questions Sign in Order an audit Write to us
01 Blog · 6 min

Your analytics tell you where. They cannot tell you why.

Your dashboard is right about which step loses people. It has no way of knowing what happened on that screen, and no amount of extra instrumentation will change that.

Open any analytics tool and you will see the same shape every time. A funnel that starts wide and ends narrow, with one step where the drop is steeper than the rest. Cart to shipping. Sign-up to confirmation. Search to product page.

That number is accurate. It is also the outer limit of what the tool knows. It records that a session stopped at a given URL. It cannot record that the visitor read the shipping cost, said a word we will not print here, and closed the tab.

The problem is documented. Its cause on your site is not.

Nobody needs convincing that this is a real problem.

70.22% Average cart abandonment rate, and it has stayed remarkably flat for over a decade, through ten years of heavy investment in conversion optimisation. Baymard Institute, meta-analysis of 50 studies

Read that second line again. If the rate does not move while everyone optimises, the problem is not a shortage of measurement. It is that industry averages say nothing about your particular screen.

Baymard does identify a dominant reason among the fixable abandonments: costs revealed too late, whether that is shipping, sales tax, or handling fees. Useful as a lead. Useless as a diagnosis. Knowing that costs surprise people does not tell you whether, on your page, the culprit is a missing tax line, a badly worded free-shipping threshold, or a delivery date that only appears one step later. Three different causes, three different fixes, one identical number on your dashboard.

Why measuring harder does not close the gap

The standard reflex is to instrument more aggressively. Heatmaps, session recordings, scroll depth, rage clicks. Each one adds resolution, but always to the same kind of evidence: observable behaviour.

A heatmap will show you that eleven people clicked a label that is not a button. It will not tell you they clicked it because the actual button says Confirm and they had no idea what they would be confirming. That sentence exists only inside the visitor's head.

That is the gap, and it is structural. These tools produce hypotheses, never causes. Your numbers are a smoke alarm. Excellent at telling you there is a fire, silent on what is burning.

What a person adds: the reasoning, live

Put a real visitor in front of the same screen, give them a goal, and ask them to say everything that goes through their head as they go. The method has a name and a literature behind it: the think-aloud protocol, formalised by Ericsson and Simon in the early 1980s. It remains the backbone of qualitative usability testing.

Within thirty seconds you get something no dashboard contains. Not user hesitated for four seconds, but I can't tell if tax is included in this, and I've been burned before.

The second sentence is fixable. The first never is. It names the fear, and the fix follows directly from it: show the tax line earlier. One change, one afternoon.

One point of intellectual honesty, because it changes how the study is run. What people say does not always match what they do. Nisbett and Wilson established this back in 1977: we happily rationalise after the fact about decisions whose real causes we cannot access. That is precisely why you have people narrate during the task rather than afterwards, and why you always cross-check what was said against what happened on screen. A tester who says that was clear after clicking the wrong thing three times has just given you two pieces of information, not one.

How many people? What the five-user rule actually says

You have met the number before: five testers is enough. It comes from Nielsen and Landauer (1993), popularised by Jakob Nielsen's Why You Only Need to Test with 5 Users (2000).

85% of a journey's usability problems surface with five testers, on an average 31% chance that any one participant meets any one problem. Nielsen and Landauer, ACM INTERCHI'93

The rule is sound, provided you state its conditions, which almost everyone omits.

  • Five people per segment. A B2B tool with admins, end users and buyers needs fifteen participants, not five.
  • Five people to discover, not to measure. A task success rate or a comparable score needs 20 to 40 participants. That is a different exercise entirely.
  • The average conceals real spread. Faulkner's replication work (2003) confirms the roughly 85% average but shows that some groups of five land well below it. Rare problems slip past small samples.
  • Three rounds of five beat one round of fifteen. You fix between rounds, which means you test the fix instead of assuming it worked.

In short: five people on one specific journey with one homogeneous audience is the best signal-to-cost ratio in user research. It is not a magic number that holds everywhere.

Both, in the right order

None of this devalues your analytics. It makes them a first step rather than a complete method.

  1. Look at the funnel, find the worst step. Ten minutes, no cost.
  2. Send five people through exactly that step, narrating. Listen.
  3. Fix it, then measure again on the same funnel.

The numbers say where to point the study, the study says what to change, the numbers say whether it worked. Doing it the other way round means testing a screen that was never the problem.

Sources

Every figure above comes from the Baymard Institute, from peer-reviewed academic work, or from an official standard. No statistics aggregators, no vendor blogs. Baymard is an independent research organisation specialising in ecommerce usability, founded in 2009; its findings come from moderated one-to-one usability testing and a benchmark database of manually reviewed sites.

Primary ecommerce usability research

Peer-reviewed academic work

  • Nielsen, J. and Landauer, T. K., "A mathematical model of the finding of usability problems", Proceedings of ACM INTERCHI'93, Amsterdam, 1993, pp. 206-213. The 31% and 85% model.
  • Faulkner, L., "Beyond the five-user assumption: Benefits of increased sample sizes in usability testing", Behavior Research Methods, Instruments & Computers, 35(3), 2003, pp. 379-383. The spread around that average.
  • Ericsson, K. A. and Simon, H. A., Protocol Analysis: Verbal Reports as Data, MIT Press, 1984. The think-aloud protocol.
  • Nisbett, R. E. and Wilson, T. D., "Telling more than we can know: Verbal reports on mental processes", Psychological Review, 84(3), 1977, pp. 231-259. Rationalisation after the fact.

Further reading

  • Nielsen, J., Why You Only Need to Test with 5 Users, Nielsen Norman Group, 2000. The popularisation of Nielsen and Landauer, not the source itself, and the readable way in.

Figures verified August 2026. Abandonment reason percentages come from quantitative surveys of US online shoppers, calculated after excluding respondents who reported they were just browsing. Read them as sector-level orders of magnitude, not as values that transfer directly to any one site.

Find out what happens at the step where you lose people.

Real people matching your customer profile, an annotated report, 7 to 10 business days. You pick the journey, we find where it breaks and why.