3. Metric, threshold, decision rule
One metric per experiment. One threshold, the number that means yes, written before the experiment runs. And a decision rule that says what the team will do at each outcome, also written in advance. Deciding these after the data arrives is how teams talk themselves into whatever they hoped. Deciding them before is what makes an experiment an experiment.
- How do you choose a metric for a product experiment?
- What is a decision rule in experimentation?
- Why set a success threshold before running a test?
§The question: what one number will we read, what value means yes, and what will we do either way, all decided before we look?
§The names a prediction with a number. This chapter turns that into the three commitments that make an experiment trustworthy, and all three are made before the experiment runs.
§ 3.1One metric#
§One. The experiment moves one number, and that number is chosen before the test. Alistair Croll and Benjamin Yoskovitz's one-metric-that-matters is the discipline: at any moment a team is trying to move one thing, and an experiment that watches five metrics will find one that moved and call it a win.
§Choosing the metric is choosing what the experiment is about. If the hypothesis is about whether recognition drives return, the metric is repeat-visit rate for the recognized cohort, not sign-ups, not app opens, not satisfaction. Ronny Kohavi's teams call the chosen measure the overall evaluation criterion, and insist it be agreed before the test, because a metric chosen after the data is a metric chosen to give the answer you want.
§The other numbers are not ignored; they are guardrails. You watch them to make sure the thing you moved did not break something else. But one metric decides the experiment.
§ 3.2One threshold#
§The number that means yes, written down before the run. Not "we'll see if it goes up." Up by how much, and why that much.
§The threshold comes from the business, not from statistical convenience. If the need a 60% profile completion rate to work, 60% is the threshold, and 55% is a fail even though it is higher than today. If a 2% lift would not change any decision, then 2% is not worth testing for, and the experiment should be designed to detect something that matters or not run at all.
§ 3.3One decision rule#
§What the team does at each outcome, decided in advance. Above the threshold: what happens, ship it, scale it, build the next thing on it. Below: what happens, kill it, iterate once and retest, go back to the drawing board. In the ambiguous middle, if there is one: what happens.
§The decision rule is the antidote to what Kahneman calls outcome bias, the tendency to judge a decision by how it turned out and to rewrite what we intended once we know the result. A rule written before the data cannot be rewritten by the data. The team that agreed "below 55% we cut the feature" and then sees 48% has already decided, and the decision is clean because it was made when nobody knew the answer.
| Weak setup | Strong setup |
|---|---|
| "Let's ship the new onboarding and see if engagement improves." | "Metric: day-one profile completion. Threshold: 60%, because the model needs it. Rule: at or above 60% we keep it; below 55% we revert and try a different cut; between, we run it two more weeks." |
§ 3.4The three together#
§The three are one act: before the experiment, the team writes the metric, the threshold and the rule on a card, and everyone agrees. That card is the experiment's contract with itself. When the data comes, there is no meeting to decide what it means, because the meaning was decided when the card was written. The meeting is only to confirm the number and execute the rule.
§ 3.5What you leave with#
§One card per experiment: the metric, the threshold with its reason, the decision rule for each outcome, and the guardrail metrics to watch. Signed off before the experiment starts. The next chapter is how long to run it and how few users is too few.
One metric, one threshold written before the run, one rule for each outcome. Decide what the number means before you see it.
- Ronny Kohavi, Diane Tang and Ya Xu, Trustworthy Online Controlled Experiments (2020), on overall evaluation criteria. experimentguide.com
- Alistair Croll and Benjamin Yoskovitz, Lean Analytics (2013), on the one metric that matters. leananalyticsbook.com
- Daniel Kahneman, Thinking, Fast and Slow (2011), on outcome bias. www.penguinrandomhouse.com/books/89308/thinking-fast-and-slow-by-daniel-kahneman