19. Failure modes of experimentation
Five ways experimentation fails. Testing everything and learning nothing. Metrics that move without meaning. Growth before retention. Statistical theater on tiny samples. And experiments nobody wrote down. Each has a tell, and each is a way of feeling scientific while producing noise, which is more dangerous than not experimenting at all, because it comes with confidence.
- What are common mistakes in product experimentation?
- What is a vanity metric?
- Why is a badly run experiment worse than none?
§The question: where does experimentation go wrong, and how do you catch it feeling like science while producing noise?
§The danger of experimentation is that its failures wear the costume of rigor. A team can run experiments diligently and learn nothing, and feel productive the whole time, which is worse than not experimenting, because a confident wrong answer is more expensive than an admitted unknown. Five failures, each with a tell.
§ 19.1Testing everything and learning nothing#
§What it looks like. A high volume of experiments, mostly on trivia, mostly with no clear or threshold. Why it happens. Velocity is mistaken for learning, and small tests are easy to run and feel like progress. The tell. High experiment count, and the team cannot state what it has learned about its users. The defense. Chapter 1 and chapter 2: every experiment starts with a falsifiable hypothesis aimed at the riskiest assumption. Velocity counts only complete, honest loops.
§ 19.2Metrics that move without meaning#
§What it looks like. The team tracks and celebrates numbers that go up but do not connect to value or revenue: total sign-ups, page views, downloads, time in app. Why it happens. These numbers are easy to move and always go up with effort, so they feel like success. Eric Ries named them vanity metrics. The tell. The celebrated metric has risen and revenue, retention and activation have not. The defense. Every metric traces to and to value. Cohorts, not cumulative totals. The one-metric-that-matters discipline from chapter 3.
§ 19.3Growth before retention#
§What it looks like. Acquisition is turned on while the retention curve still decays. Why it happens. The pressure to show growth is strongest exactly before the has earned it, and acquisition makes the top-line number move. The tell. Sign-ups rising, active users flat or falling, budget draining. The defense. The whole of Part II. No growth spending until retention flattens, activation is reliable, and the target segment has paid. This is the failure the whitepaper is most designed to prevent.
§ 19.4Statistical theater#
§What it looks like. A/B tests on tiny samples reported with confidence, differences of a few users called significant, experiments stopped when they look good. Why it happens. The forms of rigor are copied without the substance, and small teams rarely have the sample the tests require. The tell. "Conversion doubled" on twelve users. P-values on samples of dozens. Results that evaporate in the next . The defense. Chapter 4: know what your sample can answer, run to a fixed stopping point, and when the sample is too small, read qualitatively and say so rather than dressing a coin flip as a test.
§ 19.5Experiments nobody wrote down#
§What it looks like. Experiments run, decisions made, and nothing recorded. A year later the team re-runs old experiments and cannot say why it killed features. Why it happens. Documentation feels like overhead when the team is small and moving fast. The tell. A new hire asks what the team has learned and no one can produce it. The same experiment appears in the backlog twice, months apart. The defense. Chapter 16: one entry per experiment, searchable, and chapter 18: results turned into durable insights.
§ 19.6The root#
§Every failure here is the same one the whole collection keeps naming: something happened without a falsifiable claim behind it. Testing without a hypothesis. Celebrating a metric with no connection to value. Growing without the retention that justifies it. Reporting a result the sample cannot support. Deciding without recording why. Each feels like work, and each produces confidence instead of learning, which is the worst possible trade.
§ 19.7What comes out#
§The five, pinned by the backlog and read at the weekly meeting. Before each experiment, the check: is there a falsifiable hypothesis, aimed at the riskiest assumption, with a threshold set in advance and a sample that can support the read? After each, the check: was it written down, and what did we learn about the users? The conclusion is what AI changes about all of this.
A badly run experiment is worse than none, because it produces a confident wrong answer. Every failure here feels scientific and teaches nothing.
- Eric Ries, The Lean Startup (2011), on vanity versus actionable metrics. theleanstartup.com
- Ronny Kohavi, Diane Tang and Ya Xu, Trustworthy Online Controlled Experiments (2020). experimentguide.com
- Evan Miller, How Not to Run an A/B Test (2010). www.evanmiller.org/how-not-to-run-an-ab-test.html