With AI
AI generates experiment variants in minutes, analyzes results in the room, and drafts hypotheses from the funnel, which raises experiment velocity enormously. That is the danger as much as the gift: when running an experiment costs nothing, teams test because they can, and the discipline of choosing the riskiest assumption and setting a threshold in advance matters more, not less. The riskiest assumption is still chosen by a person.
- How does AI change product experimentation?
- Can AI run experiments and analyze results?
- What parts of experimentation still need a human?
§The whitepaper has described a discipline that existed before AI and exists after. This conclusion is about what changed, and, as with the and strategy, the change is concentrated in the middle and absent at the ends.
§ 20.1What got cheap#
§Variants. Generating the alternatives to test, copy, layouts, onboarding flows, email versions, is now minutes of work. A team can produce ten variants of an onboarding as fast as it once produced one, and build the winning one from a prompt.
§Analysis. Reading a result, pulling the , checking the guardrails, running the honest statistics, is now something a model does in the room while the team watches, rather than a data pull that takes a day. The weekly review can read more experiments because reading each is faster.
§Hypothesis drafting. A model given and the backlog will propose hypotheses, score them on impact, confidence and ease, and suggest thresholds. The proposals are a strong starting point, and the scoring conversation is faster for having a draft.
§Together these raise experiment velocity, the metric chapter 15 named, by a large factor. A team can complete far more per week than before, and since learning compounds, that is a real and significant gain.
§ 20.2Why cheap is dangerous#
§Here is the trap, and it is the same trap the sprint's conclusion named for prototypes. When running an experiment costs almost nothing, teams test because they can, not because the test answers the riskiest assumption. The backlog fills with cheap experiments on trivia, velocity soars, and learning does not, because velocity of the wrong experiments is just fast noise.
§The discipline that mattered before matters more now, precisely because the friction that used to enforce it is gone. When an experiment cost a week to build and a day to analyze, the cost forced the team to choose carefully. Now nothing forces the choice, so the team must impose it: the riskiest-assumption filter from chapter 2, the threshold set in advance from chapter 3, the honest sample read from chapter 4. These were always the discipline; now they are the only thing standing between a team and a firehose of meaningless green results.
§ 20.3What did not get cheap#
§Choosing the riskiest assumption. A model can list assumptions and even guess at their consequence and uncertainty. It cannot know which one, if false, would end this particular company, because that judgment depends on the founders' conviction, the runway, and the market's specifics in a way no model has access to. The single most important act in experimentation, pointing it at the thing that matters most, is still human.
§Reading behavior for meaning. A model reports that retention rose. It cannot tell you that it rose because a different kind of user arrived that week, or that it rose for a reason that will not persist, unless a person who knows the market and the users asks the question. The of chapter 18, the durable statement about who these users are, is a human synthesis of what the numbers and the qualitative reads mean together.
§The honesty. Kohavi's warnings about trustworthy experiments, about peeking, about data quality, about underpowered tests, apply with more force when a tool will happily generate a confident analysis of a sample too small to support it. A model asked whether a result is significant will often produce a plausible answer regardless of whether the sample permits one. The person who knows to ask "can our sample even answer this?" is the guardrail, and there is no prompt for skepticism.
§ 20.4The human loop inside the machine loop#
§Ethan Mollick's frame closes this whitepaper as it closed the others: give the tireless collaborator the tireless work, generating variants, pulling cohorts, drafting hypotheses, and keep the judgment. In experimentation the judgment is three acts: choosing the riskiest assumption, setting the threshold that means yes before the data arrives, and reading the result for what it means about the users. Those three are the experiment. Everything else is now nearly free, and nearly free is exactly why the three must be held deliberately.
§ 20.5What comes out#
§A validated , a set of growth loops, a system that learns weekly, and a body of insight about the users that compounds and feeds back into strategy and the model. The seventh and final whitepaper steps back from the process to the essay: what it means to own product in an era when building is cheap and the thinking, the choosing, the reading, the deciding, is the whole of the work.
AI makes experiments nearly free to run, which makes choosing the right one, and reading it honestly, the whole job.
- Ethan Mollick, Co-Intelligence (2024). www.penguinrandomhouse.com/books/741805/co-intelligence-by-ethan-mollick
- Ronny Kohavi, Diane Tang and Ya Xu, Trustworthy Online Controlled Experiments (2020). experimentguide.com
- Eric Ries, The Lean Startup (2011). theleanstartup.com