Jul 27, 2023 · 1h 23m · lennys-podcast
The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon)
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
Experimentation authority Ronny Kohavi joins Lenny Rachitsky to deliver a comprehensive guide on building trustworthy A/B testing platforms, avoiding statistical pitfalls, and balancing incremental compounding with high-risk product bets.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Lenny holds 20.2% of the talking time here. How this is scored →
speaking balance: gold is Lenny, purple is the guest (3 minute bins)
Ronny directly disputes the host's premise that Airbnb's top-down, anti-experimentation shift was validated by business success, arguing counterfactually that the company would be significantly larger had it continued rigorous testing.
Hardest push from Lenny ▶ 38:42 Pushing for large redesigns to escape local maximaLenny challenges Ronny's strict incrementalism, arguing that product teams cannot always iterate one factor at a time and sometimes must take radical swings to break out of local maxima.
Biggest teaching moment ▶ 1:04:30 Deconstructing p-values and false positive riskRonny dismantles the common misunderstanding of p-values, using Bayesian reasoning to show how an 8% prior success rate turns a p<0.05 result into a startling 26% false positive risk.
Lenny holds their own ▶ 9:01 Lenny citing Airbnb's new-tab search winLenny demonstrates his deep domain experience by introducing the exact search listing new-tab experiment from his tenure at Airbnb, matching Ronny's enterprise testing examples with a concrete insider case study.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Lenny as informed peer | Guest teaching | Guest disagreement | Lenny pushing back | Why |
|---|---|---|---|---|---|---|
| Bing's $100M Ad Headline Experiment | 4 | 7 | 1 | 0 | Lenny opens with a prompt about surprising A/B test results, prompting Ronny to explain the famous Bing headline experiment in detail. Lenny demonstrates familiarity with the mechanism by summarizing the exact UI shift. Ronny educates the audience and host on how counterintuitive even small wins can be. | |
| Opening Links in New Tabs and Institutional Memory | 6 | 6 | 1 | 0 | Lenny brings up a parallel experiment from his time at Airbnb regarding opening search listings in new tabs. Ronny builds on it by detailing the lineage of that exact test dating back to 2008 at MSN and Hotmail, schooling the host on institutional memory and failure rates. | |
| Baseline Failure Rates Across Tech Companies | 4 | 8 | 2 | 0 | Lenny asks whether a 92% failure rate is typical across tech companies. Ronny breaks down benchmark failure rates across Microsoft, Bing, Booking, and Google Ads, disabusing PMs of the notion that their success rates are unusually high. | |
| Codifying UI Patterns with GoodUI and Rules of Thumb | 5 | 7 | 1 | 0 | Lenny asks about codified repositories of UI patterns, suggesting feeding them into LLMs for PM roadmaps. Ronny introduces GoodUI and his Microsoft rules-of-thumb paper, explaining the mechanics of institutional memory and surprising negative results like the Windows indexer. | |
| Portfolio Strategy and the Bing Social Search Failure | 5 | 6 | 2 | 1 | Lenny raises the classic critique that experimentation leads to local maxima and micro-optimizations. Ronny counters with portfolio theory and recounts Bing's 100-person-year failed social search bet, while Lenny reinforces with similar failures at Netflix and Airbnb. | |
| When Startups Should Start Experimenting | 4 | 8 | 2 | 0 | Lenny asks for tactical startup rules of thumb on when to begin experimentation. Ronny provides a strict statistical cutoff, explaining why math breaks down below tens of thousands of users and naming 200k users as the threshold for true velocity. | |
| Designing the Overall Evaluation Criterion (OEC) | 5 | 8 | 2 | 0 | Lenny prompts Ronny on the Overall Evaluation Criterion (OEC) framework. Ronny explains why single revenue metrics incentivize degradation of UX and shows how constrained optimization and guardrail metrics enforce long-term health. | |
| Lifetime Value Models and Amazon's Unsubscribe Lesson | 4 | 8 | 1 | 0 | Lenny asks about methods to capture long-term metrics and lifetime value. Ronny provides a detailed case study from Amazon's email recommendation team where modeling the dollar cost of unsubscribes reversed over half of the team's seemingly positive campaigns. | |
| The Pitfalls of Radical Redesigns vs. OFAT | 5 | 7 | 3 | 2 | Lenny notes how radical redesigns almost always hurt metrics and asks how to handle passionate teams wanting a full overhaul. Ronny advocates for One Factor At a Time (OFAT), highlighting Microsoft PM denial where teams claimed they were too smart to fail 50% of the time. | |
| Sponsor Break: Eppo | 4 | 6 | 2 | 1 | Following the Eppo ad read, Lenny asks whether radical redesigns are ever worth the risk to escape local maxima. Ronny reiterates the 80% failure rule and firmly rejects shipping flat experiments due to codebase complexity and maintenance overhead. | |
| Airbnb's Product Direction, COVID Shifts, and Counterfactuals | 6 | 8 | 5 | 2 | Lenny questions Ronny on Airbnb's recent top-down product shift under Brian Chesky and the turning off of performance marketing during COVID. Ronny pushes back firmly, arguing counterfactually that Airbnb would be significantly bigger today had it maintained rigorous experimentation. | |
| Experiment Platform Trust and Optimizely's P-Value Flaw | 4 | 8 | 4 | 0 | Lenny asks why Ronny emphasizes trust over velocity. Ronny delivers a detailed breakdown of Optimizely's early statistical flaws with real-time p-value peeking, showing how false positive rates inflated to 30% and eroded market credibility. | |
| Diagnosing Sample Ratio Mismatches (SRM) | 4 | 8 | 3 | 0 | Lenny asks how to spot invalid tests and whether incorrect assignment points are the primary culprit. Ronny explains Sample Ratio Mismatch (SRM) diagnostics, citing bots and pipeline issues, and humorously recounts how PMs bypassed warnings until scorecards were struck through with red lines. | |
| Twyman's Law: Why Extreme Results Are Usually Wrong | 4 | 9 | 3 | 0 | Lenny asks for an explanation of Twyman's Law and p-values. Ronny explains why extreme results are almost always flaws and breaks down Bayes' rule, revealing that with low base success rates (like Airbnb's 8%), a p<0.05 result carries a 26% false positive risk. | |
| Implementing Experimentation: Build vs. Buy and Team Selection | 4 | 7 | 2 | 0 | Lenny asks how to start experimenting and how to change anti-testing cultures. Ronny discusses build vs buy decisions, advising practitioners to establish beachheads with fast-shipping teams, and illustrates the danger of poorly defined OECs using a Microsoft support site example. | |
| Scaling Experimentation Platforms and Metric Maturity | 3 | 8 | 2 | 0 | Lenny asks about scaling platforms to avoid one-off experiments. Ronny outlines the crawl-walk-run-fly maturity framework, noting how Bing scaled to 10,000 metrics and contrasting that with Airbnb's reliance on data scientist headcount to make up for platform gaps. | |
| Accelerating Experiment Velocity with Variance Reduction | 5 | 7 | 1 | 0 | Lenny asks about accelerating test velocity, leading Ronny to detail variance reduction techniques like metric capping and CUPED before transitioning smoothly through the lightning round book recommendations and closing banter. |