Jul 27, 2023 · 1h 23m · lennys-podcast

The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon)

Ronny Kohavi · 58m spoken Lenny Rachitsky · 14m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Experimentation authority Ronny Kohavi joins Lenny Rachitsky to deliver a comprehensive guide on building trustworthy A/B testing platforms, avoiding statistical pitfalls, and balancing incremental compounding with high-risk product bets.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Lenny holds 20.2% of the talking time here. How this is scored →

Lenny as informed peer 4.5 Guest teaching 7.4 Guest disagreement 2.2 Lenny pushing back 0.3
05100:0020:0040:001:00:001:20:004:30–9:00 · Lenny as informed peer 4/10 Bing's $100M Ad Headline Experiment Lenny opens with a prompt about surprising A/B test results, prompting Ronny to explain the famous Bing headline experiment in detail. Lenny demonstrates familiarity with the mechanism by summarizing the exact UI shift. Ronny educates the audience and host on how counterintuitive even small wins can be.9:01–13:17 · Lenny as informed peer 6/10 Opening Links in New Tabs and Institutional Memory Lenny brings up a parallel experiment from his time at Airbnb regarding opening search listings in new tabs. Ronny builds on it by detailing the lineage of that exact test dating back to 2008 at MSN and Hotmail, schooling the host on institutional memory and failure rates.13:17–15:36 · Lenny as informed peer 4/10 Baseline Failure Rates Across Tech Companies Lenny asks whether a 92% failure rate is typical across tech companies. Ronny breaks down benchmark failure rates across Microsoft, Bing, Booking, and Google Ads, disabusing PMs of the notion that their success rates are unusually high.15:36–20:44 · Lenny as informed peer 5/10 Codifying UI Patterns with GoodUI and Rules of Thumb Lenny asks about codified repositories of UI patterns, suggesting feeding them into LLMs for PM roadmaps. Ronny introduces GoodUI and his Microsoft rules-of-thumb paper, explaining the mechanics of institutional memory and surprising negative results like the Windows indexer.20:45–24:48 · Lenny as informed peer 5/10 Portfolio Strategy and the Bing Social Search Failure Lenny raises the classic critique that experimentation leads to local maxima and micro-optimizations. Ronny counters with portfolio theory and recounts Bing's 100-person-year failed social search bet, while Lenny reinforces with similar failures at Netflix and Airbnb.24:48–28:00 · Lenny as informed peer 4/10 When Startups Should Start Experimenting Lenny asks for tactical startup rules of thumb on when to begin experimentation. Ronny provides a strict statistical cutoff, explaining why math breaks down below tens of thousands of users and naming 200k users as the threshold for true velocity.28:01–32:42 · Lenny as informed peer 5/10 Designing the Overall Evaluation Criterion (OEC) Lenny prompts Ronny on the Overall Evaluation Criterion (OEC) framework. Ronny explains why single revenue metrics incentivize degradation of UX and shows how constrained optimization and guardrail metrics enforce long-term health.32:43–36:30 · Lenny as informed peer 4/10 Lifetime Value Models and Amazon's Unsubscribe Lesson Lenny asks about methods to capture long-term metrics and lifetime value. Ronny provides a detailed case study from Amazon's email recommendation team where modeling the dollar cost of unsubscribes reversed over half of the team's seemingly positive campaigns.36:31–41:47 · Lenny as informed peer 5/10 The Pitfalls of Radical Redesigns vs. OFAT Lenny notes how radical redesigns almost always hurt metrics and asks how to handle passionate teams wanting a full overhaul. Ronny advocates for One Factor At a Time (OFAT), highlighting Microsoft PM denial where teams claimed they were too smart to fail 50% of the time.41:48–45:38 · Lenny as informed peer 4/10 Sponsor Break: Eppo Following the Eppo ad read, Lenny asks whether radical redesigns are ever worth the risk to escape local maxima. Ronny reiterates the 80% failure rule and firmly rejects shipping flat experiments due to codebase complexity and maintenance overhead.45:41–50:06 · Lenny as informed peer 6/10 Airbnb's Product Direction, COVID Shifts, and Counterfactuals Lenny questions Ronny on Airbnb's recent top-down product shift under Brian Chesky and the turning off of performance marketing during COVID. Ronny pushes back firmly, arguing counterfactually that Airbnb would be significantly bigger today had it maintained rigorous experimentation.50:06–55:26 · Lenny as informed peer 4/10 Experiment Platform Trust and Optimizely's P-Value Flaw Lenny asks why Ronny emphasizes trust over velocity. Ronny delivers a detailed breakdown of Optimizely's early statistical flaws with real-time p-value peeking, showing how false positive rates inflated to 30% and eroded market credibility.55:28–1:00:44 · Lenny as informed peer 4/10 Diagnosing Sample Ratio Mismatches (SRM) Lenny asks how to spot invalid tests and whether incorrect assignment points are the primary culprit. Ronny explains Sample Ratio Mismatch (SRM) diagnostics, citing bots and pipeline issues, and humorously recounts how PMs bypassed warnings until scorecards were struck through with red lines.1:00:45–1:06:20 · Lenny as informed peer 4/10 Twyman's Law: Why Extreme Results Are Usually Wrong Lenny asks for an explanation of Twyman's Law and p-values. Ronny explains why extreme results are almost always flaws and breaks down Bayes' rule, revealing that with low base success rates (like Airbnb's 8%), a p<0.05 result carries a 26% false positive risk.1:06:28–1:10:18 · Lenny as informed peer 4/10 Implementing Experimentation: Build vs. Buy and Team Selection Lenny asks how to start experimenting and how to change anti-testing cultures. Ronny discusses build vs buy decisions, advising practitioners to establish beachheads with fast-shipping teams, and illustrates the danger of poorly defined OECs using a Microsoft support site example.1:10:19–1:12:21 · Lenny as informed peer 3/10 Scaling Experimentation Platforms and Metric Maturity Lenny asks about scaling platforms to avoid one-off experiments. Ronny outlines the crawl-walk-run-fly maturity framework, noting how Bing scaled to 10,000 metrics and contrasting that with Airbnb's reliance on data scientist headcount to make up for platform gaps.1:12:24–1:21:34 · Lenny as informed peer 5/10 Accelerating Experiment Velocity with Variance Reduction Lenny asks about accelerating test velocity, leading Ronny to detail variance reduction techniques like metric capping and CUPED before transitioning smoothly through the lightning round book recommendations and closing banter.4:30–9:00 · Guest teaching 7/10 Bing's $100M Ad Headline Experiment Lenny opens with a prompt about surprising A/B test results, prompting Ronny to explain the famous Bing headline experiment in detail. Lenny demonstrates familiarity with the mechanism by summarizing the exact UI shift. Ronny educates the audience and host on how counterintuitive even small wins can be.9:01–13:17 · Guest teaching 6/10 Opening Links in New Tabs and Institutional Memory Lenny brings up a parallel experiment from his time at Airbnb regarding opening search listings in new tabs. Ronny builds on it by detailing the lineage of that exact test dating back to 2008 at MSN and Hotmail, schooling the host on institutional memory and failure rates.13:17–15:36 · Guest teaching 8/10 Baseline Failure Rates Across Tech Companies Lenny asks whether a 92% failure rate is typical across tech companies. Ronny breaks down benchmark failure rates across Microsoft, Bing, Booking, and Google Ads, disabusing PMs of the notion that their success rates are unusually high.15:36–20:44 · Guest teaching 7/10 Codifying UI Patterns with GoodUI and Rules of Thumb Lenny asks about codified repositories of UI patterns, suggesting feeding them into LLMs for PM roadmaps. Ronny introduces GoodUI and his Microsoft rules-of-thumb paper, explaining the mechanics of institutional memory and surprising negative results like the Windows indexer.20:45–24:48 · Guest teaching 6/10 Portfolio Strategy and the Bing Social Search Failure Lenny raises the classic critique that experimentation leads to local maxima and micro-optimizations. Ronny counters with portfolio theory and recounts Bing's 100-person-year failed social search bet, while Lenny reinforces with similar failures at Netflix and Airbnb.24:48–28:00 · Guest teaching 8/10 When Startups Should Start Experimenting Lenny asks for tactical startup rules of thumb on when to begin experimentation. Ronny provides a strict statistical cutoff, explaining why math breaks down below tens of thousands of users and naming 200k users as the threshold for true velocity.28:01–32:42 · Guest teaching 8/10 Designing the Overall Evaluation Criterion (OEC) Lenny prompts Ronny on the Overall Evaluation Criterion (OEC) framework. Ronny explains why single revenue metrics incentivize degradation of UX and shows how constrained optimization and guardrail metrics enforce long-term health.32:43–36:30 · Guest teaching 8/10 Lifetime Value Models and Amazon's Unsubscribe Lesson Lenny asks about methods to capture long-term metrics and lifetime value. Ronny provides a detailed case study from Amazon's email recommendation team where modeling the dollar cost of unsubscribes reversed over half of the team's seemingly positive campaigns.36:31–41:47 · Guest teaching 7/10 The Pitfalls of Radical Redesigns vs. OFAT Lenny notes how radical redesigns almost always hurt metrics and asks how to handle passionate teams wanting a full overhaul. Ronny advocates for One Factor At a Time (OFAT), highlighting Microsoft PM denial where teams claimed they were too smart to fail 50% of the time.41:48–45:38 · Guest teaching 6/10 Sponsor Break: Eppo Following the Eppo ad read, Lenny asks whether radical redesigns are ever worth the risk to escape local maxima. Ronny reiterates the 80% failure rule and firmly rejects shipping flat experiments due to codebase complexity and maintenance overhead.45:41–50:06 · Guest teaching 8/10 Airbnb's Product Direction, COVID Shifts, and Counterfactuals Lenny questions Ronny on Airbnb's recent top-down product shift under Brian Chesky and the turning off of performance marketing during COVID. Ronny pushes back firmly, arguing counterfactually that Airbnb would be significantly bigger today had it maintained rigorous experimentation.50:06–55:26 · Guest teaching 8/10 Experiment Platform Trust and Optimizely's P-Value Flaw Lenny asks why Ronny emphasizes trust over velocity. Ronny delivers a detailed breakdown of Optimizely's early statistical flaws with real-time p-value peeking, showing how false positive rates inflated to 30% and eroded market credibility.55:28–1:00:44 · Guest teaching 8/10 Diagnosing Sample Ratio Mismatches (SRM) Lenny asks how to spot invalid tests and whether incorrect assignment points are the primary culprit. Ronny explains Sample Ratio Mismatch (SRM) diagnostics, citing bots and pipeline issues, and humorously recounts how PMs bypassed warnings until scorecards were struck through with red lines.1:00:45–1:06:20 · Guest teaching 9/10 Twyman's Law: Why Extreme Results Are Usually Wrong Lenny asks for an explanation of Twyman's Law and p-values. Ronny explains why extreme results are almost always flaws and breaks down Bayes' rule, revealing that with low base success rates (like Airbnb's 8%), a p<0.05 result carries a 26% false positive risk.1:06:28–1:10:18 · Guest teaching 7/10 Implementing Experimentation: Build vs. Buy and Team Selection Lenny asks how to start experimenting and how to change anti-testing cultures. Ronny discusses build vs buy decisions, advising practitioners to establish beachheads with fast-shipping teams, and illustrates the danger of poorly defined OECs using a Microsoft support site example.1:10:19–1:12:21 · Guest teaching 8/10 Scaling Experimentation Platforms and Metric Maturity Lenny asks about scaling platforms to avoid one-off experiments. Ronny outlines the crawl-walk-run-fly maturity framework, noting how Bing scaled to 10,000 metrics and contrasting that with Airbnb's reliance on data scientist headcount to make up for platform gaps.1:12:24–1:21:34 · Guest teaching 7/10 Accelerating Experiment Velocity with Variance Reduction Lenny asks about accelerating test velocity, leading Ronny to detail variance reduction techniques like metric capping and CUPED before transitioning smoothly through the lightning round book recommendations and closing banter.4:30–9:00 · Guest disagreement 1/10 Bing's $100M Ad Headline Experiment Lenny opens with a prompt about surprising A/B test results, prompting Ronny to explain the famous Bing headline experiment in detail. Lenny demonstrates familiarity with the mechanism by summarizing the exact UI shift. Ronny educates the audience and host on how counterintuitive even small wins can be.9:01–13:17 · Guest disagreement 1/10 Opening Links in New Tabs and Institutional Memory Lenny brings up a parallel experiment from his time at Airbnb regarding opening search listings in new tabs. Ronny builds on it by detailing the lineage of that exact test dating back to 2008 at MSN and Hotmail, schooling the host on institutional memory and failure rates.13:17–15:36 · Guest disagreement 2/10 Baseline Failure Rates Across Tech Companies Lenny asks whether a 92% failure rate is typical across tech companies. Ronny breaks down benchmark failure rates across Microsoft, Bing, Booking, and Google Ads, disabusing PMs of the notion that their success rates are unusually high.15:36–20:44 · Guest disagreement 1/10 Codifying UI Patterns with GoodUI and Rules of Thumb Lenny asks about codified repositories of UI patterns, suggesting feeding them into LLMs for PM roadmaps. Ronny introduces GoodUI and his Microsoft rules-of-thumb paper, explaining the mechanics of institutional memory and surprising negative results like the Windows indexer.20:45–24:48 · Guest disagreement 2/10 Portfolio Strategy and the Bing Social Search Failure Lenny raises the classic critique that experimentation leads to local maxima and micro-optimizations. Ronny counters with portfolio theory and recounts Bing's 100-person-year failed social search bet, while Lenny reinforces with similar failures at Netflix and Airbnb.24:48–28:00 · Guest disagreement 2/10 When Startups Should Start Experimenting Lenny asks for tactical startup rules of thumb on when to begin experimentation. Ronny provides a strict statistical cutoff, explaining why math breaks down below tens of thousands of users and naming 200k users as the threshold for true velocity.28:01–32:42 · Guest disagreement 2/10 Designing the Overall Evaluation Criterion (OEC) Lenny prompts Ronny on the Overall Evaluation Criterion (OEC) framework. Ronny explains why single revenue metrics incentivize degradation of UX and shows how constrained optimization and guardrail metrics enforce long-term health.32:43–36:30 · Guest disagreement 1/10 Lifetime Value Models and Amazon's Unsubscribe Lesson Lenny asks about methods to capture long-term metrics and lifetime value. Ronny provides a detailed case study from Amazon's email recommendation team where modeling the dollar cost of unsubscribes reversed over half of the team's seemingly positive campaigns.36:31–41:47 · Guest disagreement 3/10 The Pitfalls of Radical Redesigns vs. OFAT Lenny notes how radical redesigns almost always hurt metrics and asks how to handle passionate teams wanting a full overhaul. Ronny advocates for One Factor At a Time (OFAT), highlighting Microsoft PM denial where teams claimed they were too smart to fail 50% of the time.41:48–45:38 · Guest disagreement 2/10 Sponsor Break: Eppo Following the Eppo ad read, Lenny asks whether radical redesigns are ever worth the risk to escape local maxima. Ronny reiterates the 80% failure rule and firmly rejects shipping flat experiments due to codebase complexity and maintenance overhead.45:41–50:06 · Guest disagreement 5/10 Airbnb's Product Direction, COVID Shifts, and Counterfactuals Lenny questions Ronny on Airbnb's recent top-down product shift under Brian Chesky and the turning off of performance marketing during COVID. Ronny pushes back firmly, arguing counterfactually that Airbnb would be significantly bigger today had it maintained rigorous experimentation.50:06–55:26 · Guest disagreement 4/10 Experiment Platform Trust and Optimizely's P-Value Flaw Lenny asks why Ronny emphasizes trust over velocity. Ronny delivers a detailed breakdown of Optimizely's early statistical flaws with real-time p-value peeking, showing how false positive rates inflated to 30% and eroded market credibility.55:28–1:00:44 · Guest disagreement 3/10 Diagnosing Sample Ratio Mismatches (SRM) Lenny asks how to spot invalid tests and whether incorrect assignment points are the primary culprit. Ronny explains Sample Ratio Mismatch (SRM) diagnostics, citing bots and pipeline issues, and humorously recounts how PMs bypassed warnings until scorecards were struck through with red lines.1:00:45–1:06:20 · Guest disagreement 3/10 Twyman's Law: Why Extreme Results Are Usually Wrong Lenny asks for an explanation of Twyman's Law and p-values. Ronny explains why extreme results are almost always flaws and breaks down Bayes' rule, revealing that with low base success rates (like Airbnb's 8%), a p<0.05 result carries a 26% false positive risk.1:06:28–1:10:18 · Guest disagreement 2/10 Implementing Experimentation: Build vs. Buy and Team Selection Lenny asks how to start experimenting and how to change anti-testing cultures. Ronny discusses build vs buy decisions, advising practitioners to establish beachheads with fast-shipping teams, and illustrates the danger of poorly defined OECs using a Microsoft support site example.1:10:19–1:12:21 · Guest disagreement 2/10 Scaling Experimentation Platforms and Metric Maturity Lenny asks about scaling platforms to avoid one-off experiments. Ronny outlines the crawl-walk-run-fly maturity framework, noting how Bing scaled to 10,000 metrics and contrasting that with Airbnb's reliance on data scientist headcount to make up for platform gaps.1:12:24–1:21:34 · Guest disagreement 1/10 Accelerating Experiment Velocity with Variance Reduction Lenny asks about accelerating test velocity, leading Ronny to detail variance reduction techniques like metric capping and CUPED before transitioning smoothly through the lightning round book recommendations and closing banter.4:30–9:00 · Lenny pushing back 0/10 Bing's $100M Ad Headline Experiment Lenny opens with a prompt about surprising A/B test results, prompting Ronny to explain the famous Bing headline experiment in detail. Lenny demonstrates familiarity with the mechanism by summarizing the exact UI shift. Ronny educates the audience and host on how counterintuitive even small wins can be.9:01–13:17 · Lenny pushing back 0/10 Opening Links in New Tabs and Institutional Memory Lenny brings up a parallel experiment from his time at Airbnb regarding opening search listings in new tabs. Ronny builds on it by detailing the lineage of that exact test dating back to 2008 at MSN and Hotmail, schooling the host on institutional memory and failure rates.13:17–15:36 · Lenny pushing back 0/10 Baseline Failure Rates Across Tech Companies Lenny asks whether a 92% failure rate is typical across tech companies. Ronny breaks down benchmark failure rates across Microsoft, Bing, Booking, and Google Ads, disabusing PMs of the notion that their success rates are unusually high.15:36–20:44 · Lenny pushing back 0/10 Codifying UI Patterns with GoodUI and Rules of Thumb Lenny asks about codified repositories of UI patterns, suggesting feeding them into LLMs for PM roadmaps. Ronny introduces GoodUI and his Microsoft rules-of-thumb paper, explaining the mechanics of institutional memory and surprising negative results like the Windows indexer.20:45–24:48 · Lenny pushing back 1/10 Portfolio Strategy and the Bing Social Search Failure Lenny raises the classic critique that experimentation leads to local maxima and micro-optimizations. Ronny counters with portfolio theory and recounts Bing's 100-person-year failed social search bet, while Lenny reinforces with similar failures at Netflix and Airbnb.24:48–28:00 · Lenny pushing back 0/10 When Startups Should Start Experimenting Lenny asks for tactical startup rules of thumb on when to begin experimentation. Ronny provides a strict statistical cutoff, explaining why math breaks down below tens of thousands of users and naming 200k users as the threshold for true velocity.28:01–32:42 · Lenny pushing back 0/10 Designing the Overall Evaluation Criterion (OEC) Lenny prompts Ronny on the Overall Evaluation Criterion (OEC) framework. Ronny explains why single revenue metrics incentivize degradation of UX and shows how constrained optimization and guardrail metrics enforce long-term health.32:43–36:30 · Lenny pushing back 0/10 Lifetime Value Models and Amazon's Unsubscribe Lesson Lenny asks about methods to capture long-term metrics and lifetime value. Ronny provides a detailed case study from Amazon's email recommendation team where modeling the dollar cost of unsubscribes reversed over half of the team's seemingly positive campaigns.36:31–41:47 · Lenny pushing back 2/10 The Pitfalls of Radical Redesigns vs. OFAT Lenny notes how radical redesigns almost always hurt metrics and asks how to handle passionate teams wanting a full overhaul. Ronny advocates for One Factor At a Time (OFAT), highlighting Microsoft PM denial where teams claimed they were too smart to fail 50% of the time.41:48–45:38 · Lenny pushing back 1/10 Sponsor Break: Eppo Following the Eppo ad read, Lenny asks whether radical redesigns are ever worth the risk to escape local maxima. Ronny reiterates the 80% failure rule and firmly rejects shipping flat experiments due to codebase complexity and maintenance overhead.45:41–50:06 · Lenny pushing back 2/10 Airbnb's Product Direction, COVID Shifts, and Counterfactuals Lenny questions Ronny on Airbnb's recent top-down product shift under Brian Chesky and the turning off of performance marketing during COVID. Ronny pushes back firmly, arguing counterfactually that Airbnb would be significantly bigger today had it maintained rigorous experimentation.50:06–55:26 · Lenny pushing back 0/10 Experiment Platform Trust and Optimizely's P-Value Flaw Lenny asks why Ronny emphasizes trust over velocity. Ronny delivers a detailed breakdown of Optimizely's early statistical flaws with real-time p-value peeking, showing how false positive rates inflated to 30% and eroded market credibility.55:28–1:00:44 · Lenny pushing back 0/10 Diagnosing Sample Ratio Mismatches (SRM) Lenny asks how to spot invalid tests and whether incorrect assignment points are the primary culprit. Ronny explains Sample Ratio Mismatch (SRM) diagnostics, citing bots and pipeline issues, and humorously recounts how PMs bypassed warnings until scorecards were struck through with red lines.1:00:45–1:06:20 · Lenny pushing back 0/10 Twyman's Law: Why Extreme Results Are Usually Wrong Lenny asks for an explanation of Twyman's Law and p-values. Ronny explains why extreme results are almost always flaws and breaks down Bayes' rule, revealing that with low base success rates (like Airbnb's 8%), a p<0.05 result carries a 26% false positive risk.1:06:28–1:10:18 · Lenny pushing back 0/10 Implementing Experimentation: Build vs. Buy and Team Selection Lenny asks how to start experimenting and how to change anti-testing cultures. Ronny discusses build vs buy decisions, advising practitioners to establish beachheads with fast-shipping teams, and illustrates the danger of poorly defined OECs using a Microsoft support site example.1:10:19–1:12:21 · Lenny pushing back 0/10 Scaling Experimentation Platforms and Metric Maturity Lenny asks about scaling platforms to avoid one-off experiments. Ronny outlines the crawl-walk-run-fly maturity framework, noting how Bing scaled to 10,000 metrics and contrasting that with Airbnb's reliance on data scientist headcount to make up for platform gaps.1:12:24–1:21:34 · Lenny pushing back 0/10 Accelerating Experiment Velocity with Variance Reduction Lenny asks about accelerating test velocity, leading Ronny to detail variance reduction techniques like metric capping and CUPED before transitioning smoothly through the lightning round book recommendations and closing banter.

speaking balance: gold is Lenny, purple is the guest (3 minute bins)

0:00 · Lenny 65.2% · guest 34.8%0:00 · Lenny 65.2% · guest 34.8%3:00 · Lenny 68.3% · guest 31.7%3:00 · Lenny 68.3% · guest 31.7%6:00 · Lenny 2.4% · guest 97.6%6:00 · Lenny 2.4% · guest 97.6%9:00 · Lenny 21.9% · guest 78.1%9:00 · Lenny 21.9% · guest 78.1%12:00 · Lenny 8.1% · guest 91.9%12:00 · Lenny 8.1% · guest 91.9%15:00 · Lenny 13.3% · guest 86.7%15:00 · Lenny 13.3% · guest 86.7%18:00 · Lenny 14.1% · guest 85.9%18:00 · Lenny 14.1% · guest 85.9%21:00 · Lenny 4.1% · guest 95.9%21:00 · Lenny 4.1% · guest 95.9%24:00 · Lenny 23% · guest 77%24:00 · Lenny 23% · guest 77%27:00 · Lenny 7.6% · guest 92.4%27:00 · Lenny 7.6% · guest 92.4%30:00 · Lenny 15.4% · guest 84.6%30:00 · Lenny 15.4% · guest 84.6%33:00 · Lenny 0% · guest 100%33:00 · Lenny 0% · guest 100%36:00 · Lenny 30.7% · guest 69.3%36:00 · Lenny 30.7% · guest 69.3%39:00 · Lenny 21.1% · guest 78.9%39:00 · Lenny 21.1% · guest 78.9%42:00 · Lenny 42% · guest 58%42:00 · Lenny 42% · guest 58%45:00 · Lenny 41.1% · guest 58.9%45:00 · Lenny 41.1% · guest 58.9%48:00 · Lenny 29.7% · guest 70.3%48:00 · Lenny 29.7% · guest 70.3%51:00 · Lenny 8.9% · guest 91.1%51:00 · Lenny 8.9% · guest 91.1%54:00 · Lenny 11.5% · guest 88.5%54:00 · Lenny 11.5% · guest 88.5%57:00 · Lenny 4.1% · guest 95.9%57:00 · Lenny 4.1% · guest 95.9%1:00:00 · Lenny 11.8% · guest 88.2%1:00:00 · Lenny 11.8% · guest 88.2%1:03:00 · Lenny 9% · guest 91%1:03:00 · Lenny 9% · guest 91%1:06:00 · Lenny 16.2% · guest 83.8%1:06:00 · Lenny 16.2% · guest 83.8%1:09:00 · Lenny 7.4% · guest 92.6%1:09:00 · Lenny 7.4% · guest 92.6%1:12:00 · Lenny 15.8% · guest 84.2%1:12:00 · Lenny 15.8% · guest 84.2%1:15:00 · Lenny 27.9% · guest 72.1%1:15:00 · Lenny 27.9% · guest 72.1%1:18:00 · Lenny 11.7% · guest 88.3%1:18:00 · Lenny 11.7% · guest 88.3%1:21:00 · Lenny 40.5% · guest 59.5%1:21:00 · Lenny 40.5% · guest 59.5%
Sharpest disagreement ▶ 46:15 Rejecting Airbnb's top-down narrative

Ronny directly disputes the host's premise that Airbnb's top-down, anti-experimentation shift was validated by business success, arguing counterfactually that the company would be significantly larger had it continued rigorous testing.

Hardest push from Lenny ▶ 38:42 Pushing for large redesigns to escape local maxima

Lenny challenges Ronny's strict incrementalism, arguing that product teams cannot always iterate one factor at a time and sometimes must take radical swings to break out of local maxima.

Biggest teaching moment ▶ 1:04:30 Deconstructing p-values and false positive risk

Ronny dismantles the common misunderstanding of p-values, using Bayesian reasoning to show how an 8% prior success rate turns a p<0.05 result into a startling 26% false positive risk.

Lenny holds their own ▶ 9:01 Lenny citing Airbnb's new-tab search win

Lenny demonstrates his deep domain experience by introducing the exact search listing new-tab experiment from his tenure at Airbnb, matching Ronny's enterprise testing examples with a concrete insider case study.

the scores for every segment, with the reasoning behind each
ChapterTopicLenny as informed peerGuest teachingGuest disagreementLenny pushing backWhy
Bing's $100M Ad Headline Experiment 4710 Lenny opens with a prompt about surprising A/B test results, prompting Ronny to explain the famous Bing headline experiment in detail. Lenny demonstrates familiarity with the mechanism by summarizing the exact UI shift. Ronny educates the audience and host on how counterintuitive even small wins can be.
Opening Links in New Tabs and Institutional Memory 6610 Lenny brings up a parallel experiment from his time at Airbnb regarding opening search listings in new tabs. Ronny builds on it by detailing the lineage of that exact test dating back to 2008 at MSN and Hotmail, schooling the host on institutional memory and failure rates.
Baseline Failure Rates Across Tech Companies 4820 Lenny asks whether a 92% failure rate is typical across tech companies. Ronny breaks down benchmark failure rates across Microsoft, Bing, Booking, and Google Ads, disabusing PMs of the notion that their success rates are unusually high.
Codifying UI Patterns with GoodUI and Rules of Thumb 5710 Lenny asks about codified repositories of UI patterns, suggesting feeding them into LLMs for PM roadmaps. Ronny introduces GoodUI and his Microsoft rules-of-thumb paper, explaining the mechanics of institutional memory and surprising negative results like the Windows indexer.
Portfolio Strategy and the Bing Social Search Failure 5621 Lenny raises the classic critique that experimentation leads to local maxima and micro-optimizations. Ronny counters with portfolio theory and recounts Bing's 100-person-year failed social search bet, while Lenny reinforces with similar failures at Netflix and Airbnb.
When Startups Should Start Experimenting 4820 Lenny asks for tactical startup rules of thumb on when to begin experimentation. Ronny provides a strict statistical cutoff, explaining why math breaks down below tens of thousands of users and naming 200k users as the threshold for true velocity.
Designing the Overall Evaluation Criterion (OEC) 5820 Lenny prompts Ronny on the Overall Evaluation Criterion (OEC) framework. Ronny explains why single revenue metrics incentivize degradation of UX and shows how constrained optimization and guardrail metrics enforce long-term health.
Lifetime Value Models and Amazon's Unsubscribe Lesson 4810 Lenny asks about methods to capture long-term metrics and lifetime value. Ronny provides a detailed case study from Amazon's email recommendation team where modeling the dollar cost of unsubscribes reversed over half of the team's seemingly positive campaigns.
The Pitfalls of Radical Redesigns vs. OFAT 5732 Lenny notes how radical redesigns almost always hurt metrics and asks how to handle passionate teams wanting a full overhaul. Ronny advocates for One Factor At a Time (OFAT), highlighting Microsoft PM denial where teams claimed they were too smart to fail 50% of the time.
Sponsor Break: Eppo 4621 Following the Eppo ad read, Lenny asks whether radical redesigns are ever worth the risk to escape local maxima. Ronny reiterates the 80% failure rule and firmly rejects shipping flat experiments due to codebase complexity and maintenance overhead.
Airbnb's Product Direction, COVID Shifts, and Counterfactuals 6852 Lenny questions Ronny on Airbnb's recent top-down product shift under Brian Chesky and the turning off of performance marketing during COVID. Ronny pushes back firmly, arguing counterfactually that Airbnb would be significantly bigger today had it maintained rigorous experimentation.
Experiment Platform Trust and Optimizely's P-Value Flaw 4840 Lenny asks why Ronny emphasizes trust over velocity. Ronny delivers a detailed breakdown of Optimizely's early statistical flaws with real-time p-value peeking, showing how false positive rates inflated to 30% and eroded market credibility.
Diagnosing Sample Ratio Mismatches (SRM) 4830 Lenny asks how to spot invalid tests and whether incorrect assignment points are the primary culprit. Ronny explains Sample Ratio Mismatch (SRM) diagnostics, citing bots and pipeline issues, and humorously recounts how PMs bypassed warnings until scorecards were struck through with red lines.
Twyman's Law: Why Extreme Results Are Usually Wrong 4930 Lenny asks for an explanation of Twyman's Law and p-values. Ronny explains why extreme results are almost always flaws and breaks down Bayes' rule, revealing that with low base success rates (like Airbnb's 8%), a p<0.05 result carries a 26% false positive risk.
Implementing Experimentation: Build vs. Buy and Team Selection 4720 Lenny asks how to start experimenting and how to change anti-testing cultures. Ronny discusses build vs buy decisions, advising practitioners to establish beachheads with fast-shipping teams, and illustrates the danger of poorly defined OECs using a Microsoft support site example.
Scaling Experimentation Platforms and Metric Maturity 3820 Lenny asks about scaling platforms to avoid one-off experiments. Ronny outlines the crawl-walk-run-fly maturity framework, noting how Bing scaled to 10,000 metrics and contrasting that with Airbnb's reliance on data scientist headcount to make up for platform gaps.
Accelerating Experiment Velocity with Variance Reduction 5710 Lenny asks about accelerating test velocity, leading Ronny to detail variance reduction techniques like metric capping and CUPED before transitioning smoothly through the lightning round book recommendations and closing banter.

Statements from this episode (32)

Assertion Supported
Moving Bing ad text to the headline generated $100 million
“This thing was worth a hundred million dollars at the time when Bing was a lot smaller.”
Ronny Kohavi Jul 27, 2023 ▶ 7:17
Insight
Product teams are consistently poor at predicting experiment outcomes
“We are often humbled by how bad we are at predicting the outcome of experiments.”
Ronny Kohavi Jul 27, 2023 ▶ 8:53
Assertion Not checkable as stated
Bing's search relevance team targets 2% annual compounding improvement
“One is at Bing, the relevance team. Hundreds of people all working to improve Bing relevance. They have a metric. We'll talk about, oh, we see the overall evaluation criterion, but they have a metric that their goal is to improve it by two percent every year. …”
Ronny Kohavi Jul 27, 2023 ▶ 12:00
Assertion Not checkable as stated
Airbnb's search team drove a 6% revenue gain across 250 experiments
“Another example that I am allowed to speak about from Airbnb Is the fact that we ran some 250 experiments in my tenure there in search relevance. And again, small improvements added up. So this became overall a six percent improvement to revenue.”
Ronny Kohavi Jul 27, 2023 ▶ 12:28
Assertion Supported
Microsoft's overall experiment failure rate is 66%, reaching 85% at Bing
“So overall at Microsoft, About 66%, two-thirds of ideas fail, right? And don't think the 66 is accurate. Like, you know, it's about two-thirds. At Bing, which is a much more optimized domain after we've been optimizing it for a while, the failure rate was arou…”
Ronny Kohavi Jul 27, 2023 ▶ 13:36
Assertion Supported
Airbnb's experiment failure rate was 92%, while Google Ads hit 80-90%
“And then at Airbnb this 92% number is You know, the highest failure rate that I've observed. Now I've quoted other sources that, you know, it's not that I worked at groups that were particularly bad booking Google ads, other companies published numbers. There …”
Ronny Kohavi Jul 27, 2023 ▶ 14:03
Insight
Approximately 10% of experiments are aborted on day one due to bugs
“In fact, 10% of experiments tend to be aborted on the first date. Those are usually not that the idea is bad, but that there is an implementation issue or something we haven't thought about that forces an abort.”
Ronny Kohavi Jul 27, 2023 ▶ 14:48
Assertion Supported
GoodUI.org tracks roughly 140 crowdsourced experiment patterns and win rates
“There's probably like a 140 patterns, I think at this point. And then for each pattern, he says well, who hasn't helped? How many times and by how much? So you have an idea of, you know, this worked three out of five times and it was a huge win.”
Ronny Kohavi Jul 27, 2023 ▶ 16:36
Assertion Supported
By 2019, Microsoft launched roughly 100 new experimental treatments daily
“At Microsoft, just to let you know, when I left in 2019, we were on a rate of about 20 to 25,000 experiments every year. So every working day, we were starting something like a hundred new treatments.”
Ronny Kohavi Jul 27, 2023 ▶ 20:02
Assertion Supported
Bing wasted 100 person-years integrating social feeds before shutting it down
“We hooked into the Twitter Firehose feed, and we hooked into Facebook, and we spent a hundred person years on this idea, and it failed. You don't see it anymore. It existed for about a year and a half, and all the experiments were just negative to flat.”
Ronny Kohavi Jul 27, 2023 ▶ 23:02
Assertion Not checkable as stated
An early Airbnb feature showing friends' stays had no measurable impact
“At Airbnb early on. There was a big social attempt to make like, here's your friends have stayed at these Airbnb's completely not, had no impact.”
Lenny Rachitsky Jul 27, 2023 ▶ 24:11
Insight
A/B testing statistics generally fail without tens of thousands of users
“Unless you have at least tens of thousands of users, The math, the statistics just don't work out for most of the metrics that you're interested in.”
Ronny Kohavi Jul 27, 2023 ▶ 26:52
Insight
Comprehensive A/B testing requires a minimum of approximately 200,000 users
“So you ask for rule of thumb, 200,000 users, you're magical. Below that, start building the culture, start building the platform, start integrating, so that as you scale, you start to see the value.”
Ronny Kohavi Jul 27, 2023 ▶ 27:25
Insight
An experiment's primary metric must causally predict customer lifetime value
“To me, the key here, the key word is lifetime value, which is you have to define the OEC such that it is causally predictive of the lifetime value of the user.”
Ronny Kohavi Jul 27, 2023 ▶ 32:05
Assertion Not checkable as stated
Over half of Amazon's recommendation emails had negative ROI due to unsubscribes
“And then when we started to incorporate this formula, more than half the campaigns that were being sent were negative.”
Ronny Kohavi Jul 27, 2023 ▶ 35:05
Insight
Bundling multiple redesign changes is more likely to fail than incremental testing
“If you believe in that statistics that I published, then doing 17 changes together is more likely to be negative. Do them in smaller increments. Learn from, it's called OFAT, one factor at a time. Do one factor, learn from it, and adjust. Of the 17, maybe you …”
Ronny Kohavi Jul 27, 2023 ▶ 37:52
Insight
Major product redesigns fail approximately 80% of the time
“80% of the time you will fail. So be ready for that. Right. What people usually expect is my redesign is going to work. No, you're most likely going to fail, but if you do succeed, it's a breakthrough.”
Ronny Kohavi Jul 27, 2023 ▶ 43:20
Insight
Never ship flat experiment results due to the hidden maintenance overhead
“Flat to me, if something is not Statsig, that's a no ship because you've just introduced more code. There is a maintenance overhead. To shipping your stuff. I've heard people say, look, we already spent all this time. The team will be demotivated if we don't s…”
Ronny Kohavi Jul 27, 2023 ▶ 44:29
Assertion Not publicly verifiable
Airbnb tested its entire rebrand and homepage redesign with a holdout group
“When Airbnb launched the rebrand, even that they ran as an experiment with the entire homepage redesign, the new logo and all that. And I think there was a long-term holdout even, and I think it was positive in the end from what I remember.”
Lenny Rachitsky Jul 27, 2023 ▶ 45:27
Disclosure
Airbnb's search relevance team launched nothing without an A/B test
“One is in my team in search relevance, everything was A-B tested. So while Brian can focus on some of the design aspects, the people who are actually doing, you know, the neural networks and the search, Everything was AB tested to help. So nothing was launchin…”
Ronny Kohavi Jul 27, 2023 ▶ 46:23
What-if
Kohavi: Airbnb would be in a better state today with more experiments
“I believe that had Airbnb kept people like Greg really, which was pushing for a lot more data driven and had Airbnb run more experiments. It would have been in a better state than today, but it's the counterfactual. We don't know.”
Ronny Kohavi Jul 27, 2023 ▶ 47:00
What-if
Kohavi: Airbnb's revenue would have grown just as much without COVID pivots
“I think if Airbnb stayed the course, did nothing, the revenue would have gone up in the same way.”
Ronny Kohavi Jul 27, 2023 ▶ 49:41
Opinion
Airbnb's Online Experiences had poor initial data and became a footnote
“In fact, if you look at one investment, one big investment that was done at the time was online experiences, and the initial data wasn't very promising, and I think today it's a footnote.”
Ronny Kohavi Jul 27, 2023 ▶ 49:49
Assertion Partly supported
Roughly 8% of Microsoft's online experiments had sample ratio mismatches
“And where we share that at Microsoft, even though we'd be running experiments for a while is around eight percent of experiments that suffered from the sample ratio of dispatch.”
Ronny Kohavi Jul 27, 2023 ▶ 57:19
Opinion
Bots are the most common cause of sample ratio mismatches in experiments
“So when you say most common, I think the most common is bots.”
Ronny Kohavi Jul 27, 2023 ▶ 58:40
Assertion Not checkable as stated
Nine times out of ten, surprising A/B test wins contain hidden flaws
“And I will say that nine out of 10, when we call out Twyman's Law, it is the case that we find some flaw in the experiment.”
Ronny Kohavi Jul 27, 2023 ▶ 1:01:39
Assertion Supported
A statistically significant Airbnb search result still carried a 26% false-positive risk
“If you're at Airbnb where the success rate of, or Airbnb search where the success rate is only eight percent, if you get a statistically significant result with a p-value less than point oh five, there is a 26% chance that this is a false positive result. Righ…”
Ronny Kohavi Jul 27, 2023 ▶ 1:04:35
Disclosure
Airbnb required teams to replicate A/B tests with borderline p-values
“When I worked at Airbnb, one of the things we did is we said, okay, if you're less than .5, but above .1, rerun, replicate. When you replicate, you can combine the two experiments and get a combined p-value using something called Fisher's method or Stauffer's …”
Ronny Kohavi Jul 27, 2023 ▶ 1:04:57
Assertion Supported
Bing allowed experimenters to define up to 10,000 metrics per scorecard
“At Bing, you could define 10,000 metrics that you wanted to be in your scorecard.”
Ronny Kohavi Jul 27, 2023 ▶ 1:10:55
Opinion
Airbnb's weak experimentation platform forced heavy hiring of data scientists
“I think, you know, one of the things that I will say at Airbnb is the analysis was relatively weak. And so lots of data scientists were hired to be able to compensate for the fact that the platform didn't do enough.”
Ronny Kohavi Jul 27, 2023 ▶ 1:11:25
Insight
Automated scorecards are the primary way to accelerate experimentation speed
“One is if your platform is good, then when the experiment finishes, you should have a scorecard soon after. Maybe it takes a day, but it shouldn't be that you have to wait a week for the data scientists. To me, this is the number one way to speed up things.”
Ronny Kohavi Jul 27, 2023 ▶ 1:12:42
Assertion Supported
The CUPED technique reduces metric variance and sample size requirements without bias
“A third technique is called Cupid, which is an article that we published. Again, I can give it in the notes, which uses the pre-experiment data to adjust the result. And we can show that you get the result as unbiased, but with lower variance in health, hence …”
Ronny Kohavi Jul 27, 2023 ▶ 1:13:53
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.