Ronny Kohavi

14 statements across 1 episodes · 4 bullish · 1 bearish · 1 people on the record · first statement Jul 27, 2023 by Ronny Kohavi · said 3 times in 3 episodes since 2023 · across every show →

On the record as a speaker too: Ronny Kohavi's record, appearances and statements → this page counts the times other people say the name.

Mentions by year

brought up most by Ramesh Johari (1), Lenny Rachitsky (1), Itamar Gilad (1)

tap a year for its mentions
0022332023episodesmentions
0232023episodes it came up in
000.51.5132023episodesmentions per episode

every mention, scene by scene, with the transcript →

Everything said about Ronny Kohavi, oldest first

Jul 27, 2023 neutral
Insight
Major product redesigns fail approximately 80% of the time
“80% of the time you will fail. So be ready for that. Right. What people usually expect is my redesign is going to work. No, you're most likely going to fail, but if you do succeed, it's a breakthrough.”
Ronny Kohavi Jul 27, 2023 ▶ 43:20 The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon)
Jul 27, 2023 positive
Assertion Not checkable as stated
Airbnb's search team drove a 6% revenue gain across 250 experiments
“Another example that I am allowed to speak about from Airbnb Is the fact that we ran some 250 experiments in my tenure there in search relevance. And again, small improvements added up. So this became overall a six percent improvement to revenue.”
Ronny Kohavi Jul 27, 2023 ▶ 12:28 The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon)
Jul 27, 2023 neutral
Insight
A/B testing statistics generally fail without tens of thousands of users
“Unless you have at least tens of thousands of users, The math, the statistics just don't work out for most of the metrics that you're interested in.”
Ronny Kohavi Jul 27, 2023 ▶ 26:52 The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon)
Jul 27, 2023 positive
Insight
An experiment's primary metric must causally predict customer lifetime value
“To me, the key here, the key word is lifetime value, which is you have to define the OEC such that it is causally predictive of the lifetime value of the user.”
Ronny Kohavi Jul 27, 2023 ▶ 32:05 The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon)
Jul 27, 2023 neutral
Disclosure
Airbnb's search relevance team launched nothing without an A/B test
“One is in my team in search relevance, everything was A-B tested. So while Brian can focus on some of the design aspects, the people who are actually doing, you know, the neural networks and the search, Everything was AB tested to help. So nothing was launchin…”
Ronny Kohavi Jul 27, 2023 ▶ 46:23 The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon)
Jul 27, 2023 neutral
Assertion Supported
By 2019, Microsoft launched roughly 100 new experimental treatments daily
“At Microsoft, just to let you know, when I left in 2019, we were on a rate of about 20 to 25,000 experiments every year. So every working day, we were starting something like a hundred new treatments.”
Ronny Kohavi Jul 27, 2023 ▶ 20:02 The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon)
Jul 27, 2023 neutral
Insight
Comprehensive A/B testing requires a minimum of approximately 200,000 users
“So you ask for rule of thumb, 200,000 users, you're magical. Below that, start building the culture, start building the platform, start integrating, so that as you scale, you start to see the value.”
Ronny Kohavi Jul 27, 2023 ▶ 27:25 The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon)
Jul 27, 2023
Insight
Bundling multiple redesign changes is more likely to fail than incremental testing
“If you believe in that statistics that I published, then doing 17 changes together is more likely to be negative. Do them in smaller increments. Learn from, it's called OFAT, one factor at a time. Do one factor, learn from it, and adjust. Of the 17, maybe you …”
Ronny Kohavi Jul 27, 2023 ▶ 37:52 The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon)
Jul 27, 2023 neutral
Insight
Approximately 10% of experiments are aborted on day one due to bugs
“In fact, 10% of experiments tend to be aborted on the first date. Those are usually not that the idea is bad, but that there is an implementation issue or something we haven't thought about that forces an abort.”
Ronny Kohavi Jul 27, 2023 ▶ 14:48 The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon)
Jul 27, 2023 positive
Assertion Not checkable as stated
Bing's search relevance team targets 2% annual compounding improvement
“One is at Bing, the relevance team. Hundreds of people all working to improve Bing relevance. They have a metric. We'll talk about, oh, we see the overall evaluation criterion, but they have a metric that their goal is to improve it by two percent every year. …”
Ronny Kohavi Jul 27, 2023 ▶ 12:00 The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon)
Jul 27, 2023 positive
Insight
Automated scorecards are the primary way to accelerate experimentation speed
“One is if your platform is good, then when the experiment finishes, you should have a scorecard soon after. Maybe it takes a day, but it shouldn't be that you have to wait a week for the data scientists. To me, this is the number one way to speed up things.”
Ronny Kohavi Jul 27, 2023 ▶ 1:12:42 The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon)
Jul 27, 2023 negative
Insight
Never ship flat experiment results due to the hidden maintenance overhead
“Flat to me, if something is not Statsig, that's a no ship because you've just introduced more code. There is a maintenance overhead. To shipping your stuff. I've heard people say, look, we already spent all this time. The team will be demotivated if we don't s…”
Ronny Kohavi Jul 27, 2023 ▶ 44:29 The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon)
Jul 27, 2023
Disclosure
Airbnb required teams to replicate A/B tests with borderline p-values
“When I worked at Airbnb, one of the things we did is we said, okay, if you're less than .5, but above .1, rerun, replicate. When you replicate, you can combine the two experiments and get a combined p-value using something called Fisher's method or Stauffer's …”
Ronny Kohavi Jul 27, 2023 ▶ 1:04:57 The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon)
Jul 27, 2023 neutral
Insight
Product teams are consistently poor at predicting experiment outcomes
“We are often humbled by how bad we are at predicting the outcome of experiments.”
Ronny Kohavi Jul 27, 2023 ▶ 8:53 The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon)
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.