Everything Ronny Kohavi said on any show that made the record, most notable first. Each card names its show and opens the statement there.
Kohavi: Airbnb would be in a better state today with more experiments
“I believe that had Airbnb kept people like Greg really, which was pushing for a lot more data driven and had Airbnb run more experiments. It would have been in a better state than today, but it's the counterfactual. We don't know.”
Kohavi: Airbnb's revenue would have grown just as much without COVID pivots
“I think if Airbnb stayed the course, did nothing, the revenue would have gone up in the same way.”
Never ship flat experiment results due to the hidden maintenance overhead
“Flat to me, if something is not Statsig, that's a no ship because you've just introduced more code. There is a maintenance overhead. To shipping your stuff. I've heard people say, look, we already spent all this time. The team will be demotivated if we don't s…”
Airbnb's weak experimentation platform forced heavy hiring of data scientists
“I think, you know, one of the things that I will say at Airbnb is the analysis was relatively weak. And so lots of data scientists were hired to be able to compensate for the fact that the platform didn't do enough.”
Microsoft's overall experiment failure rate is 66%, reaching 85% at Bing
“So overall at Microsoft, About 66%, two-thirds of ideas fail, right? And don't think the 66 is accurate. Like, you know, it's about two-thirds. At Bing, which is a much more optimized domain after we've been optimizing it for a while, the failure rate was arou…”
Airbnb's experiment failure rate was 92%, while Google Ads hit 80-90%
“And then at Airbnb this 92% number is You know, the highest failure rate that I've observed. Now I've quoted other sources that, you know, it's not that I worked at groups that were particularly bad booking Google ads, other companies published numbers. There …”
A/B testing statistics generally fail without tens of thousands of users
“Unless you have at least tens of thousands of users, The math, the statistics just don't work out for most of the metrics that you're interested in.”
Comprehensive A/B testing requires a minimum of approximately 200,000 users
“So you ask for rule of thumb, 200,000 users, you're magical. Below that, start building the culture, start building the platform, start integrating, so that as you scale, you start to see the value.”
An experiment's primary metric must causally predict customer lifetime value
“To me, the key here, the key word is lifetime value, which is you have to define the OEC such that it is causally predictive of the lifetime value of the user.”
Airbnb's Online Experiences had poor initial data and became a footnote
“In fact, if you look at one investment, one big investment that was done at the time was online experiences, and the initial data wasn't very promising, and I think today it's a footnote.”
Moving Bing ad text to the headline generated $100 million
“This thing was worth a hundred million dollars at the time when Bing was a lot smaller.”
Product teams are consistently poor at predicting experiment outcomes
“We are often humbled by how bad we are at predicting the outcome of experiments.”
Airbnb's search team drove a 6% revenue gain across 250 experiments
“Another example that I am allowed to speak about from Airbnb Is the fact that we ran some 250 experiments in my tenure there in search relevance. And again, small improvements added up. So this became overall a six percent improvement to revenue.”
Bing wasted 100 person-years integrating social feeds before shutting it down
“We hooked into the Twitter Firehose feed, and we hooked into Facebook, and we spent a hundred person years on this idea, and it failed. You don't see it anymore. It existed for about a year and a half, and all the experiments were just negative to flat.”
Over half of Amazon's recommendation emails had negative ROI due to unsubscribes
“And then when we started to incorporate this formula, more than half the campaigns that were being sent were negative.”
Bundling multiple redesign changes is more likely to fail than incremental testing
“If you believe in that statistics that I published, then doing 17 changes together is more likely to be negative. Do them in smaller increments. Learn from, it's called OFAT, one factor at a time. Do one factor, learn from it, and adjust. Of the 17, maybe you …”
Major product redesigns fail approximately 80% of the time
“80% of the time you will fail. So be ready for that. Right. What people usually expect is my redesign is going to work. No, you're most likely going to fail, but if you do succeed, it's a breakthrough.”
Airbnb's search relevance team launched nothing without an A/B test
“One is in my team in search relevance, everything was A-B tested. So while Brian can focus on some of the design aspects, the people who are actually doing, you know, the neural networks and the search, Everything was AB tested to help. So nothing was launchin…”
Roughly 8% of Microsoft's online experiments had sample ratio mismatches
“And where we share that at Microsoft, even though we'd be running experiments for a while is around eight percent of experiments that suffered from the sample ratio of dispatch.”
Bots are the most common cause of sample ratio mismatches in experiments
“So when you say most common, I think the most common is bots.”
Nine times out of ten, surprising A/B test wins contain hidden flaws
“And I will say that nine out of 10, when we call out Twyman's Law, it is the case that we find some flaw in the experiment.”
A statistically significant Airbnb search result still carried a 26% false-positive risk
“If you're at Airbnb where the success rate of, or Airbnb search where the success rate is only eight percent, if you get a statistically significant result with a p-value less than point oh five, there is a 26% chance that this is a false positive result. Righ…”
Automated scorecards are the primary way to accelerate experimentation speed
“One is if your platform is good, then when the experiment finishes, you should have a scorecard soon after. Maybe it takes a day, but it shouldn't be that you have to wait a week for the data scientists. To me, this is the number one way to speed up things.”
Bing's search relevance team targets 2% annual compounding improvement
“One is at Bing, the relevance team. Hundreds of people all working to improve Bing relevance. They have a metric. We'll talk about, oh, we see the overall evaluation criterion, but they have a metric that their goal is to improve it by two percent every year. …”