What-if
Kohavi: Airbnb would be in a better state today with more experiments
“I believe that had Airbnb kept people like Greg really, which was pushing for a lot more data driven and had Airbnb run more experiments. It would have been in a better state than today, but it's the counterfactual. We don't know.”
What-if
Kohavi: Airbnb's revenue would have grown just as much without COVID pivots
“I think if Airbnb stayed the course, did nothing, the revenue would have gone up in the same way.”
Insight
Never ship flat experiment results due to the hidden maintenance overhead
“Flat to me, if something is not Statsig, that's a no ship because you've just introduced more code. There is a maintenance overhead. To shipping your stuff. I've heard people say, look, we already spent all this time. The team will be demotivated if we don't s…”
Opinion
Airbnb's weak experimentation platform forced heavy hiring of data scientists
“I think, you know, one of the things that I will say at Airbnb is the analysis was relatively weak. And so lots of data scientists were hired to be able to compensate for the fact that the platform didn't do enough.”
Assertion Supported
Microsoft's overall experiment failure rate is 66%, reaching 85% at Bing
“So overall at Microsoft, About 66%, two-thirds of ideas fail, right? And don't think the 66 is accurate. Like, you know, it's about two-thirds. At Bing, which is a much more optimized domain after we've been optimizing it for a while, the failure rate was arou…”
Assertion Supported
Airbnb's experiment failure rate was 92%, while Google Ads hit 80-90%
“And then at Airbnb this 92% number is You know, the highest failure rate that I've observed. Now I've quoted other sources that, you know, it's not that I worked at groups that were particularly bad booking Google ads, other companies published numbers. There …”
Insight
A/B testing statistics generally fail without tens of thousands of users
“Unless you have at least tens of thousands of users, The math, the statistics just don't work out for most of the metrics that you're interested in.”
Insight
Comprehensive A/B testing requires a minimum of approximately 200,000 users
“So you ask for rule of thumb, 200,000 users, you're magical. Below that, start building the culture, start building the platform, start integrating, so that as you scale, you start to see the value.”
Insight
An experiment's primary metric must causally predict customer lifetime value
“To me, the key here, the key word is lifetime value, which is you have to define the OEC such that it is causally predictive of the lifetime value of the user.”
Opinion
Airbnb's Online Experiences had poor initial data and became a footnote
“In fact, if you look at one investment, one big investment that was done at the time was online experiences, and the initial data wasn't very promising, and I think today it's a footnote.”
Assertion Supported
Moving Bing ad text to the headline generated $100 million
“This thing was worth a hundred million dollars at the time when Bing was a lot smaller.”
Insight
Product teams are consistently poor at predicting experiment outcomes
“We are often humbled by how bad we are at predicting the outcome of experiments.”
Assertion Not checkable as stated
Airbnb's search team drove a 6% revenue gain across 250 experiments
“Another example that I am allowed to speak about from Airbnb Is the fact that we ran some 250 experiments in my tenure there in search relevance. And again, small improvements added up. So this became overall a six percent improvement to revenue.”
Assertion Supported
Bing wasted 100 person-years integrating social feeds before shutting it down
“We hooked into the Twitter Firehose feed, and we hooked into Facebook, and we spent a hundred person years on this idea, and it failed. You don't see it anymore. It existed for about a year and a half, and all the experiments were just negative to flat.”
Assertion Not checkable as stated
Over half of Amazon's recommendation emails had negative ROI due to unsubscribes
“And then when we started to incorporate this formula, more than half the campaigns that were being sent were negative.”
Insight
Bundling multiple redesign changes is more likely to fail than incremental testing
“If you believe in that statistics that I published, then doing 17 changes together is more likely to be negative. Do them in smaller increments. Learn from, it's called OFAT, one factor at a time. Do one factor, learn from it, and adjust. Of the 17, maybe you …”
Insight
Major product redesigns fail approximately 80% of the time
“80% of the time you will fail. So be ready for that. Right. What people usually expect is my redesign is going to work. No, you're most likely going to fail, but if you do succeed, it's a breakthrough.”
Disclosure
Airbnb's search relevance team launched nothing without an A/B test
“One is in my team in search relevance, everything was A-B tested. So while Brian can focus on some of the design aspects, the people who are actually doing, you know, the neural networks and the search, Everything was AB tested to help. So nothing was launchin…”
Assertion Partly supported
Roughly 8% of Microsoft's online experiments had sample ratio mismatches
“And where we share that at Microsoft, even though we'd be running experiments for a while is around eight percent of experiments that suffered from the sample ratio of dispatch.”
Opinion
Bots are the most common cause of sample ratio mismatches in experiments
“So when you say most common, I think the most common is bots.”
Assertion Not checkable as stated
Nine times out of ten, surprising A/B test wins contain hidden flaws
“And I will say that nine out of 10, when we call out Twyman's Law, it is the case that we find some flaw in the experiment.”
Assertion Supported
A statistically significant Airbnb search result still carried a 26% false-positive risk
“If you're at Airbnb where the success rate of, or Airbnb search where the success rate is only eight percent, if you get a statistically significant result with a p-value less than point oh five, there is a 26% chance that this is a false positive result. Righ…”
Insight
Automated scorecards are the primary way to accelerate experimentation speed
“One is if your platform is good, then when the experiment finishes, you should have a scorecard soon after. Maybe it takes a day, but it shouldn't be that you have to wait a week for the data scientists. To me, this is the number one way to speed up things.”
Assertion Not checkable as stated
Bing's search relevance team targets 2% annual compounding improvement
“One is at Bing, the relevance team. Hundreds of people all working to improve Bing relevance. They have a metric. We'll talk about, oh, we see the overall evaluation criterion, but they have a metric that their goal is to improve it by two percent every year. …”