Michael Jordan, computer science and statistics professor at UC Berkeley, explains why classic bootstrap resampling methods cannot scale in modern distributed cloud architecture.
“Okay, but, gotcha, big gotcha, which is you can't do this on a terabyte of data, alright, because each resampling of the original data set on, if you have a terabyte, it's about 632 gigabytes. So you're sitting there on your terabyte of data at a central computer, You sample with replacement a few thousands of times, that's what you need to do, and you get these 632 gigabyte data sets, thousands of them, you're sending on your network to get out to your machines. Hopeless. Way too slow. Alright, so our big principle in statistics, which is, really is the bootstrap of how to get error bars in some generality without priors, just can't, it doesn't scale.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Michael Jordan
PredictionNot checkable as stated
Jordan predicts principled personalized big data systems remain decades away
“So I think we're decades away from being able to do what this boss is asking us to do in some principle way. You can occasionally build a one-off system that does some of these things, but we're decades from having the real principles.”
Jordan says personalization business models fail due to statistical limits
“A lot of these business models are failing. People actually can't personalize very well, and it's Because of statistical issues. You've got huge amounts of data about some people, and very little about lots of people, and you don't know how to transfer the sta…”
Jordan warns high-dimensional Bayesian inference is overly sensitive to unknown priors
“And a lot of times you have no idea what the prior should be. You don't know what the tails should be in particular. And you're in high dimensions, you really have no idea how the tail behavior should be. And the whole inference is highly sensitive to the tail…”
This entire site, over 1,000 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.