Stanford Professor Chris Ré describes how DeepDive uses probabilistic inference to identify which parts of a data pipeline need work, yielding massive efficiency gains.
“In contrast, we can build a pipeline first, which may not be very high quality, and then use probabilistic inference as a way to say, where should we spend our effort next? And the reduction in effort can sometimes be orders of magnitude, as we've seen in some of the applications we've built in our lab.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Chris Ré
Opinion
Classical models of sequential evaluation can be completely thrown away
“That is, why it's a new trade-off space and classical trade-off space don't sort of suffice, and to do that, I'm going to show you one result which just illustrates the classical models of sequential evaluation you can completely throw away.”
Relaxing hardware memory consistency yields up to 1000x speedups over competitors
“We've basically done a couple of parlor tricks, and those parlor tricks have allowed us to get, you know, 10, a hundred, a thousand times faster than competitor systems, because we're actually taking advantage of what the hardware is giving us.”
Automated dark data pipelines can outperform human data extraction quality
“There are a number of systems that have shown that it's actually possible to build these ETL pipelines, these pipelines that extract, transform, and load information with higher quality than humans in some very simple settings.”
Traditional data pipelines over-extract and over-clean due to unmeasurable impact
“So in contrast, if you think about the way that people build these extraction and integration and cleaning systems, they tend to over-extract, over-integrate, and over-clean Because they have no idea if their cleaning or extraction is actually gonna prove the …”
Modern scientific literature is universally accessible but impossible for humans to read
“For even the narrowest of scientific questions, They couldn't possibly read all the information that was relevant to them, or even a significant fraction of it. So this is why I say that scientific knowledge is accessible in a way like never before, but it's n…”
This entire site, over 1,000 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.