Data Pipeline

topic on 4 shows · 8 statements across 8 episodes

Innovators & Investors Latent Space the MAD Podcast 20VC

8 statements about Data Pipeline, every show

Ethan He: Pipeline bug fixes drive more model gains than new algorithms
“And often I find that this is kind of boring, but like a lot of the improvements does not come from new algorithms. It comes from finding small bugs here and there in the data pipeline, in the model training pipeline. Those gave the biggest boost to the model …”
Ethan He Jun 1, 2026 ▶ 7:40 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
20VC Assertion Not checkable as stated
Feldman: Many AI projects fail due to bad data, not AI
“I think most, many AI projects fail on those fronts. Have nothing to do with the AI. They fail because the data was a disaster.”
Andrew Feldman Oct 6, 2025 ▶ 1:10:17 Cerebras CEO, Andrew Feldman on Why Raise $1BN and Delay the IPO & Why NVIDIA’s Worried About Growth · 20VC with Harry Stebbings
El-Moursi: Data preparation and pipelines are the key prerequisite for AI agents
“Because once you actually solve for that on-ramp of how to get your data on, how to cleanse it, how to protect it, how to make it usable and create the data pipelines then the agentic piece, you can ask any question, and that's what you're engaging with today …”
Suzanne El-Moursi May 28, 2025 ▶ 11:51 How BrightHive is Unlocking the Power of AI to Transform Data-Driven Work
MAD Assertion Not checkable as stated
Ben Rogojan: Data type issues still routinely break data pipelines
“But to this day, people are still having like a data type issue break a pipeline that then causes everyone to have to kind of spend a little bit of time fixing them.”
Ben Rogojan (Seattle Data Guy) Jan 23, 2025 ▶ 12:49 Understanding Data Engineering in 2025 | Ben Rogojan, Seattle Data Guy
MAD Insight
Moses: Validating data at a single pipeline point is no longer sufficient
“And so making sure that your data is accurate at only one point of the pipeline is just no longer sufficient”
Barr Moses Apr 12, 2022 ▶ 9:52 Unlocking Data Observability with Monte Carlo's Barr Moses
MAD Insight
Dada argues building data pipelines is harder than machine learning itself
“And this is where we believe that machine learning is not really the hard part, but building and maintaining the data pipeline is”
Kedro Product Manager Feb 17, 2021 ▶ 3:18 Introducing Kedro
MAD Assertion Not checkable as stated
Sherman: Timber reduced its data pipeline costs by about 90%
“And we actually reduce the cost of our pipeline by about 90% by doing this.”
Zach Sherman Dec 5, 2018 ▶ 9:55 A New Kind of Logging System // Zach Sherman & Ben Johnson, Timber (FirstMark's Data Driven NYC)
MAD Insight
Joffe: Modifying pre-aggregated data cube pipelines takes weeks to months
“You have to go and change the pipeline. For those of you that have been there, done that, know that this is something you stay away from as much as you can, because this takes on the order of weeks to months.”
Nitay Joffe Jun 16, 2016 ▶ 12:56 Why Marketing is All About Data // Nitay Joffe, ActionIQ [FirstMark's Data Driven]

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.