Data Pipeline
topic on 4 shows · 8 statements across 8 episodes
Innovators & Investors
Latent Space
the MAD Podcast
20VC
8 statements about Data Pipeline, every show
Ethan He: Pipeline bug fixes drive more model gains than new algorithms
“And often I find that this is kind of boring, but like a lot of the improvements does not come from new algorithms. It comes from finding small bugs here and there in the data pipeline, in the model training pipeline. Those gave the biggest boost to the model …”
Feldman: Many AI projects fail due to bad data, not AI
“I think most, many AI projects fail on those fronts. Have nothing to do with the AI. They fail because the data was a disaster.”
El-Moursi: Data preparation and pipelines are the key prerequisite for AI agents
“Because once you actually solve for that on-ramp of how to get your data on, how to cleanse it, how to protect it, how to make it usable and create the data pipelines then the agentic piece, you can ask any question, and that's what you're engaging with today …”
Ben Rogojan: Data type issues still routinely break data pipelines
“But to this day, people are still having like a data type issue break a pipeline that then causes everyone to have to kind of spend a little bit of time fixing them.”
Moses: Validating data at a single pipeline point is no longer sufficient
“And so making sure that your data is accurate at only one point of the pipeline is just no longer sufficient”
Dada argues building data pipelines is harder than machine learning itself
“And this is where we believe that machine learning is not really the hard part, but building and maintaining the data pipeline is”