Datasets
topic on 6 shows · 11 statements across 11 episodes
Latent Space
No Priors
Capital Allocators
the MAD Podcast
the a16z Podcast
20VC
11 statements about Datasets, every show
Hershberg: Drug development alpha relies on proprietary datasets and modeling tools
“To actually make something more beautiful and more potent at this stage of the game, the alpha probably is in really interesting data sets, new modeling tools and capabilities.”
Morris: AI paradigm shifts are driven by novel datasets, not architectures
“I think, like, all of the things that I would consider paradigm shifts in the Kuhnian sense came from a new technique, but trained on new data, and I think the new data is super, super important”
Swyx: Model papers at NeurIPS are dead as research shifts to datasets
“The focus I saw is that model papers at NeurIPS are kind of dead. No one really presents models anymore. It's just data sets because it's all the grad students are working on.”
Karpathy: Model architecture is no longer the fundamental bottleneck in AI
“I don't think that the neural network architecture is like holding us back fundamentally anymore. It's like not the bottleneck, whereas I think in the previous, before Transformer, it was a bottleneck, but now it's not the bottleneck. So now we're talking a lo…”
2022 DeepMind paper showed dataset size matters more than parameter count
“In fact, in twenty-twenty-two, a pivotal paper came out that changed the way that many people in the research community thought about this very calculus. And it demonstrated that datasets were actually more important than just the sheer size of the model.”
Emad Mostaque: Every Nation Will Need Sovereign Datasets and Open Models
“As part of that, every nation will need their own data sets, which again, have from broadcaster data. They will need their own open models that can stimulate innovation internally as well.”
Braga: Systematica seeks to process 50 to 100 new datasets yearly
“Just this morning I was having a meeting here about how do we enlarge our capability to process 50 to a hundred new data sets per year.”
Data teams should focus on computation processes rather than physical datasets
“People think of data sets as physical things, right? Like a table and a database. Right. But in the modern world where you're really applying software engineering processes to data, all that data ends up being computed. And so what we thought really is that yo…”
Johnson: Twitter holds an incredibly valuable but underutilized dataset
“I think Twitter is a great example where it's not that they're not doing nothing, they're not doing, they're doing some things with the data, it's just they could do a lot more. I think it's one of the most valuable data sets on the planet, but one would think…”
CrowdFlower claims to have the largest dataset determining if images are funny
“We have a gigantic, maybe the biggest data set available on, isn't image funny?”
In an era of data abundance, the main risk is opportunity cost
“The real danger in an era of abundant data is really opportunity cost. If you spend too much time playing with a data set that you can't extract value from, you missed an opportunity to play with another data set that you might actually get value from.”