Synthetic Data Generation
topic on 4 shows · 5 statements across 5 episodes
BG2 Pod
Latent Space
No Priors
Invest Like the Best
5 statements about Synthetic Data Generation, every show
Zaharia: Customizing AI models will get significantly easier over time
“My feeling is, like customizing models is actually going to get way easier over time. That's what we're finding, because The base models are smarter, so they generate better traces in RL already, and then RL is about learning from your own past traces, and the…”
Patel: AI's primary challenge is domain data generation, not model size
“The challenge today is not necessarily make the model bigger. The challenge is how do I generate and create data that is in useful domains so that the model gets better at them.”
Patel: AI models may improve faster over the next 6-12 months
“We may actually see models improve faster in the next six months to a year than we saw them improve in the last year. Because there's this new axis of synthetic data generation and the amount of compute we can throw at it is, we're still right here in the scal…”
Karpathy: Synthetic data risks silent distribution collapse without injected entropy
“When you're doing synthetic data generation, this is a problem, because you actually really want that entropy. You want the diversity and richness in your data set. Otherwise, you're getting collapsed data sets, and you can't see it when you look at any indivi…”