Cosmopedia
product on 1 show · 2 statements across 1 episodes · said 4 times in 1 episodes since 2024
Mentions by year, every show
tap a year for its mentions
Latent Space 4
every mention on every show, scene by scene, with the transcript →
2 statements about Cosmopedia, every show
Ben Allal: Prompt diversity is essential for scaling synthetic training data
“The key ingredient to getting a good data set that is synthetic is trying as much as possible to keep it diverse, because if you just throw the same prompts as your model, like generate, like, a textbook about linear algebra, and even if you change the tempera…”
Ben Allal: LLMs can be trained with entirely synthetic pipelines
“Today you can train an LLM with like an entirely synthetic pipeline. For example, you can use our Cosmopedia data sets and you can train a one B model on like a hundred and fifty billion tokens. Those are a hundred percent synthetic, and those are also of good…”