Insight certainty 4/5 debate potential 3/5

Soldani: Frontier LLM pre-training requires at least 50,000 GPUs

Luca Soldani · Best of 2024: Open Models [LS LIVE! at NeurIPS 2024] · Dec 23, 2024 · at 11:11

Luca Soldani of the Allen Institute for AI outlines the escalating compute tiers required across different stages of LLM development at NeurIPS 2024.

0:00 / 0:22exact quote · 22.8s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“To give you a sense of, like, how I personally think about research budget for each part of the language model pipeline is, like, on the pre-training side, you can maybe do something with a thousand GPUs. Really, you want 10,000. And, like, if you want real estate of the art, you know, your DeepSeq minimum is like 50,000.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Luca Soldani

Opinion
Soldani: Open Model Bio-Risk Warnings Were a Lobbying Ploy
“You know, if you remember the beginning of this year, it was all about bio-risk of these open models. The whole thing fizzled out because there's been, finally there's been, like, rigorous research, not just this paper from coherent folks, but there's been rig…”
Luca Soldani Dec 23, 2024 ▶ 23:09 Best of 2024: Open Models [LS LIVE! at NeurIPS 2024]
Insight
Soldani: Replicating OpenAI's o1 requires roughly 10,000 GPUs
“If you're interested in you know, your, Open replication of what OpenAI's O-one is you're gonna be on the 10 K spectrum of our GPUs.”
Luca Soldani Dec 23, 2024 ▶ 12:08 Best of 2024: Open Models [LS LIVE! at NeurIPS 2024]
Opinion
Soldani: Web blocking disproportionately benefits incumbent closed AI labs
“And I think the problem is this blocking or ideas really, it impacts people in different ways. It disproportionately helps companies that have a head start, which are usually the closed labs, and it hurts incoming newcomer players where you either have now to …”
Luca Soldani Dec 23, 2024 ▶ 20:34 Best of 2024: Open Models [LS LIVE! at NeurIPS 2024]
Opinion
Soldani: Local models beat closed models in retrieval applications
“There are some applications where local models just blow closed models out of the water. So, like, retrieval is a very clear example.”
Luca Soldani Dec 23, 2024 ▶ 3:32 Best of 2024: Open Models [LS LIVE! at NeurIPS 2024]
Assertion Supported
Soldani: Llama and Qwen models fail OSI open source AI definition
“Under this definition, for example, Lama or some of the Quen models are not open source because the license says you can, you can't use this model for this, or it says if you use this model, you have to name the output this way or derivative needs to be named …”
Luca Soldani Dec 23, 2024 ▶ 7:07 Best of 2024: Open Models [LS LIVE! at NeurIPS 2024]
Opinion
Soldani: AI Risks Are Standard Software Issues, Not Existential Threats
“The thing that it's like to me is sorry, it's ingenuous, is like just putting this AI on a pedestal and calling it like an unknown alien technology that has like new and undiscovered potentials to destroy humanity. When in reality, all the dangers I think are …”
Luca Soldani Dec 23, 2024 ▶ 21:56 Best of 2024: Open Models [LS LIVE! at NeurIPS 2024]
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.