Ai2's Luca Soldani discusses the compute scale needed for post-training advanced reasoning models at NeurIPS 2024.
Opinion
Soldani: Open Model Bio-Risk Warnings Were a Lobbying Ploy
“You know, if you remember the beginning of this year, it was all about bio-risk of these open models. The whole thing fizzled out because there's been, finally there's been, like, rigorous research, not just this paper from coherent folks, but there's been rig…”
Insight
Soldani: Frontier LLM pre-training requires at least 50,000 GPUs
“To give you a sense of, like, how I personally think about research budget for each part of the language model pipeline is, like, on the pre-training side, you can maybe do something with a thousand GPUs. Really, you want 10,000. And, like, if you want real es…”
Opinion
Soldani: Web blocking disproportionately benefits incumbent closed AI labs
“And I think the problem is this blocking or ideas really, it impacts people in different ways. It disproportionately helps companies that have a head start, which are usually the closed labs, and it hurts incoming newcomer players where you either have now to …”
Opinion
Soldani: Local models beat closed models in retrieval applications
“There are some applications where local models just blow closed models out of the water. So, like, retrieval is a very clear example.”
Assertion Supported
Soldani: Llama and Qwen models fail OSI open source AI definition
“Under this definition, for example, Lama or some of the Quen models are not open source because the license says you can, you can't use this model for this, or it says if you use this model, you have to name the output this way or derivative needs to be named …”
Opinion
Soldani: AI Risks Are Standard Software Issues, Not Existential Threats
“The thing that it's like to me is sorry, it's ingenuous, is like just putting this AI on a pedestal and calling it like an unknown alien technology that has like new and undiscovered potentials to destroy humanity. When in reality, all the dangers I think are …”