Kedrosky: Investors will pressure AI labs to slash massive pre-training spend
“Once investors look under the hood and see more and more of this, they'll be questioning, why are we spending so much on pre-training? Why are you doing billion dollar training runs anymore? If most of the gains and models are coming from post training and RLH…”
Kant: Reinforcement learning will move earlier into LLM pre-training
“I have I would say a not commonly held opinion that reinforcement learning will move earlier and earlier into pre-training.”
Kant: RL compute cannot scale like pre-training due to task batch constraints
“And RL is batch size constraint, right? So like you are ultimately in your batch size constraint because you don't have infinite tasks, right? When you've got the entire web, you can be much more flexible in scaling up your batch size because you've got the en…”
Sierra will not pre-train foundation models, leaving capex to major AI labs
“And we're not, you know, we're not doing our own pre-training. We'll leave the capital expense there to, you know, the labs and the larger companies.”
OpenAI's three research pillars are pre-training, RL, and alignment
“At the very highest level, right, we have an org that focuses on pre-training, right, which is, you know, giving models a lot of world knowledge. We focus on RL, like, teaching the models how to reason with that knowledge, how to chain the little insights toge…”
Chen: Pre-training is not dead and remains underrated in AI research
“Well, I think if you still have a pre-training is dead view of the world I think pre-training is definitely yeah, yeah, not, not dead. It's underrated.”
AI reinforcement learning energy demand probably already exceeds traditional pre-training
“And this is an area that is becoming huge in terms of energy demand. It'll, it will, The probably already is bigger than what we have historically considered training, you know, pre-training”
Hassabis: Current AI paradigms will be part of final AGI architecture
“The components that you just mentioned, I'm pretty sure will be part of the final architecture for AGI. So I think they've come such a long way now and we've proven out so many things about what they can do. I can't see a world in which we will sort of realize…”
AI model progress will alternate between pre-training and post-training breakthroughs
“We're going to be having a bit of a swing back and forth between pre-training and post-training.”
Post-training techniques cannot compensate for a weak base AI model
“Pre-training is still the foundation and like, you can never post-train your way out of a week-based model.”
New pre-training techniques will drastically boost base AI model capabilities
“The way that we used to do pre-training, maybe, like, you know, like two, a year ago or two years ago maybe, like, you know, diminishing return is, like, obvious, but I can see how new ideas are bringing, like, you know, fresh, fresh energy into the pre-traini…”
Brockman: Pre-Training Capability Multiplies Through the Entire AI Model Pipeline
“Every single step of the model production pipeline multiplies. And so you want to improve all of them. And the thing that we see is we prove the pre-training. It makes all the other steps much easier. And it makes sense because it's a model is able to learn fa…”
Lample: Mistral is far from reaching pre-training saturation
“We are still working a lot on the pre-training side. We are very, very far from any sort of situation on the pre-training.”
LLMs favor CLI tools over APIs due to massive pre-training data volumes
“I think that in pre-training, there's just an enormous amount of command line data. Like even let's ignore, let's like, let's ignore RL. Like you're doing no harness post training. Just the amount of like CLI versus API documentation for just like navigating t…”
Core AI capabilities must be built during pre-training, not just fine-tuned
“If there is a core capability that you actually care about, that capability should be part of the foundation and not a fine-tuned artifact.”
Pre-training is no longer where the low-hanging AI gains lie
“Pre-training is not dead, but pre-training is boring. So it's not where the low hanging fruit is anymore.”
Izmailov: AI researchers cannot reliably trace model behaviors to pre-training sources
“We don't really know what's the source of this type of behaviors, but that's also true for a lot of other behaviors in the models with, like, even the good ones. We don't really, we cannot always pin down, like, where they come from in the pre-training.”
Bourgeau: Architecture and data innovation currently matter more than scale
“The other parts are architecture and data innovation. These also play a really, really important part in the Performance of pre-training and probably even more so than pure scale these days, but scaling is still an important factor as well.”
Kaiser: Reasoning yields far greater AI capability gains per dollar than pre-training
“With the new paradigm of reasoning, you can get much more gains for the same amount of money because it's on this like lower and like, there are just discoveries to be made and these discoveries unlock insane capabilities.”
Kaiser: Pre-training consumes the most GPUs of any AI development stage
“Currently, pre-training just uses the most GPUs of all the parts, so it needs the most GPUs, right?”
Kaiser: Test-time compute increases AI capabilities faster than pre-training
“Using more tokens to think increases your capability, and it increases it, given the computation, way faster than pre-training, right?”
Łukasz Kaiser: AI pre-training expands stored knowledge rather than generalization
“Pre-training is a little different, right? Because it increases the data together with your increase in model size. So it doesn't necessarily increase generalization. It just uses more knowledge.”
David Owen: AI pre-training receives less focus due to post-training progress
“It seems as if pre-training is comparatively less of a focus than it was before, partly because, like, you have this exciting new direction of, well, new, newish direction of post-training where they've done so much about reasoning”
David Owen: Post-training usage data generates feedback loops for pre-training
“A lot of this stuff is quite synergistic. You develop a better model. You, like, use post-training stuff to make it better. You get a load of data of the model actually being used successfully or not. A lot of that can probably go into pre-training next time.”
Kant: Next-token pre-training gains hit a sigmoidal curve and slowed down
“The first paradigm of kind of pre-training of predicting the next token on the web was becoming sigmoidal and was slowing down in terms of the gains that it had.”
Smarter pre-training methods will reduce massive AI data spending requirements
“So I think we're just getting smarter about how to do pre-training rather than shoving everything we have into a bucket and like seeing what happens. And so as a result of that, you might not necessarily have to spend the exact same amount of money to get a ca…”
Nadella: Pre-training Remains More Efficient Than RL Due to Amortization
“Pre-training is a more efficient form of training. Because you can advertise it.”
Future AI models will continue to rely on pre-training data
“Personally, I think that's unlikely. Not, not because pre-training is strictly necessary. I think we may well be able to train something completely from scratch, as we've been able to do in other domains, but more because pre-training on this vast data sets th…”
Schrittwieser: AI pre-training risks over-restricting an agent's exploration search space
“I think the main, you know, the main challenge or the main thing you need to watch out for is that you don't over encode or you don't restrict your search space too much. If your pre-training, if your prior knowledge prevents you from exploring something that …”
Reinforcement learning scaling yields returns on compute similar to pre-training
“If you look at all the RL literature over time, we see very similar returns on compute in pre-training and in RL, where we can invest exponentially more compute in RL and keep getting benefits.”
Huyen: Internet data is maxed out, making post-training the key AI differentiator
“At some point, we are actually, like, have kind of maxed out on, like, internet data, right? And then people, like, text data, people max out. I think a lot of people are doing, like, with other data, like audios and videos, and, like, everyone's trying to thi…”
Tworek: Pre-Training on Unlabeled Data Yields Far More Intelligence Than Supervised Mapping
“There are many more bits usually in the targets than in the labels and studying the structure of targets itself. It yields much more learning and much more intelligence than learning the mapping itself. So like spending a whole compute on just learning the dat…”
Jerry Tworek: Pre-training AI models is mathematically simple compared to RL
“The first thing that is important to know and understand, RL is hard. Like, conceptually, if you think about it, and there's still a lot of depth to it, but very conceptually, mathematically speaking, pre-training is dead simple.”
Tworek: Reinforcement learning and pre-training require each other to succeed
“And like, I don't like in terms of a pure RL, I don't think like really pure RL makes sense. RL needs Pre-training to be successful. And I think pre-training, as I said before, needs RL to be successful as well.”
Harris: Best AI innovations happen in post-training as data runs out
“It does seem like post training is where the best innovations are happening now and the pre-training and the amount of data, like they've, we've used up a lot of the data. They're trying to create synthetic data to try to improve model performance.”
Patel: AI models cannot learn external memory usage from pre-training alone
“How do I train it to interact with these databases and these word documents that it writes to? Because it's never gonna learn that from pre-training. Has to learn that from an environment.”
Morcos: Post-training techniques are better applied in pre- and mid-training
“Most of what we do in post-training is better
were done in pre and mid training and earlier on in training in general.”
Morcos: Post-training alignment is ineffective long-term compared to pre-training alignment
“Like fundamentally, I think alignment and post training doesn't really make sense as a long-term solution. If you can easily align a model through post training, you can easily misalign a model through post training. If it's easy to put it in, it's easy to tak…”
Sharma: Industry spending on AI post-training will eventually surpass pre-training
“Like, I believe we will see, you know, just as much money spent on post-training as we will on pre-training, and in the future, more on post-training.”
Lord: AI pre-training gains asymptoted 18 to 24 months ago
“And about 18 months ago, 24 months ago, we started to really see, like, an asymptoting of gains coming from, because they had essentially, like, sucked up all of the knowledge on the internet. And so labs really shifted towards most of the gains now coming fro…”
Pre-Training Proprietary Models Does Not Build Moats for AI Application Startups
“Training your own model, certainly pre-training it, not helpful. Yes, you need data to fine tune, but it's not actually a ton of data. And so a lot of people have access to that. And so in many ways the differentiation that I think will continue to compound is…”
Kaplan: Compute scaling drives AI progress more than researcher cleverness
“Basically you can Scale up the compute in both pre-training and RL and get better and better performance. And I think that's sort of the fundamental thing that is driving AI progress. It's not that AI researchers are really smart or they suddenly got smart. It…”
Laskin: RL requires far fewer FLOPs than pre-training for frontier models
“We're in this brief period in history right now where the RL flops are still manageable. Like you can really have a best in class product if you're focused. And yes, you'll need to put, you know, you still need a decent amount of GPUs, but from a flops perspec…”
Patel: Pre-training returns are plateauing; GPT-4.5 was unimpressive and deprecated
“Pre-training seems to have been giving us these plateauing returns. We make these models bigger. GPT-Fort .5 didn't seem to be all that impressive. They had to deprecate it.”
Kantrowitz: Pre-training data walls justify $100M+ packages for top AI talent
“And this is a strength, a sound strategy because you have everybody talking about how pre-training is hitting diminishing returns. You have everybody talking about how data is hitting a wall. And so what do you need? You just need these algorithmic development…”
Patel: Pre-training scaling is seeing diminishing returns
“Pre-training, which is this idea that you just make the model bigger that has had diminishing returns.”
Ameisen: Deceptive Backward Reasoning Exists in Base Pre-Trained Models
“I bet, I don't know how much I bet a hundred bucks. So somebody can like, they would get a hundred bucks from me if they prove that I'm wrong, that this behavior for a model that does a drink fine tuning, it also does it post pre-training.”
Kilpatrick: Pre-Training Isn't Dead; Gains Multiply Through Post-Training and RL
“And this is why, like, I don't subscribe to the, like pre-training is, you know, dead and all that stuff, because the more work that you can do at the pre-training level, those capabilities, as you do post-training and as you give the models RL capability, it'…”
Patel: AI models ingest dangerous data during pre-training for world knowledge
“So you don't want to just filter out everything so that the model doesn't know anything about it but at the same time, you don't want it to output, you know, how to build a bomb so there's like a fine balance here, and that's why pre-training is defined as pre…”
Chen: AI models cannot learn reasoning from scratch without pre-trained knowledge
“You need knowledge in order to build reasoning on top of it.
Right.
a model can't kind of go in blind and just learn reasoning from scratch.
So we find these two paradigms to be fairly complementary and we think, you know, they have feedback loops on each oth…”
Patel: AI pre-training gains are becoming logarithmically more expensive
“So, the whole paradigm of training, you know, pre-training is, is, is not slowing down. It's just, it's logarithmically more expensive each, for each generation, for each incremental improvement.”
Fine-Tuning Cannot Effectively Add New Languages to Large Language Models
“There's just no way you can do that without intervening on pre-training. You can't like fine tune or post train Japanese into a model effectively. And so you have to start from scratch.”
Huang: AI Inference and Post-Training Are Now Just as Hard as Pre-Training
“People used to think that pre-training was hard, and inference was easy. Now everything is hard.”
Bret Taylor: AI pre-training outside AGI labs wastes capital
“Unless you are an AGI research lab, doing pre-training on a model I believe is just burning capital.”
Taylor: AI pre-training will consolidate into a small number of frontier model builders
“I think that will probably play out with the frontier models. We'll end up with a relatively small number of companies doing pre-training you know, which is the really capital intensive part of model building not because, you know, there's not, they're the onl…”
Howard: Training stages form a continuum allowing deep modification of pre-trained models
“Sorry, it wasn't the end of fine-tuning, but more that we should treat it as a continuum, and we should have much higher expectations of how much you can do with an already trained model. You can really add a lot of behavior to it. You can change its behavior.…”
Vinyals predicts pre-training compute will drop to ~50% as RL expands
“So to me, that balance feels correct, like some on pre-training, and here we, we're trying to learn every task. So certainly that's going to be, you know, let's say it can be as high as 50%, not as high as over 90 like today. And then the rest mostly on reinfo…”
Lambert: RLHF provides a richer signal per unit of compute than pre-training
“As reinforcement learning is so much less compute, like it, Is a richer signal in terms of its impact, because if they could do what RLHF is doing at pre-training, they would, but they don't know how to have that effect in like a stable manner. Otherwise, ever…”
Liang: Controlling AI hallucination is easier once models understand the concept
“So I think there's pre-training, which is predicting the next word and developing a world model, so to speak. And with those capabilities, then you can, you still have to say don't hallucinate, but it will be much easier to control that model if it has a notio…”