Reyes: Companies will post-train internal commodity models for specialized tasks
“Businesses that will say, well, you know, we do a lot of commodity tasks, but there's a couple of very high volume specialized tasks. That only we do. And for those, your commodity model won't be good enough. Your frontier model will be too expensive. And so t…”
Trojanowski: Context engineering is runtime training with far less data
“I think the framework I like is thinking about it as training data. Except you're just training the model at runtime. It is training data. And because you, the model's learning at inference time, the total amount of training data is far lower, right? Like the …”
Kant: Big model post-training recipes transfer down to small models, not up
“It's not very helpful to have a post training recipe for a smaller model and try to apply it to a bigger model. Yeah. It just, in all cases, you're gonna have to rethink most of the recipe. But recipe for post training for a bigger model applied to a smaller m…”
Bubna: Modal multi-node training targets post-training, not large-scale pre-training
“And we're not going for obviously like large scale pre-training runs. The thing that we've built multi-handle training for is we see a lot of smaller scale post-training like people are post-training like medium-sized fun models so they can get higher quality …”
Zico Kolter: Reinforcement learning is now the foundation of all AI post-training
“RL is now the foundation of really all post training. It's all done by RL.”
Srivastava: Startups should not do post-training before achieving product-market fit
“Hey, go find, go prove to yourself with the best in class model that you have something worth optimizing. And I think, you know, A lot of, you know, if a customer comes to us, was that meme, which was like, it was like two years ago, it feels like there's no G…”
AI model progress will alternate between pre-training and post-training breakthroughs
“We're going to be having a bit of a swing back and forth between pre-training and post-training.”
Post-training techniques cannot compensate for a weak base AI model
“Pre-training is still the foundation and like, you can never post-train your way out of a week-based model.”
Yi Tay: 'Reasoning' Technically Just Means Post-Training RL with Thinking Trajectories
“So I think the actual, like, technical definition of reasoning is making models better with thinking and post-training. Ok? Yeah. So basically, like, RL-ing the model to think better.”
Bourgeau: Recent continual learning progress has mostly occurred via post-training search tools
“First, I think a lot of progress has been made on this front since in the last few years. I think this is mostly around post-training, around search, use search tools and then make search calls, then they would have access to that new information.”
Post-training scaling laws drove all AI benchmark progress since October 2024
“And so all the progress we've had immense progress since October, 24 through today was based entirely on these two new scaling laws.”
Chen: AI post-training is an art driven by taste, not pure science
“One of the things I often think about is that there's a, it's almost like there's an art to post training. It's not purely a science. Like when you were deciding what kind of model you're trying to create and what it's good at. There's this notion of taste and…”
Sherman Wu: Heavy compute for text model post-training bottlenecks verticalization
“For the text models, there's always going to be this like really big fat free training step that like you have to invest in here. And then even the post training side is like, You know, it's not the, it's not like the easiest thing. Like it's, you know we all,…”
David Owen: AI pre-training receives less focus due to post-training progress
“It seems as if pre-training is comparatively less of a focus than it was before, partly because, like, you have this exciting new direction of, well, new, newish direction of post-training where they've done so much about reasoning”
David Owen: Post-training usage data generates feedback loops for pre-training
“A lot of this stuff is quite synergistic. You develop a better model. You, like, use post-training stuff to make it better. You get a load of data of the model actually being used successfully or not. A lot of that can probably go into pre-training next time.”
Lambert: Scaling AI 10x alters post-training, not pre-training methods
“If like, if we were to train a model that was 10 times as big, like all this post-training stuff would change. But the pre-training And mid training and long contacts, I think would actually become looking pretty similar.”
Huyen: Internet data is maxed out, making post-training the key AI differentiator
“At some point, we are actually, like, have kind of maxed out on, like, internet data, right? And then people, like, text data, people max out. I think a lot of people are doing, like, with other data, like audios and videos, and, like, everyone's trying to thi…”
Harris: Best AI innovations happen in post-training as data runs out
“It does seem like post training is where the best innovations are happening now and the pre-training and the amount of data, like they've, we've used up a lot of the data. They're trying to create synthetic data to try to improve model performance.”
Labenz: AI post-training reasoning currently yields higher ROI than raw scaling
“And it just seems like we're getting more benefit from the post training and the reasoning paradigm than scaling. But I don't think either one is I definitely don't think either one is, is dead.”
Microsoft uses idle nighttime inference capacity for global AI post-training
“What's nice about post training is that you don't have to do it in one large data center in one location. And so part of the technique that we've been focused on is how do we take this inferencing capacity around the world? And a lot of it is idle at night as …”
Anthropic's Joseph: Do Everything Possible in Post-Training Over Pre-Training
“The way I usually think about it is anything you can do in post training, you probably should, because your iteration loop, like the ability to make progress is really fast. You can try something, you can try it again, you can try it again.”
Morcos: Post-training techniques are better applied in pre- and mid-training
“Most of what we do in post-training is better
were done in pre and mid training and earlier on in training in general.”
Morcos: Post-training alignment is ineffective long-term compared to pre-training alignment
“Like fundamentally, I think alignment and post training doesn't really make sense as a long-term solution. If you can easily align a model through post training, you can easily misalign a model through post training. If it's easy to put it in, it's easy to tak…”
Sharma: Industry spending on AI post-training will eventually surpass pre-training
“Like, I believe we will see, you know, just as much money spent on post-training as we will on pre-training, and in the future, more on post-training.”
Lord: AI pre-training gains asymptoted 18 to 24 months ago
“And about 18 months ago, 24 months ago, we started to really see, like, an asymptoting of gains coming from, because they had essentially, like, sucked up all of the knowledge on the internet. And so labs really shifted towards most of the gains now coming fro…”
Kim: AI post-training functions more like art than traditional research
“For post-training, what's really f- or one of the reasons I really like post-training is it feels more like an art than maybe even, like, other areas of research, because you kind of have to make all these trade-offs, right?”
Lightcap: Post-training scaling will dominate AI development for the next 1–2 years
“You know, the O series of models, which were kind of the previous reasoning models were really just the beginning of us starting to explore what's possible in that post-training regime. And I think that's going to be kind of the dominant theme here for the nex…”
Nadella: AI product creation will center on data feedback loops for post-training
“But the interesting thing is the feedback loop, the data path inside the product that then is used in order to post train, in order to be able to do the right tool you know, selection that seems to be the place where product creation is all gonna happen.”
Deng: Partnering post-training researchers with product drives AI breakthroughs
“I think that really that close, tight knit relationship between at any of these large model companies between post training and product is going to produce some really incredible stuff.”
Brown: OpenAI models undergo mid-training and post-training before release
“For open AI models, like, they go through a mid-training step, and then they go through a post-training step, and then they're released, and they're a lot more useful. Like, frankly, if you interacted with the only pre-trained model, it would be super difficul…”
Brin: Post-training and thinking models are a huge, uncapped step forward
“And more recently, the post-training, especially as the thinking models have come around. And that's been, you know, another huge step up in general in AI. So you know, we don't really know what the ceiling is.”
Kilpatrick: Pre-Training Isn't Dead; Gains Multiply Through Post-Training and RL
“And this is why, like, I don't subscribe to the, like pre-training is, you know, dead and all that stuff, because the more work that you can do at the pre-training level, those capabilities, as you do post-training and as you give the models RL capability, it'…”
Patel: Generating pre-answer reasoning tokens yields superior AI performance
“Models now will think for some time before they answer. And this enables much better performance on all sorts of tasks, whether it be coding or math or understanding science or understanding complex Social dilemmas, right? All sorts of different topics they're…”
Agarwal: Optimal Post-Training Pipeline Combines Heavy Distillation Followed by RL
“So, so I would think maybe an optimal pipeline would look like you do distillation heavily, but then you still do some RL afterwards, because maybe there's still something you can get out of your reward functions or whatever your post-training stack is.”
Nguyen: Post-training scaling avoids data walls through infinite learnable tasks
“The scaling in post-chaining itself is not hitting the wall, and that's because Basically, we went from, like, raw data sets from pre-trained models to infinite amount of tasks that you can teach the model in the post-training world via reinforcement learning.…”
Chip Huyen notes major AI labs keep post-training research proprietary
“And unfortunately, a lot of labs that are doing it are not quite, like, publishing papers about it.”
Chip Huyen argues post-training is what differentiates frontier AI models
“So, so I do think that post-training is what makes this, like, really big lab models are, like, different.”