Sankar: Agentic AI use cases took off after reasoning models dropped
“All kinds of agentic use cases, which really started becoming a thing after the reasoning model dropped.”
George says post-reasoning AI models drove an adoption takeoff for Harvey
“Now, fast forward, post reasoning models, like that totally flipped. And you could see, you know, absolute takeoff of adoption, right? And so a bunch of different things happen at the same time. Lawyers got way more value out of the product. You could see it i…”
Sawhney: Backtracking is a general-purpose reasoning tool, not math-specific
“A lot of these behaviors that we're describing mathematically, like backtracking, or kind of starting again, I mean, these are not really specific to mathematics. I mean, we're seeing them specifically in mathematics in these examples, but kind of, they're gen…”
Kolter: Reasoning models are much harder to jailbreak via probability optimization
“Reasoning models were much more effective because you can't really do the same trick of optimizing for a probability with a reasoning model that has a whole trace of reasoning that happens in the middle and kind of reflect a bit more. So it's much harder to br…”
Kolter: Major AI breakthroughs require both massive scale and luck
“Reasoning models were the next big breakthrough. Those are rare. They do take kind of a, you know, both, both a massive scale and kind of a bit of luck to get there.”
May: Reasoning AI tactics increase token usage 20% of the time
“Sometimes when you try to use like one of these reasoning models where you add chain of thought or one of these tactics, they'll actually use more tokens about 20% of the time. It'll be more expensive.”
Lopopolo: Reasoning models eliminate need for rigid state-machine scaffolding
“And this I think is like the fundamental difference between reasoning models and the four ones and four O's of the past where these models could not think. So you kind of had to put them in boxes with a predefined set of state transitions. Whereas here we have…”
Bose: Enterprise context grows more valuable as AI reasoning models improve
“The real value we're providing, again, is with the enterprise-weight context and the shared memory. And so that becomes instantly more valuable as the reasoning model gets better.”
Fedus: Reasoning and coding agents connect software AI to physical domains
“And I think those were foundational technologies necessary To then connect these systems to the physical world. Like it was just not impossible, not possible with like the AI technology of.”
Turley: ChatGPT reasoning currently serves only power users but is transformative
“When you look at reasoning in ChatGPT today, it's irrelevant to a very small group of people. It's relevant for the people who are trying to get the most out of ChatGPT, but I fundamentally believe that reasoning, it's transformative”
Turley: Early OpenAI reasoning model swore after making a mistake
“We were showing this chain of thought as it was streaming out of the model and the model swore and said, like, oh, damn it, may have to adjust, because it realized it had made a mistake in the puzzle. And the fact that it did that, But in particular, the fact …”
LeCroix: Reasoning models are a major priority for Mistral AI
“Reasoning is a big priority”
Pineau: Reasoning models fail at multi-tier hierarchical planning
“That's the part that the reasoning models don't do. They do really well at like one level of granularity... But the going back and forth between different levels of sort of resolution of action, it's really hard. So on the technical terms, we call it hierarchi…”
Reasoning models consume 10x more tokens on average than non-reasoning models
“So, earlier this year, and probably when you and George last spoke for the AI engineers world's fair, we had this great slide that was super easy, where we would show that the average reasoning model is using 10 times the number of tokens per query in our inte…”
AI progress would have completely stalled in 2024 without reasoning models
“Had reasoning not come along, there would have been no AI progress from mid 24 Through, essentially, Gemini three. There would have been none. Everything would have stalled.”
Kaiser: Reasoning models are the second major milestone after Transformers
“One point was, of course, the Transformers when it started, but the other point was reasoning models.”
Kaiser: OpenAI began working on reasoning models around three years ago
“So we started working on it maybe three years ago”
Łukasz Kaiser: Reasoning models require verifiable data, excelling in math and coding
“So currently, and current for at least the Most basic ways we use it currently, it needs to be fairly verifiable. So there is an, is your answer correct or not? You prepare data for that. You can do that in mathematics, coding very well. You can do this in sci…”
Kaiser: Test-time compute increases AI capabilities faster than pre-training
“Using more tokens to think increases your capability, and it increases it, given the computation, way faster than pre-training, right?”
Lenz: Most enterprises avoid reasoning models due to high latency
“Most enterprises don't really want to use reasoning models. The latencies is too high”
Zelikman: Training reasoning models on just positive examples causes a plateau
“So if you only train on like the positive examples, then you end up in this kind of like potential minimum where there's just no more data that it can actually solve.”
Altman: OpenAI continues achieving fundamental breakthroughs in deep learning and reasoning
“And deep learning has been this miracle that keeps on giving, and we have kept finding, like, breakthrough after breakthrough. Again, when we got the reasoning model breakthrough, like, I also thought that was like, we're never gonna get another one like that.…”
Dwivedi: Deterministic workflows cannot solve complex enterprise incident debugging
“No amount of workflows will suffice for a big enterprise. Like you have to link together some of the missing pieces, some of the poorly instrumented data, and that requires world knowledge and a few iterations with the world knowledge.”
Patel: Reasoning models are driving a surge in AI inference demand
“Inference demand has been skyrocketing this year, right? These reasoning models, the revenue it's been skyrocketing this year”
Swix: Reasoning Models Have Better Context Utilization Than Standard LLMs
“I have a theory also that reasoning models have better context utilization because they can loop back. Normal auto-aggressive models, they just kind of go left to right, but reasoning models, in theory, they can loop back and look for things that they needed c…”
Lightcap: Most free ChatGPT users haven't used reasoning models yet
“Most of them have actually not experienced the power of the reasoning models. They mostly are using GPT-IV-O and, you know, they mostly are kind of using it for this very kind of you know, turn-based kind of like very quick you know, back and forth, almost sea…”
Dwarkesh: Reasoning models will outperform GPT-4o on real-world deductive tasks
“I think a reasoning model, I think a reasoning model will be more reliable and be better at solving that kind of problem than Poirot.”
Lambert: The RL algorithm is not the most important component in reasoning models
“I definitely don't think the algorithm tends to be the most important thing.”
Lambert: North star of reasoning models is dynamic token budget calibration
“I think that has to be the north star for most people working on reasoning, which is the model will just Spend the right amount of tokens on it.”
Nadella: Reasoning models paired with human synthesis are the true frontier
“Now, like having a sophisticated reasoning model and your prefrontal cortex work together whereas a lot of the mundane Stuff is getting done by even some core agent or what have you. That I think is definitely the frontier.”
Brown: Deep Research proves reasoning models work in unverifiable domains
“And that is very clearly a domain where you don't have an easily verifiable metric for success. It's very like, what is the best research report that you could generate? And yet these models are doing extremely well at this domain. So I think that's like an ex…”
Chen: AI reasoning only emerges at large scale requiring massive compute
“When you look at reasoning you just don't see that happen at small scale, right? There's like a certain scale at which it starts becoming signal bearing and that requires you to have resources, right?”
Gomez: Creating AI reasoning models is dramatically cheaper than pre-training
“It's easy to create a reasoning model. It's dramatically cheaper than pre-training. And so it's accessible. And so there's this huge intelligence uplift that comes for really quite little effort.”
Gomez: AI reasoning models will expand into medicine and physical sciences
“We've just scratched the surface at the moment. It's mostly focused on, you know math problems and this sort of thing. There is a whole world of applications that we need to make it work work in medicine, you know, everything from the pure sciences, physics, c…”
Will Brown: AI reasoning models are merely a stepping stone toward autonomous agents
“The thing that's going to make the next wave of stuff be powerful is just, like, everyone wants better agents. Everyone wants models that can, like, go off and do stuff. And, like, reasoning was kind of, like, a precursor to that a little bit.”
Reasoning models acting as reward models are key to agent RL
“And the most, one of the most promising ways, I think, towards doing this is having the reward models also be able to answer harder questions by themselves being reasoning models.”
Marcus: 'Reasoning' AI models only mimic patterns without genuine abstractions
“Now, the reason I wouldn't call them reasoning models, though you're right that many people do, is what I think they're doing is basically copying patterns of human reasoning. They're getting data about how humans reason certain things, but the depth of reason…”
McKinzie: Tools prevent reasoning models from degrading during test-time compute
“We've in the past for our reasoning models talked a lot about test time scaling, and I think for a lot of problems you know, without tools, test time scaling might occasionally work and, but at some point the model is just kind of ranting in its internal chain…”
Mitchell: AI reasoning improvements will not be limited to math and code
“So like there, I think there's some reason for spikiness, but I think some people will probably go too far with this and saying like, oh yes, these models will only be really good at math and code. And like, not, you know, like everything else is like, you can…”
Factorio benchmark results show reasoning models underperform expectations on extended planning
“One thing we have found in preliminary results is that the reasoning models don't seem to do as well as you'd expect in this setting. And I think that's probably because the way we set this up, it's a bit like we're already making it do reasoning traces over a…”
Fulford: Training reasoning models on math and coding generalizes to writing
“So I think in general you will always get a model better, better at a specific task if you train on that task, but we also see a lot of generalization from training on one kind of task to, you know, other domains. So you can train a reasoning model on mostly m…”
Patel: Generating pre-answer reasoning tokens yields superior AI performance
“Models now will think for some time before they answer. And this enables much better performance on all sorts of tasks, whether it be coding or math or understanding science or understanding complex Social dilemmas, right? All sorts of different topics they're…”
Taylor: Reasoning models generate net new ideas, breaking the data wall
“What's really interesting about, you know, reasoning and reasoning models is I think I feel really optimistic these models are generating net new ideas, and so it really affords the opportunity to break through some of these, the data wall as well.”
Pokrass: Pair reasoning models for planning with smaller models for execution
“I do think reasoning models for planning and using kind of more targeted models to execute is definitely a good architecture.”
GPT-4.1 excels at exploring repositories, while reasoning models dominate targeted file changes
“Basically, where GPT, 4.1, can it kind of explore, go through a repo? It's been trained to do that particularly well. Whereas you know, to just get some code and produce a change, a reasoning model might do better because it can kind of reason over the entire …”
OpenAI currently restricts reinforcement fine-tuning exclusively to its reasoning models
“No, that's reinforcement fine tuning is only for reasoning models.”
Suleyman: Reasoning models taking desktop action will be cheap and abundant within 15 years
“That's the transition that's going to happen over the next 15 years is that it is going to be a cheap and basically abundant resource to have these reasoning models that can take action in your workplace, that can orchestrate your apps and get things done for …”
Hendrycks: Recent AI reasoning models score in 90th percentile on wet-lab guidance
“We are finding that with the most recent reasoning models quite unlike the models from two years ago, like the initial GPT-IV, the most recent reasoning models are getting around 90th percentile compared to these expert level virologists in their area of exper…”
Srinivas: Many reasoning models will exist, but few products will integrate personal context well
“There's gonna be a bunch of great reasoning models, but there's not gonna be a hundred products that really package, personal context all the API integration, services integrations native integration to your phone to be an assistant really well.”
Acharya: Voice plus reasoning models will rapidly eliminate unwanted hallucinations
“This is where I think the capability is just going to get better and better faster than we appreciate. You know, with the language models, they're prone to hallucination, and there are certain conversations like the therapy one that benefit from the hallucinat…”
Shankar: Evals are necessary to train AI reasoning models
“You need evals to train your reasoning models.”
Dohmke: Improved model reasoning will push SWE-bench scores near 100%
“As the models get better in reasoning we're going to get closer to a hundred percent of this VBench, which is that benchmark out of 12 repos open source Python repos a team in Princeton identified 2200 or so issue pull request pairs. Effectively, all the model…”
Appenzeller: Reasoning models now dominate top AI model rankings
“If you look at the slide here that shows the current ranking of one of the best AI models that we have today, you'll see that pretty much the whole top of the rankings has been taken over by reasoning models.”
Appenzeller: Switching to reasoning models would increase inference compute needs 20x
“Very roughly, if we all, if everybody would switch tomorrow from whatever they have today to a reasoning model, we would need 20 times more inference.”
Hendrycks: Reasoning models have reached expert-level virology capabilities
“The AIs are getting very good at STEM PhD level types of topics, and that includes virology. So I think that they are sort of rounding the corner on being able to provide expert level capabilities in terms of their knowledge of the literature, Or even helping …”
Kilpatrick: Reasoning will solve multi-item retrieval in long context
“And like, it feels like, again, like back to this, the thread around these capabilities, like it feels like long context with reasoning is like finally going to be that thing where like, it actually just like blows the lid off of it. And like, it makes the use…”
Chen: AI models cannot learn reasoning from scratch without pre-trained knowledge
“You need knowledge in order to build reasoning on top of it.
Right.
a model can't kind of go in blind and just learn reasoning from scratch.
So we find these two paradigms to be fairly complementary and we think, you know, they have feedback loops on each oth…”
Chen: GPT-4.5 outshines reasoning models like o1 in creative writing
“And, you know, we find that in a lot of areas like creative writing, for instance
Again, this is stuff that we want to test over the next one or two months but we find that there are areas like creative writing where this model outshines reasoning models.”
Roucher: Visual AI models will likely jump the reasoning S-curve in 2025
“But as with text agents, we've really found that we made a jump on the S curve with reasoning models. I think it's going to be the same with the next visual models, basically better base models just allow you to jump over this S curve. And probably I think it'…”
Karina Nguyen: Verification difficulty makes alignment crucial for reasoning models
“The question of like alignment is actually more important for this like complex reasoning models to like, how do we help humans to like verify the outputs of these models is quite important.”