reasoning models

also referred to as: reasoning model

18 statements across 15 episodes · 11 bullish · 3 bearish · 14 people on the record · first statement Feb 1, 2025 by Karina Nguyen · across every show →

Everything said about reasoning models, oldest first

Feb 1, 2025 neutral
Opinion
Karina Nguyen: Verification difficulty makes alignment crucial for reasoning models
“The question of like alignment is actually more important for this like complex reasoning models to like, how do we help humans to like verify the outputs of these models is quite important.”
Karina Nguyen Feb 1, 2025 ▶ 21:41 The Agent Reasoning Interface: Claude, ChatGPT Canvas, Tasks, Operator — with Karina Nguyen, OpenAI
Feb 13, 2025 bullish
Prediction Not checkable as stated
Roucher: Visual AI models will likely jump the reasoning S-curve in 2025
“But as with text agents, we've really found that we made a jump on the S curve with reasoning models. I think it's going to be the same with the next visual models, basically better base models just allow you to jump over this S curve. And probably I think it'…”
Aymeric (Emmerich) Feb 13, 2025 ▶ 21:21 smol agents are all you need
Feb 28, 2025 positive
Prediction Not checkable as stated
Kilpatrick: Reasoning will solve multi-item retrieval in long context
“And like, it feels like, again, like back to this, the thread around these capabilities, like it feels like long context with reasoning is like finally going to be that thing where like, it actually just like blows the lid off of it. And like, it makes the use…”
Logan Kilpatrick Feb 28, 2025 ▶ 11:43 Gemini 2.0 Flash and Flash Thinking: the new SOTA models for the agentic era
Mar 13, 2025 positive
Insight
Shankar: Evals are necessary to train AI reasoning models
“You need evals to train your reasoning models.”
Shreya Shankar Mar 13, 2025 ▶ 27:13 [Lightning Pod] Evals: How to Improve AI Consistently — with Hamel Husain and Shreya Shankar
Apr 15, 2025 positive
Insight
Pokrass: Pair reasoning models for planning with smaller models for execution
“I do think reasoning models for planning and using kind of more targeted models to execute is definitely a good architecture.”
Michelle Pokrass Apr 15, 2025 ▶ 28:48 GPT 4.1: The New OpenAI Workhorse
Apr 15, 2025 positive
Insight
GPT-4.1 excels at exploring repositories, while reasoning models dominate targeted file changes
“Basically, where GPT, 4.1, can it kind of explore, go through a repo? It's been trained to do that particularly well. Whereas you know, to just get some code and produce a change, a reasoning model might do better because it can kind of reason over the entire …”
Michelle Pokrass Apr 15, 2025 ▶ 31:20 GPT 4.1: The New OpenAI Workhorse
Apr 15, 2025 neutral
Assertion Supported
OpenAI currently restricts reinforcement fine-tuning exclusively to its reasoning models
“No, that's reinforcement fine tuning is only for reasoning models.”
Michelle Pokrass Apr 15, 2025 ▶ 39:32 GPT 4.1: The New OpenAI Workhorse
Apr 27, 2025 negative
Assertion Supported
Factorio benchmark results show reasoning models underperform expectations on extended planning
“One thing we have found in preliminary results is that the reasoning models don't seem to do as well as you'd expect in this setting. And I think that's probably because the way we set this up, it's a bit like we're already making it do reasoning traces over a…”
Jack Hopkins Apr 27, 2025 ▶ 12:12 ⚡️Factorio Learning Environment: the ultimate Game Agent Eval — Jack Hopkins
May 9, 2025 bullish
Insight
Reasoning models acting as reward models are key to agent RL
“And the most, one of the most promising ways, I think, towards doing this is having the reward models also be able to answer harder questions by themselves being reasoning models.”
Will Brown May 9, 2025 ▶ 11:54 ⚡️Open Questions in Agentic RL — Will Brown (Prime Intellect)
May 23, 2025 positive
Insight
Will Brown: AI reasoning models are merely a stepping stone toward autonomous agents
“The thing that's going to make the next wave of stuff be powerful is just, like, everyone wants better agents. Everyone wants models that can, like, go off and do stuff. And, like, reasoning was kind of, like, a precursor to that a little bit.”
Will Brown May 23, 2025 ▶ 1:29 ⚡️Multi-Turn RL for Multi-Hour Agents — with Will Brown, Prime Intellect
Jun 19, 2025 positive
Insight
Brown: Deep Research proves reasoning models work in unverifiable domains
“And that is very clearly a domain where you don't have an easily verifiable metric for success. It's very like, what is the best research report that you could generate? And yet these models are doing extremely well at this domain. So I think that's like an ex…”
Noam Brown Jun 19, 2025 ▶ 7:32 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Jul 31, 2025 positive
Insight
Lambert: North star of reasoning models is dynamic token budget calibration
“I think that has to be the north star for most people working on reasoning, which is the model will just Spend the right amount of tokens on it.”
Nathan Lambert Jul 31, 2025 ▶ 20:44 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Jul 31, 2025
Insight
Lambert: The RL algorithm is not the most important component in reasoning models
“I definitely don't think the algorithm tends to be the most important thing.”
Nathan Lambert Jul 31, 2025 ▶ 19:28 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Aug 19, 2025 positive
Opinion
Swix: Reasoning Models Have Better Context Utilization Than Standard LLMs
“I have a theory also that reasoning models have better context utilization because they can loop back. Normal auto-aggressive models, they just kind of go left to right, but reasoning models, in theory, they can loop back and look for things that they needed c…”
Shawn Wang Aug 19, 2025 ▶ 18:28 Long Live Context Engineering - with Jeff Huber of Chroma
Oct 5, 2025 negative
Opinion
Dwivedi: Deterministic workflows cannot solve complex enterprise incident debugging
“No amount of workflows will suffice for a big enterprise. Like you have to link together some of the missing pieces, some of the poorly instrumented data, and that requires world knowledge and a few iterations with the world knowledge.”
Raaz Dwivedi Oct 5, 2025 ▶ 32:39 ⚡️Traversal: Causal ML and Reinforcement Learning
Oct 11, 2025 bearish
Opinion
Lenz: Most enterprises avoid reasoning models due to high latency
“Most enterprises don't really want to use reasoning models. The latencies is too high”
Barak Lenz Oct 11, 2025 ▶ 33:15 Building Jamba 3B: the tiny Hybrid Transformer State Space Reasoning Model - Barak Lenz, CTO of AI21
Jan 9, 2026 neutral
Assertion Supported
Reasoning models consume 10x more tokens on average than non-reasoning models
“So, earlier this year, and probably when you and George last spoke for the AI engineers world's fair, we had this great slide that was super easy, where we would show that the average reasoning model is using 10 times the number of tokens per query in our inte…”
Micah Hill-Smith Jan 9, 2026 ▶ 1:06:27 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Apr 7, 2026 positive
Insight
Lopopolo: Reasoning models eliminate need for rigid state-machine scaffolding
“And this I think is like the fundamental difference between reasoning models and the four ones and four O's of the past where these models could not think. So you kind of had to put them in boxes with a predefined set of state transitions. Whereas here we have…”
Ryan Lopopolo Apr 7, 2026 ▶ 11:59 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.