May 1, 2025 · 38m · no-priors

No Priors Ep. 113 | With OpenAI's Eric Mitchell and Brandon McKinzie

Eric Mitchell · 14m spoken Brandon McKinzie · 10m spoken Sarah Guo · 6m spoken Elad Gil · 4m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of No Priors, OpenAI research scientists Eric Mitchell and Brandon McKinzie discuss the technical breakthroughs behind the o3 reasoning model, detailing how reinforcement learning, test-time compute scaling, and autonomous tool integration drive modern agentic workflows. They also explore the computational constraints of physical simulation, architectural unification, and the urgent need for uncontaminated evaluation benchmarks.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 29.4% of the talking time here. How this is scored →

The hosts as informed peer 5.5 Guest teaching 5.4 Guest disagreement 1.5 The hosts pushing back 1.9
05100:0010:0020:0030:000:29–3:19 · The hosts as informed peer 4/10 Deliberative Reasoning and Autonomous Tool Integration in o3 Elad opens by asking the guests to explain the core architectural and conceptual differences between o3 and previous foundation models. Eric explains how deliberative test-time thinking and autonomous tool use like code execution enable higher-level problem solving without prescriptive user prompting.3:20–5:23 · The hosts as informed peer 5/10 Reinforcement Learning and Test-Time Compute Scaling Brandon explains that reinforcement learning underpins o3's ability to take time on difficult problems rather than merely predicting tokens. Elad demonstrates domain familiarity by referencing OpenAI's published curves on compute duration versus accuracy.5:24–8:52 · The hosts as informed peer 6/10 Model Unification, Steerability, and Uncertainty Estimation Sarah pushes back against the model unification thesis, asking why research should merge reasoning with pre-training when consumers only care about fast accuracy and developers want granular API control. Brandon and Eric argue that ideal models should internally assess their own uncertainty and become steerable based on context.8:52–11:05 · The hosts as informed peer 5/10 Why Tool Execution Enhances Test-Time Scaling Sarah references previous guest Noam Brown while asking for intuition on why tool access improves test-time scaling efficiency. Brandon explains image-cropping tools in visual reasoning, while Eric points out that executing deterministic code avoids wasteful token-level verification cycles.11:05–13:53 · The hosts as informed peer 5/10 Training Deep Research Agents and Browsing Workflows Elad inquires into the specialized RL objectives and training data used to construct Deep Research. Eric details why web browsing serves as an ideal testbed for long-horizon behaviors and how RL rewards are aligned to user tolerance for rollout latency.13:53–15:55 · The hosts as informed peer 6/10 Accelerating AI Research and Software Engineering Capabilities Brandon highlights software engineering and accelerating internal AI research velocity as premier application areas. Elad contributes his own synthesis on AI self-improvement acting as the fastest bootstrap mechanism toward superintelligence.15:55–19:09 · The hosts as informed peer 5/10 Autonomous Computer Use and Operational Safety Sandboxes Brandon shares his enthusiasm for ambient computer use while Sarah notes the remaining gaps in complex business workflows. Eric provides a measured reality check, emphasizing asymmetric operational risks and the necessity of keeping agents in sandboxes.19:09–21:53 · The hosts as informed peer 6/10 Framework for Environmental Uncertainty and Task Simulation Sarah asks for an organizing mental model to categorize tasks amenable to test-time scaling and RL. Eric delivers a masterclass, framing task difficulty around environmental uncertainty versus internal memorization, and physical latency bottlenecks versus digital simulatability.21:54–25:21 · The hosts as informed peer 7/10 Robotics Convergence, Physical Latency, and Perception Biases Elad draws on historical precedents like Codex merging into GPT to ask whether standalone robotics foundation models will be subsumed into general weights. When Eric notes real-world real-time physics constraints, Elad hits back with biological examples showing simple organisms like ants and frogs compute physics with minimal parameters.25:21–29:25 · The hosts as informed peer 6/10 Simulating Human Teamwork and Multi-Agent Reinforcement Learning Sarah questions how simulation environments can accurately reflect complex human software collaboration. Brandon floats multi-agent RL simulation as a spicy solution, while Eric humorously characterizes humans as high-uncertainty, high-latency tool calls.29:25–32:22 · The hosts as informed peer 6/10 Evaluating Domain Spikiness Versus Algorithmic Generalization Sarah posits that post-training RL will cause model capability advances to be spikier across domains compared to pre-training. Eric and Brandon clarify her definition and push back against the assumption that reasoning models will only improve in math and code, pointing to creative writing transfer.32:22–34:39 · The hosts as informed peer 5/10 The Critical Need for Uncontaminated Benchmark Evals When Sarah asks what dream training dataset the team would request, Eric intentionally dodges to prioritize uncontaminated evaluation benchmarks. Brandon adds the need for long-horizon multi-turn interaction logs across large codebases.34:39–37:48 · The hosts as informed peer 6/10 Power-User Prompting Distributions and Work Delegation Strategies Eric expresses frustration with social media users running single-shot evals and ignoring the wide response distribution of reasoning models. Sarah and Elad pitch concrete product ideas, including best-of-n auto-ranking and cross-sample response synthesis.0:29–3:19 · Guest teaching 5/10 Deliberative Reasoning and Autonomous Tool Integration in o3 Elad opens by asking the guests to explain the core architectural and conceptual differences between o3 and previous foundation models. Eric explains how deliberative test-time thinking and autonomous tool use like code execution enable higher-level problem solving without prescriptive user prompting.3:20–5:23 · Guest teaching 5/10 Reinforcement Learning and Test-Time Compute Scaling Brandon explains that reinforcement learning underpins o3's ability to take time on difficult problems rather than merely predicting tokens. Elad demonstrates domain familiarity by referencing OpenAI's published curves on compute duration versus accuracy.5:24–8:52 · Guest teaching 4/10 Model Unification, Steerability, and Uncertainty Estimation Sarah pushes back against the model unification thesis, asking why research should merge reasoning with pre-training when consumers only care about fast accuracy and developers want granular API control. Brandon and Eric argue that ideal models should internally assess their own uncertainty and become steerable based on context.8:52–11:05 · Guest teaching 6/10 Why Tool Execution Enhances Test-Time Scaling Sarah references previous guest Noam Brown while asking for intuition on why tool access improves test-time scaling efficiency. Brandon explains image-cropping tools in visual reasoning, while Eric points out that executing deterministic code avoids wasteful token-level verification cycles.11:05–13:53 · Guest teaching 5/10 Training Deep Research Agents and Browsing Workflows Elad inquires into the specialized RL objectives and training data used to construct Deep Research. Eric details why web browsing serves as an ideal testbed for long-horizon behaviors and how RL rewards are aligned to user tolerance for rollout latency.13:53–15:55 · Guest teaching 4/10 Accelerating AI Research and Software Engineering Capabilities Brandon highlights software engineering and accelerating internal AI research velocity as premier application areas. Elad contributes his own synthesis on AI self-improvement acting as the fastest bootstrap mechanism toward superintelligence.15:55–19:09 · Guest teaching 5/10 Autonomous Computer Use and Operational Safety Sandboxes Brandon shares his enthusiasm for ambient computer use while Sarah notes the remaining gaps in complex business workflows. Eric provides a measured reality check, emphasizing asymmetric operational risks and the necessity of keeping agents in sandboxes.19:09–21:53 · Guest teaching 7/10 Framework for Environmental Uncertainty and Task Simulation Sarah asks for an organizing mental model to categorize tasks amenable to test-time scaling and RL. Eric delivers a masterclass, framing task difficulty around environmental uncertainty versus internal memorization, and physical latency bottlenecks versus digital simulatability.21:54–25:21 · Guest teaching 6/10 Robotics Convergence, Physical Latency, and Perception Biases Elad draws on historical precedents like Codex merging into GPT to ask whether standalone robotics foundation models will be subsumed into general weights. When Eric notes real-world real-time physics constraints, Elad hits back with biological examples showing simple organisms like ants and frogs compute physics with minimal parameters.25:21–29:25 · Guest teaching 6/10 Simulating Human Teamwork and Multi-Agent Reinforcement Learning Sarah questions how simulation environments can accurately reflect complex human software collaboration. Brandon floats multi-agent RL simulation as a spicy solution, while Eric humorously characterizes humans as high-uncertainty, high-latency tool calls.29:25–32:22 · Guest teaching 6/10 Evaluating Domain Spikiness Versus Algorithmic Generalization Sarah posits that post-training RL will cause model capability advances to be spikier across domains compared to pre-training. Eric and Brandon clarify her definition and push back against the assumption that reasoning models will only improve in math and code, pointing to creative writing transfer.32:22–34:39 · Guest teaching 6/10 The Critical Need for Uncontaminated Benchmark Evals When Sarah asks what dream training dataset the team would request, Eric intentionally dodges to prioritize uncontaminated evaluation benchmarks. Brandon adds the need for long-horizon multi-turn interaction logs across large codebases.34:39–37:48 · Guest teaching 5/10 Power-User Prompting Distributions and Work Delegation Strategies Eric expresses frustration with social media users running single-shot evals and ignoring the wide response distribution of reasoning models. Sarah and Elad pitch concrete product ideas, including best-of-n auto-ranking and cross-sample response synthesis.0:29–3:19 · Guest disagreement 1/10 Deliberative Reasoning and Autonomous Tool Integration in o3 Elad opens by asking the guests to explain the core architectural and conceptual differences between o3 and previous foundation models. Eric explains how deliberative test-time thinking and autonomous tool use like code execution enable higher-level problem solving without prescriptive user prompting.3:20–5:23 · Guest disagreement 1/10 Reinforcement Learning and Test-Time Compute Scaling Brandon explains that reinforcement learning underpins o3's ability to take time on difficult problems rather than merely predicting tokens. Elad demonstrates domain familiarity by referencing OpenAI's published curves on compute duration versus accuracy.5:24–8:52 · Guest disagreement 2/10 Model Unification, Steerability, and Uncertainty Estimation Sarah pushes back against the model unification thesis, asking why research should merge reasoning with pre-training when consumers only care about fast accuracy and developers want granular API control. Brandon and Eric argue that ideal models should internally assess their own uncertainty and become steerable based on context.8:52–11:05 · Guest disagreement 1/10 Why Tool Execution Enhances Test-Time Scaling Sarah references previous guest Noam Brown while asking for intuition on why tool access improves test-time scaling efficiency. Brandon explains image-cropping tools in visual reasoning, while Eric points out that executing deterministic code avoids wasteful token-level verification cycles.11:05–13:53 · Guest disagreement 1/10 Training Deep Research Agents and Browsing Workflows Elad inquires into the specialized RL objectives and training data used to construct Deep Research. Eric details why web browsing serves as an ideal testbed for long-horizon behaviors and how RL rewards are aligned to user tolerance for rollout latency.13:53–15:55 · Guest disagreement 1/10 Accelerating AI Research and Software Engineering Capabilities Brandon highlights software engineering and accelerating internal AI research velocity as premier application areas. Elad contributes his own synthesis on AI self-improvement acting as the fastest bootstrap mechanism toward superintelligence.15:55–19:09 · Guest disagreement 2/10 Autonomous Computer Use and Operational Safety Sandboxes Brandon shares his enthusiasm for ambient computer use while Sarah notes the remaining gaps in complex business workflows. Eric provides a measured reality check, emphasizing asymmetric operational risks and the necessity of keeping agents in sandboxes.19:09–21:53 · Guest disagreement 1/10 Framework for Environmental Uncertainty and Task Simulation Sarah asks for an organizing mental model to categorize tasks amenable to test-time scaling and RL. Eric delivers a masterclass, framing task difficulty around environmental uncertainty versus internal memorization, and physical latency bottlenecks versus digital simulatability.21:54–25:21 · Guest disagreement 1/10 Robotics Convergence, Physical Latency, and Perception Biases Elad draws on historical precedents like Codex merging into GPT to ask whether standalone robotics foundation models will be subsumed into general weights. When Eric notes real-world real-time physics constraints, Elad hits back with biological examples showing simple organisms like ants and frogs compute physics with minimal parameters.25:21–29:25 · Guest disagreement 2/10 Simulating Human Teamwork and Multi-Agent Reinforcement Learning Sarah questions how simulation environments can accurately reflect complex human software collaboration. Brandon floats multi-agent RL simulation as a spicy solution, while Eric humorously characterizes humans as high-uncertainty, high-latency tool calls.29:25–32:22 · Guest disagreement 3/10 Evaluating Domain Spikiness Versus Algorithmic Generalization Sarah posits that post-training RL will cause model capability advances to be spikier across domains compared to pre-training. Eric and Brandon clarify her definition and push back against the assumption that reasoning models will only improve in math and code, pointing to creative writing transfer.32:22–34:39 · Guest disagreement 2/10 The Critical Need for Uncontaminated Benchmark Evals When Sarah asks what dream training dataset the team would request, Eric intentionally dodges to prioritize uncontaminated evaluation benchmarks. Brandon adds the need for long-horizon multi-turn interaction logs across large codebases.34:39–37:48 · Guest disagreement 2/10 Power-User Prompting Distributions and Work Delegation Strategies Eric expresses frustration with social media users running single-shot evals and ignoring the wide response distribution of reasoning models. Sarah and Elad pitch concrete product ideas, including best-of-n auto-ranking and cross-sample response synthesis.0:29–3:19 · The hosts pushing back 1/10 Deliberative Reasoning and Autonomous Tool Integration in o3 Elad opens by asking the guests to explain the core architectural and conceptual differences between o3 and previous foundation models. Eric explains how deliberative test-time thinking and autonomous tool use like code execution enable higher-level problem solving without prescriptive user prompting.3:20–5:23 · The hosts pushing back 1/10 Reinforcement Learning and Test-Time Compute Scaling Brandon explains that reinforcement learning underpins o3's ability to take time on difficult problems rather than merely predicting tokens. Elad demonstrates domain familiarity by referencing OpenAI's published curves on compute duration versus accuracy.5:24–8:52 · The hosts pushing back 4/10 Model Unification, Steerability, and Uncertainty Estimation Sarah pushes back against the model unification thesis, asking why research should merge reasoning with pre-training when consumers only care about fast accuracy and developers want granular API control. Brandon and Eric argue that ideal models should internally assess their own uncertainty and become steerable based on context.8:52–11:05 · The hosts pushing back 1/10 Why Tool Execution Enhances Test-Time Scaling Sarah references previous guest Noam Brown while asking for intuition on why tool access improves test-time scaling efficiency. Brandon explains image-cropping tools in visual reasoning, while Eric points out that executing deterministic code avoids wasteful token-level verification cycles.11:05–13:53 · The hosts pushing back 1/10 Training Deep Research Agents and Browsing Workflows Elad inquires into the specialized RL objectives and training data used to construct Deep Research. Eric details why web browsing serves as an ideal testbed for long-horizon behaviors and how RL rewards are aligned to user tolerance for rollout latency.13:53–15:55 · The hosts pushing back 2/10 Accelerating AI Research and Software Engineering Capabilities Brandon highlights software engineering and accelerating internal AI research velocity as premier application areas. Elad contributes his own synthesis on AI self-improvement acting as the fastest bootstrap mechanism toward superintelligence.15:55–19:09 · The hosts pushing back 2/10 Autonomous Computer Use and Operational Safety Sandboxes Brandon shares his enthusiasm for ambient computer use while Sarah notes the remaining gaps in complex business workflows. Eric provides a measured reality check, emphasizing asymmetric operational risks and the necessity of keeping agents in sandboxes.19:09–21:53 · The hosts pushing back 1/10 Framework for Environmental Uncertainty and Task Simulation Sarah asks for an organizing mental model to categorize tasks amenable to test-time scaling and RL. Eric delivers a masterclass, framing task difficulty around environmental uncertainty versus internal memorization, and physical latency bottlenecks versus digital simulatability.21:54–25:21 · The hosts pushing back 2/10 Robotics Convergence, Physical Latency, and Perception Biases Elad draws on historical precedents like Codex merging into GPT to ask whether standalone robotics foundation models will be subsumed into general weights. When Eric notes real-world real-time physics constraints, Elad hits back with biological examples showing simple organisms like ants and frogs compute physics with minimal parameters.25:21–29:25 · The hosts pushing back 3/10 Simulating Human Teamwork and Multi-Agent Reinforcement Learning Sarah questions how simulation environments can accurately reflect complex human software collaboration. Brandon floats multi-agent RL simulation as a spicy solution, while Eric humorously characterizes humans as high-uncertainty, high-latency tool calls.29:25–32:22 · The hosts pushing back 3/10 Evaluating Domain Spikiness Versus Algorithmic Generalization Sarah posits that post-training RL will cause model capability advances to be spikier across domains compared to pre-training. Eric and Brandon clarify her definition and push back against the assumption that reasoning models will only improve in math and code, pointing to creative writing transfer.32:22–34:39 · The hosts pushing back 1/10 The Critical Need for Uncontaminated Benchmark Evals When Sarah asks what dream training dataset the team would request, Eric intentionally dodges to prioritize uncontaminated evaluation benchmarks. Brandon adds the need for long-horizon multi-turn interaction logs across large codebases.34:39–37:48 · The hosts pushing back 2/10 Power-User Prompting Distributions and Work Delegation Strategies Eric expresses frustration with social media users running single-shot evals and ignoring the wide response distribution of reasoning models. Sarah and Elad pitch concrete product ideas, including best-of-n auto-ranking and cross-sample response synthesis.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 25.7% · guest 74.3%0:00 · the hosts 25.7% · guest 74.3%3:00 · the hosts 34.7% · guest 65.3%3:00 · the hosts 34.7% · guest 65.3%6:00 · the hosts 29.5% · guest 70.5%6:00 · the hosts 29.5% · guest 70.5%9:00 · the hosts 27.8% · guest 72.2%9:00 · the hosts 27.8% · guest 72.2%12:00 · the hosts 16.1% · guest 83.9%12:00 · the hosts 16.1% · guest 83.9%15:00 · the hosts 35.8% · guest 64.2%15:00 · the hosts 35.8% · guest 64.2%18:00 · the hosts 27.5% · guest 72.5%18:00 · the hosts 27.5% · guest 72.5%21:00 · the hosts 22.5% · guest 77.5%21:00 · the hosts 22.5% · guest 77.5%24:00 · the hosts 45.3% · guest 54.7%24:00 · the hosts 45.3% · guest 54.7%27:00 · the hosts 44.1% · guest 55.9%27:00 · the hosts 44.1% · guest 55.9%30:00 · the hosts 18.2% · guest 81.8%30:00 · the hosts 18.2% · guest 81.8%33:00 · the hosts 9.4% · guest 90.6%33:00 · the hosts 9.4% · guest 90.6%36:00 · the hosts 55.1% · guest 44.9%36:00 · the hosts 55.1% · guest 44.9%
Sharpest disagreement ▶ 34:47 Eric's frustration with naive single-shot model comparisons

Eric expresses strong annoyance at Twitter comparisons evaluating reasoning models on a single prompt, explaining that users fail to understand the probabilistic distribution of tool-use rollouts.

Hardest push from the hosts ▶ 7:40 Sarah challenges OpenAI's single unified model strategy

Sarah directly questions the assumption of unified model interfaces, arguing that enterprise developers require decoupled, cheap, and highly steerable task-specific models.

Biggest teaching moment ▶ 19:46 Eric's two-axis framework for task complexity

Eric provides an illuminating conceptual framework, categorizing agent difficulty along environmental uncertainty and the degree to which an environment can be simulated versus bottlenecked by real-world physical time.

The host holds their own ▶ 23:57 Elad's biological counterpoint on physical latency

When Eric argues physical real-time constraints demand heavier compute trade-offs, Elad counters by pointing out that low-compute biological systems like ants and frogs solve real-time physical navigation effortlessly.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Deliberative Reasoning and Autonomous Tool Integration in o3 4511 Elad opens by asking the guests to explain the core architectural and conceptual differences between o3 and previous foundation models. Eric explains how deliberative test-time thinking and autonomous tool use like code execution enable higher-level problem solving without prescriptive user prompting.
Reinforcement Learning and Test-Time Compute Scaling 5511 Brandon explains that reinforcement learning underpins o3's ability to take time on difficult problems rather than merely predicting tokens. Elad demonstrates domain familiarity by referencing OpenAI's published curves on compute duration versus accuracy.
Model Unification, Steerability, and Uncertainty Estimation 6424 Sarah pushes back against the model unification thesis, asking why research should merge reasoning with pre-training when consumers only care about fast accuracy and developers want granular API control. Brandon and Eric argue that ideal models should internally assess their own uncertainty and become steerable based on context.
Why Tool Execution Enhances Test-Time Scaling 5611 Sarah references previous guest Noam Brown while asking for intuition on why tool access improves test-time scaling efficiency. Brandon explains image-cropping tools in visual reasoning, while Eric points out that executing deterministic code avoids wasteful token-level verification cycles.
Training Deep Research Agents and Browsing Workflows 5511 Elad inquires into the specialized RL objectives and training data used to construct Deep Research. Eric details why web browsing serves as an ideal testbed for long-horizon behaviors and how RL rewards are aligned to user tolerance for rollout latency.
Accelerating AI Research and Software Engineering Capabilities 6412 Brandon highlights software engineering and accelerating internal AI research velocity as premier application areas. Elad contributes his own synthesis on AI self-improvement acting as the fastest bootstrap mechanism toward superintelligence.
Autonomous Computer Use and Operational Safety Sandboxes 5522 Brandon shares his enthusiasm for ambient computer use while Sarah notes the remaining gaps in complex business workflows. Eric provides a measured reality check, emphasizing asymmetric operational risks and the necessity of keeping agents in sandboxes.
Framework for Environmental Uncertainty and Task Simulation 6711 Sarah asks for an organizing mental model to categorize tasks amenable to test-time scaling and RL. Eric delivers a masterclass, framing task difficulty around environmental uncertainty versus internal memorization, and physical latency bottlenecks versus digital simulatability.
Robotics Convergence, Physical Latency, and Perception Biases 7612 Elad draws on historical precedents like Codex merging into GPT to ask whether standalone robotics foundation models will be subsumed into general weights. When Eric notes real-world real-time physics constraints, Elad hits back with biological examples showing simple organisms like ants and frogs compute physics with minimal parameters.
Simulating Human Teamwork and Multi-Agent Reinforcement Learning 6623 Sarah questions how simulation environments can accurately reflect complex human software collaboration. Brandon floats multi-agent RL simulation as a spicy solution, while Eric humorously characterizes humans as high-uncertainty, high-latency tool calls.
Evaluating Domain Spikiness Versus Algorithmic Generalization 6633 Sarah posits that post-training RL will cause model capability advances to be spikier across domains compared to pre-training. Eric and Brandon clarify her definition and push back against the assumption that reasoning models will only improve in math and code, pointing to creative writing transfer.
The Critical Need for Uncontaminated Benchmark Evals 5621 When Sarah asks what dream training dataset the team would request, Eric intentionally dodges to prioritize uncontaminated evaluation benchmarks. Brandon adds the need for long-horizon multi-turn interaction logs across large codebases.
Power-User Prompting Distributions and Work Delegation Strategies 6522 Eric expresses frustration with social media users running single-shot evals and ignoring the wide response distribution of reasoning models. Sarah and Elad pitch concrete product ideas, including best-of-n auto-ranking and cross-sample response synthesis.

Statements from this episode (21)

Assertion Supported
Mitchell: o3 autonomously executes multi-step tasks using integrated tools
“Not only is the model it's on its own smarter than our previous O series models, which is great, but it's also able to use all these tools that like further enhance its abilities and whether that's doing like research on something where you want up-to-date inf…”
Eric Mitchell May 1, 2025 ▶ 2:03
Assertion Supported
McKinzie: Reinforcement learning is the key differentiator behind o3 reasoning
“I guess the short answer is reinforcement learning is, is the biggest one. So yeah, rather than just having to predict the next token and some large pre-training corpus from, you know you know, everywhere essentially now we have a more focused goal of the mode…”
Brandon McKinzie May 1, 2025 ▶ 3:20
Insight
McKinzie: Tools prevent reasoning models from degrading during test-time compute
“We've in the past for our reasoning models talked a lot about test time scaling, and I think for a lot of problems you know, without tools, test time scaling might occasionally work and, but at some point the model is just kind of ranting in its internal chain…”
Brandon McKinzie May 1, 2025 ▶ 3:45
Disclosure
Mitchell: OpenAI plans to unify models and remove the ChatGPT switcher
“You know, I think for us, like unification of our models is something that, you know, Sam has talked about publicly that, you know, we have this big crazy model switcher in ChatGPT and there are a lot of choices and you know, we have a model that might be good…”
Eric Mitchell May 1, 2025 ▶ 5:24
Opinion
McKinzie: OpenAI is developing models with precise uncertainty understanding
“And I hope we can get to a place where our models have a more precise understanding of their own level of uncertainty. Because you know, if they already know the answer, they should just kind of tell you it. And if it takes them a day to actually figure it out…”
Brandon McKinzie May 1, 2025 ▶ 7:14
Assertion Supported
McKinzie: Tool use improves test-time scaling slopes for visual reasoning
“And we've seen exactly that, like the test time scaling slopes for without tool use and with tool use for visual reasoning specifically are very noticeably different.”
Brandon McKinzie May 1, 2025 ▶ 9:55
Insight
Mitchell: AI offloads tasks lacking comparative advantage to external tools
“I think like part of this is you can just allocate compute a lot more efficiently because you can defer stuff that the model doesn't have comparative advantage to doing to a tool that is like really well suited to doing that thing.”
Eric Mitchell May 1, 2025 ▶ 10:52
Opinion
McKinzie: OpenAI models hit an inflection point navigating internal codebases
“I think our models are getting a lot better very quickly at being actually useful. And it seems like they were kind of reaching some kind of inflection point where They are useful enough to want to reach out to and use like multiple times a day for me at least…”
Brandon McKinzie May 1, 2025 ▶ 14:32
Insight
McKinzie: AI research consists of modular tasks ripe for automated optimization
“And there's so many like different components of research too. There's, it's not just you know, sitting off in the ivory tower thinking about things, but there's like hardware there's you know, various components of training and evaluation and stuff like this.…”
Brandon McKinzie May 1, 2025 ▶ 15:34
Opinion
McKinzie: OpenAI models use external tools with weirdly human-like intuition
“It's also surprising to me how intuitively our models do use the tools we give them access to. It's like weirdly human-like, but I guess that's not too surprising given the data they've seen before,”
Brandon McKinzie May 1, 2025 ▶ 16:50
Insight
Mitchell: OpenAI limits model agency due to asymmetric error costs
“There's a reason we don't go hog wild and say, like, oh yes, here's, like, the keys to the kingdom, like, have at it. There are still, you know, asymmetric costs to, like, the time you can save and the types of errors you can make, and so we're trying to, like…”
Eric Mitchell May 1, 2025 ▶ 17:47
Insight
Mitchell: Physical time bottlenecks make AI tasks harder than simulatable domains
“Stuff that is really bottlenecked by like time, like the physical world is also, you know, just harder than stuff that we can simulate really well.”
Eric Mitchell May 1, 2025 ▶ 21:15
Opinion
McKinzie: General reasoning models could unify with robotics foundation models
“And I personally don't see any reason why we couldn't have this, these be this, the same model.”
Brandon McKinzie May 1, 2025 ▶ 22:42
Insight
Mitchell: Real-world robotics imposes strict latency constraints absent in disembodied AI
“The real world is like an interesting litmus test because at the end of the day, like there is a, you know, frame rate in the real world you need to live on. And it doesn't matter if you get the right answer after you think for two minutes, like, You know, the…”
Eric Mitchell May 1, 2025 ▶ 23:22
Assertion Not checkable as stated
McKinzie: Over 90% of internet clock images show 10:10, biasing vision models
“It's like over 90% or something like that of all clocks on the internet are 10 10.”
Brandon McKinzie May 1, 2025 ▶ 24:45
Insight
McKinzie: Multi-agent RL is a good baseline for human collaboration
“There's no reason you can't scale all this up so that models are trained to be really good at cooperating with each other. I mean, there's a lot of already existing literature on multi-agent RL and yeah, if you want the model to be good at something like colla…”
Brandon McKinzie May 1, 2025 ▶ 26:38
Insight
McKinzie: Humans are an extremely expensive tool call for AI
“Yeah, we are a super expensive tool call. You know, if you're a model, you can either ask me, you know, meat bag over here to you know, help with something and I'll try to think really slowly. In the meantime, it could have like used browser and read like a hu…”
Brandon McKinzie May 1, 2025 ▶ 28:26
Opinion
Mitchell: AI reasoning improvements will not be limited to math and code
“So like there, I think there's some reason for spikiness, but I think some people will probably go too far with this and saying like, oh yes, these models will only be really good at math and code. And like, not, you know, like everything else is like, you can…”
Eric Mitchell May 1, 2025 ▶ 31:16
Insight
Mitchell: High-quality evaluation benchmarks are underappreciated compared to training data
“I mean, yeah, like you want, you know, good data to train on and that's of course valuable for making the model better, but I think it is often neglected how also important it is to have high quality data, which is like a different definition of high quality w…”
Eric Mitchell May 1, 2025 ▶ 32:58
Assertion Not checkable as stated
McKinzie: OpenAI has run out of reliable evaluation benchmarks for recent models
“Especially with some of our recent models where we've kind of run out of Reliable evals to track because they kind of just solved a few of those.”
Brandon McKinzie May 1, 2025 ▶ 33:40
Insight
Mitchell: o3 output distribution makes single-prompt evaluations misleading
“O-three can do really cool things, like when it chains together a lot of tool calls, and then, like, sometimes for the same prompt, it won't have that, you know, moment of magic, or it will, you know, just take a little, it'll do a little less work for you, an…”
Eric Mitchell May 1, 2025 ▶ 35:21
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.