May 1, 2025 · 38m · no-priors
No Priors Ep. 113 | With OpenAI's Eric Mitchell and Brandon McKinzie
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of No Priors, OpenAI research scientists Eric Mitchell and Brandon McKinzie discuss the technical breakthroughs behind the o3 reasoning model, detailing how reinforcement learning, test-time compute scaling, and autonomous tool integration drive modern agentic workflows. They also explore the computational constraints of physical simulation, architectural unification, and the urgent need for uncontaminated evaluation benchmarks.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 29.4% of the talking time here. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Eric expresses strong annoyance at Twitter comparisons evaluating reasoning models on a single prompt, explaining that users fail to understand the probabilistic distribution of tool-use rollouts.
Hardest push from the hosts ▶ 7:40 Sarah challenges OpenAI's single unified model strategySarah directly questions the assumption of unified model interfaces, arguing that enterprise developers require decoupled, cheap, and highly steerable task-specific models.
Biggest teaching moment ▶ 19:46 Eric's two-axis framework for task complexityEric provides an illuminating conceptual framework, categorizing agent difficulty along environmental uncertainty and the degree to which an environment can be simulated versus bottlenecked by real-world physical time.
The host holds their own ▶ 23:57 Elad's biological counterpoint on physical latencyWhen Eric argues physical real-time constraints demand heavier compute trade-offs, Elad counters by pointing out that low-compute biological systems like ants and frogs solve real-time physical navigation effortlessly.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Deliberative Reasoning and Autonomous Tool Integration in o3 | 4 | 5 | 1 | 1 | Elad opens by asking the guests to explain the core architectural and conceptual differences between o3 and previous foundation models. Eric explains how deliberative test-time thinking and autonomous tool use like code execution enable higher-level problem solving without prescriptive user prompting. | |
| Reinforcement Learning and Test-Time Compute Scaling | 5 | 5 | 1 | 1 | Brandon explains that reinforcement learning underpins o3's ability to take time on difficult problems rather than merely predicting tokens. Elad demonstrates domain familiarity by referencing OpenAI's published curves on compute duration versus accuracy. | |
| Model Unification, Steerability, and Uncertainty Estimation | 6 | 4 | 2 | 4 | Sarah pushes back against the model unification thesis, asking why research should merge reasoning with pre-training when consumers only care about fast accuracy and developers want granular API control. Brandon and Eric argue that ideal models should internally assess their own uncertainty and become steerable based on context. | |
| Why Tool Execution Enhances Test-Time Scaling | 5 | 6 | 1 | 1 | Sarah references previous guest Noam Brown while asking for intuition on why tool access improves test-time scaling efficiency. Brandon explains image-cropping tools in visual reasoning, while Eric points out that executing deterministic code avoids wasteful token-level verification cycles. | |
| Training Deep Research Agents and Browsing Workflows | 5 | 5 | 1 | 1 | Elad inquires into the specialized RL objectives and training data used to construct Deep Research. Eric details why web browsing serves as an ideal testbed for long-horizon behaviors and how RL rewards are aligned to user tolerance for rollout latency. | |
| Accelerating AI Research and Software Engineering Capabilities | 6 | 4 | 1 | 2 | Brandon highlights software engineering and accelerating internal AI research velocity as premier application areas. Elad contributes his own synthesis on AI self-improvement acting as the fastest bootstrap mechanism toward superintelligence. | |
| Autonomous Computer Use and Operational Safety Sandboxes | 5 | 5 | 2 | 2 | Brandon shares his enthusiasm for ambient computer use while Sarah notes the remaining gaps in complex business workflows. Eric provides a measured reality check, emphasizing asymmetric operational risks and the necessity of keeping agents in sandboxes. | |
| Framework for Environmental Uncertainty and Task Simulation | 6 | 7 | 1 | 1 | Sarah asks for an organizing mental model to categorize tasks amenable to test-time scaling and RL. Eric delivers a masterclass, framing task difficulty around environmental uncertainty versus internal memorization, and physical latency bottlenecks versus digital simulatability. | |
| Robotics Convergence, Physical Latency, and Perception Biases | 7 | 6 | 1 | 2 | Elad draws on historical precedents like Codex merging into GPT to ask whether standalone robotics foundation models will be subsumed into general weights. When Eric notes real-world real-time physics constraints, Elad hits back with biological examples showing simple organisms like ants and frogs compute physics with minimal parameters. | |
| Simulating Human Teamwork and Multi-Agent Reinforcement Learning | 6 | 6 | 2 | 3 | Sarah questions how simulation environments can accurately reflect complex human software collaboration. Brandon floats multi-agent RL simulation as a spicy solution, while Eric humorously characterizes humans as high-uncertainty, high-latency tool calls. | |
| Evaluating Domain Spikiness Versus Algorithmic Generalization | 6 | 6 | 3 | 3 | Sarah posits that post-training RL will cause model capability advances to be spikier across domains compared to pre-training. Eric and Brandon clarify her definition and push back against the assumption that reasoning models will only improve in math and code, pointing to creative writing transfer. | |
| The Critical Need for Uncontaminated Benchmark Evals | 5 | 6 | 2 | 1 | When Sarah asks what dream training dataset the team would request, Eric intentionally dodges to prioritize uncontaminated evaluation benchmarks. Brandon adds the need for long-horizon multi-turn interaction logs across large codebases. | |
| Power-User Prompting Distributions and Work Delegation Strategies | 6 | 5 | 2 | 2 | Eric expresses frustration with social media users running single-shot evals and ignoring the wide response distribution of reasoning models. Sarah and Elad pitch concrete product ideas, including best-of-n auto-ranking and cross-sample response synthesis. |