Insight certainty 3/5 debate potential 3/5

McKinzie: Tools prevent reasoning models from degrading during test-time compute

Brandon McKinzie · No Priors Ep. 113 | With OpenAI's Eric Mitchell and Brandon McKinzie · May 1, 2025 · at 3:45

OpenAI research scientist Brandon McKinzie explains why test-time compute scaling requires tool integration rather than just internal chain-of-thought compute.

0:00 / 0:32exact quote · 32.1s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“We've in the past for our reasoning models talked a lot about test time scaling, and I think for a lot of problems you know, without tools, test time scaling might occasionally work and, but at some point the model is just kind of ranting in its internal chain of thought. And especially for like some visual perception ones, it knows that it doesn't, it's not able to see the thing that it needs and it just kind of like loses its mind and goes insane. And I think Tool use is a really important component now to continuing this like test time scaling.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Brandon McKinzie

Opinion
McKinzie: General reasoning models could unify with robotics foundation models
“And I personally don't see any reason why we couldn't have this, these be this, the same model.”
Brandon McKinzie May 1, 2025 ▶ 22:42 No Priors Ep. 113 | With OpenAI's Eric Mitchell and Brandon McKinzie
Assertion Not checkable as stated
McKinzie: Tool use noticeably changes test-time scaling for visual reasoning
“We've seen exactly that, like the test time scaling slopes for, without tool use and with tool use for visual reasoning specifically are very noticeably different.”
Brandon McKinzie Oct 31, 2025 ▶ 11:31 No Priors Ep. 138 | The Best of 2025 (So Far) with Sarah Guo and Elad Gil
Assertion Supported
McKinzie: Reinforcement learning is the key differentiator behind o3 reasoning
“I guess the short answer is reinforcement learning is, is the biggest one. So yeah, rather than just having to predict the next token and some large pre-training corpus from, you know you know, everywhere essentially now we have a more focused goal of the mode…”
Brandon McKinzie May 1, 2025 ▶ 3:20 No Priors Ep. 113 | With OpenAI's Eric Mitchell and Brandon McKinzie
Assertion Supported
McKinzie: Tool use improves test-time scaling slopes for visual reasoning
“And we've seen exactly that, like the test time scaling slopes for without tool use and with tool use for visual reasoning specifically are very noticeably different.”
Brandon McKinzie May 1, 2025 ▶ 9:55 No Priors Ep. 113 | With OpenAI's Eric Mitchell and Brandon McKinzie
Assertion Not checkable as stated
McKinzie: OpenAI has run out of reliable evaluation benchmarks for recent models
“Especially with some of our recent models where we've kind of run out of Reliable evals to track because they kind of just solved a few of those.”
Brandon McKinzie May 1, 2025 ▶ 33:40 No Priors Ep. 113 | With OpenAI's Eric Mitchell and Brandon McKinzie
Opinion
McKinzie: OpenAI is developing models with precise uncertainty understanding
“And I hope we can get to a place where our models have a more precise understanding of their own level of uncertainty. Because you know, if they already know the answer, they should just kind of tell you it. And if it takes them a day to actually figure it out…”
Brandon McKinzie May 1, 2025 ▶ 7:14 No Priors Ep. 113 | With OpenAI's Eric Mitchell and Brandon McKinzie
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.