Insight certainty 3/5 debate potential 2/5

McKinzie: Multi-agent RL is a good baseline for human collaboration

Brandon McKinzie · No Priors Ep. 113 | With OpenAI's Eric Mitchell and Brandon McKinzie · May 1, 2025 · at 26:38

OpenAI research scientist Brandon McKinzie explains to Sarah Guo how multi-agent reinforcement learning can simulate teamwork and human interaction.

0:00 / 0:17exact quote · 18.0s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“There's no reason you can't scale all this up so that models are trained to be really good at cooperating with each other. I mean, there's a lot of already existing literature on multi-agent RL and yeah, if you want the model to be good at something like collaborating with a bunch of people, like maybe a not too bad starting point is making it good with collaborating with other models.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Brandon McKinzie

Insight
McKinzie: Tools prevent reasoning models from degrading during test-time compute
“We've in the past for our reasoning models talked a lot about test time scaling, and I think for a lot of problems you know, without tools, test time scaling might occasionally work and, but at some point the model is just kind of ranting in its internal chain…”
Brandon McKinzie May 1, 2025 ▶ 3:45 No Priors Ep. 113 | With OpenAI's Eric Mitchell and Brandon McKinzie
Opinion
McKinzie: General reasoning models could unify with robotics foundation models
“And I personally don't see any reason why we couldn't have this, these be this, the same model.”
Brandon McKinzie May 1, 2025 ▶ 22:42 No Priors Ep. 113 | With OpenAI's Eric Mitchell and Brandon McKinzie
Assertion Not checkable as stated
McKinzie: Tool use noticeably changes test-time scaling for visual reasoning
“We've seen exactly that, like the test time scaling slopes for, without tool use and with tool use for visual reasoning specifically are very noticeably different.”
Brandon McKinzie Oct 31, 2025 ▶ 11:31 No Priors Ep. 138 | The Best of 2025 (So Far) with Sarah Guo and Elad Gil
Assertion Supported
McKinzie: Reinforcement learning is the key differentiator behind o3 reasoning
“I guess the short answer is reinforcement learning is, is the biggest one. So yeah, rather than just having to predict the next token and some large pre-training corpus from, you know you know, everywhere essentially now we have a more focused goal of the mode…”
Brandon McKinzie May 1, 2025 ▶ 3:20 No Priors Ep. 113 | With OpenAI's Eric Mitchell and Brandon McKinzie
Assertion Supported
McKinzie: Tool use improves test-time scaling slopes for visual reasoning
“And we've seen exactly that, like the test time scaling slopes for without tool use and with tool use for visual reasoning specifically are very noticeably different.”
Brandon McKinzie May 1, 2025 ▶ 9:55 No Priors Ep. 113 | With OpenAI's Eric Mitchell and Brandon McKinzie
Assertion Not checkable as stated
McKinzie: OpenAI has run out of reliable evaluation benchmarks for recent models
“Especially with some of our recent models where we've kind of run out of Reliable evals to track because they kind of just solved a few of those.”
Brandon McKinzie May 1, 2025 ▶ 33:40 No Priors Ep. 113 | With OpenAI's Eric Mitchell and Brandon McKinzie
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.