Disclosure certainty 4/5 debate potential 2/5

Brown: OpenAI team is scaling test-time compute to hours and days

Noam Brown · Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI · Jun 19, 2025 · at 41:59

Noam Brown, lead researcher at OpenAI, explains the research focus of his team beyond traditional multi-agent systems.

0:00 / 0:21exact quote · 21.5s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“The team, in many ways, is actually a misnomer, because we're working on more than just multi-agent. Multi-agent is one of the things we're working on. Some other things we're working on is just like being able to scale up test time compute by a ton. So how, you know, we get these models thinking for 15 minutes now, how do we get them to think for hours, how do we get them to think for days, even longer, and be able to solve incredibly difficult problems.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Noam Brown

Opinion
Brown: LLMs implicitly develop world models through scale alone
“I think it's pretty clear that as these models get bigger, they have a world model, and that world model becomes better with scale. So they are implicitly developing a world model, and I don't think it's something that you need to explicitly model.”
Noam Brown Jun 19, 2025 ▶ 52:30 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Opinion
Brown: AI models implicitly develop theory of mind through scale
“If these models become smart enough, they develop things like theory of mind. They develop an understanding that there are other agents that like can take actions and have motives and all this stuff. And these models just develop that implicitly with scale and…”
Noam Brown Jun 19, 2025 ▶ 53:31 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Assertion Not checkable as stated
Brown: GPT-4o and o3 are passing the Turing test
“So at this point, like, you know, the truth is, you know, GPT-IV-O and like O-III, these models are like passing the Turing test.”
Noam Brown Jun 19, 2025 ▶ 3:28 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Insight
Brown: Deep Research proves reasoning models work in unverifiable domains
“And that is very clearly a domain where you don't have an easily verifiable metric for success. It's very like, what is the best research report that you could generate? And yet these models are doing extremely well at this domain. So I think that's like an ex…”
Noam Brown Jun 19, 2025 ▶ 7:32 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Insight
Noam Brown: The Ideal AI Agent Harness Is No Harness
“The ideal harness is no harness. Right. I think harnesses are like a crutch that eventually we're going to be able to move beyond.”
Noam Brown Jun 19, 2025 ▶ 14:03 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Prediction Not checkable as stated
Brown: Model routers will become obsolete as unified models emerge
“We've said pretty openly that we want to move to a world where there is a single unified model. And in that world, you shouldn't need a router on top of the model. So I think that the router issue Will eventually be solved also.”
Noam Brown Jun 19, 2025 ▶ 19:00 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.