Everything Noam Brown said on any show that made the record, most notable first. Each card names its show and opens the statement there.
Overnight AI intelligence explosion unlikely due to test-time compute bottlenecks
“And I don't think we're headed to that world largely because of the fact that the models rely so much on large scale test time compute. In order to achieve their greatest intelligence. If you, if it requires so much test time compute to unlock the full capabil…”
Brown: Modern AI models can reason for weeks before plateauing
“What we're seeing today with the modern models is that 5.5 and other models can think for, if you scaffold them reasonably well, can think for weeks even before having performance plateau on some of these benchmarks.”
Noam Brown: AI Model Outputs Are Arguably More Trustworthy Than Humans
“I use it day to day for a lot of this kind of stuff, and I think they're at a point now where They've actually been at a point for a while now where I feel like I can just trust the outputs, arguably more than I could trust the output from a human.”
Brown: LLMs implicitly develop world models through scale alone
“I think it's pretty clear that as these models get bigger, they have a world model, and that world model becomes better with scale. So they are implicitly developing a world model, and I don't think it's something that you need to explicitly model.”
Brown: AI models implicitly develop theory of mind through scale
“If these models become smart enough, they develop things like theory of mind. They develop an understanding that there are other agents that like can take actions and have motives and all this stuff. And these models just develop that implicitly with scale and…”
Inference-time compute is the missing scaling dimension for AI reasoning
“This is why I'm interested in the reasoning direction, because I think there's this whole other dimension. That people are not scaling right now, which is the amount of compute at inference time.”
Brown: AI benchmarks must control for test-time compute
“And so I think the proper way to, and so my claim is the proper way to evaluate the models now is you either have some kind of budget for the benchmark, whether it's tokens or cost or time or whatever, or you plot the performance as a function of the amount of…”
Brown: Scaffolding Easily Inflates AI Benchmark Scores Without Real Gains
“It's really easy to show you can do much better than previous benchmarks or previous, previous models on benchmarks by just, for example, scaffolding a bunch of models together. So if you say, okay, well, we're going to, instead of just running this model once…”
Brown predicts AI will zero-shot his entire PhD thesis within one year
“And I wouldn't be surprised if, you know, six months or a year from now, the model is able to do zero shot an entire poker solver, basically my entire PhD thesis in one go.”
Current AI safety frameworks fail to account for test-time compute scaling
“The preparedness frameworks and responsible scaling policies, they don't really account for the amount of tests I'm computed. They just say, okay, well, what's the capability of the model? The problem is we're in a world now where the capability of the model i…”
OpenAI internal model reportedly disproved the Erdős unit distance conjecture
“We used an internal model at OpenAI a few weeks ago to disprove the unit Erdos unit distance conjecture.”
Complex task compute costs fall 10x to 100x per model release
“The model release cycle is every, every couple months we put out a new model that's even more powerful, and so the cost of disproving the Erdos unit distance gesture drops by, like, 10 or a hundred x with every model release cycle. Probably, in some cases, mor…”
OpenAI discourages researchers from using current models on open math problems
“We are trying to encourage people to not spend all their time just, like, Going through all the mathematical open problems, physics problems, and just seeing, pushing the models to their limits to see what they can prove or disprove. Because we really think th…”
Brown: AI cannot invent novel algorithms better than existing research
“Go ahead and like look at all the published work and synthesize that and then try to come up with something novel and it's not able to do it. And I can give it a lot of time and it's still not able to do it.”
Noam Brown: AI community stuck in bad equilibrium publishing static benchmark grids
“I would talk to researchers about we, it makes sense to show the benchmarks with an x-axis, whether it's tokens or cost or time, there should be an x-axis, and everybody would say, like, yeah, that makes sense, we should do that, but. Well, really, their respo…”
Brown: GPT-4o and o3 are passing the Turing test
“So at this point, like, you know, the truth is, you know, GPT-IV-O and like O-III, these models are like passing the Turing test.”
Brown: Deep Research proves reasoning models work in unverifiable domains
“And that is very clearly a domain where you don't have an easily verifiable metric for success. It's very like, what is the best research report that you could generate? And yet these models are doing extremely well at this domain. So I think that's like an ex…”
Noam Brown: The Ideal AI Agent Harness Is No Harness
“The ideal harness is no harness. Right. I think harnesses are like a crutch that eventually we're going to be able to move beyond.”
Brown: Model routers will become obsolete as unified models emerge
“We've said pretty openly that we want to move to a world where there is a single unified model. And in that world, you shouldn't need a router on top of the model. So I think that the router issue Will eventually be solved also.”
Brown: Pre-training scaling will hit economic limits before superintelligence without reasoning
“Like, we're gonna scale it, sure, we're gonna scale these things up by a few more orders of magnitude, they're gonna become more capable, but we're not gonna see superintelligence from just that. And like, yes, if we had a quadrillion dollars to train these mo…”
Noam Brown: Aligned AI will outperform human virtual assistants on effort
“And so if you have an AI model that's, like, actually really aligned, To you and your preferences, then that could end up doing a way better job than a human could. Well, not, not that it's doing a better job than a human could, but like it's doing a better jo…”
Brown: Multi-agent AI civilizations will far surpass current AI capabilities
“And I think that if you're able to have them cooperate and compete with billions of AIs over a long period of time and build up a civilization essentially, the things that they would be able to Produce and answer would be far beyond what is possible today with…”
Brown: Previous multi-agent research was heuristic and ignored Bitter Lesson
“I think that a lot of the approaches that have been taken have been very heuristic and haven't really been following like the bitter lesson approach to scaling and research.”
Brown: Game Theory Optimal fails in collaborative games like Diplomacy
“Basically, when you're playing, like, the zero-sum games, like, poker, Game Theory Optimal works really well. When you're playing a game like Diplomacy, where there's, like, you need to collaborate and compete, and you need, there's room for collaboration, the…”