Opinion
Brown: LLMs implicitly develop world models through scale alone
“I think it's pretty clear that as these models get bigger, they have a world model, and that world model becomes better with scale. So they are implicitly developing a world model, and I don't think it's something that you need to explicitly model.”
Opinion
Brown: AI models implicitly develop theory of mind through scale
“If these models become smart enough, they develop things like theory of mind. They develop an understanding that there are other agents that like can take actions and have motives and all this stuff. And these models just develop that implicitly with scale and…”
Assertion Not checkable as stated
Brown: GPT-4o and o3 are passing the Turing test
“So at this point, like, you know, the truth is, you know, GPT-IV-O and like O-III, these models are like passing the Turing test.”
Insight
Brown: Deep Research proves reasoning models work in unverifiable domains
“And that is very clearly a domain where you don't have an easily verifiable metric for success. It's very like, what is the best research report that you could generate? And yet these models are doing extremely well at this domain. So I think that's like an ex…”
Insight
Noam Brown: The Ideal AI Agent Harness Is No Harness
“The ideal harness is no harness. Right. I think harnesses are like a crutch that eventually we're going to be able to move beyond.”
Prediction Not checkable as stated
Brown: Model routers will become obsolete as unified models emerge
“We've said pretty openly that we want to move to a world where there is a single unified model. And in that world, you shouldn't need a router on top of the model. So I think that the router issue Will eventually be solved also.”
Prediction Not checkable as stated
Brown: Pre-training scaling will hit economic limits before superintelligence without reasoning
“Like, we're gonna scale it, sure, we're gonna scale these things up by a few more orders of magnitude, they're gonna become more capable, but we're not gonna see superintelligence from just that. And like, yes, if we had a quadrillion dollars to train these mo…”
Insight
Noam Brown: Aligned AI will outperform human virtual assistants on effort
“And so if you have an AI model that's, like, actually really aligned, To you and your preferences, then that could end up doing a way better job than a human could. Well, not, not that it's doing a better job than a human could, but like it's doing a better jo…”
Prediction Not checkable as stated
Brown: Multi-agent AI civilizations will far surpass current AI capabilities
“And I think that if you're able to have them cooperate and compete with billions of AIs over a long period of time and build up a civilization essentially, the things that they would be able to Produce and answer would be far beyond what is possible today with…”
Opinion
Brown: Previous multi-agent research was heuristic and ignored Bitter Lesson
“I think that a lot of the approaches that have been taken have been very heuristic and haven't really been following like the bitter lesson approach to scaling and research.”
Insight
Brown: Game Theory Optimal fails in collaborative games like Diplomacy
“Basically, when you're playing, like, the zero-sum games, like, poker, Game Theory Optimal works really well. When you're playing a game like Diplomacy, where there's, like, you need to collaborate and compete, and you need, there's room for collaboration, the…”
Opinion
Brown: Scaling self-play beyond zero-sum games will not be as easy as AlphaGo
“My point is that, like, this is where the AlphaGo analogy breaks down. And, not necessarily breaks down, but, like, it's not going to be as easy as self-play was in AlphaGo.”
Insight
Brown: Reasoning models improve primarily through compute efficiency rather than longer thinking duration
“These models are becoming more efficient in the way they're thinking, as they're able to do more with the same amount of test time compute, and I think that's a very underappreciated point, that it's not just that we're getting these models to think for longer…”
Prediction Not checkable as stated
Brown: Reasoning models will progress rapidly into agentic behavior
“I think that we're going to continue to see, as I said before, that we're going to see this paradigm continue to progress rapidly. And I think that that's true even today, that we saw that with like going from O-one preview to O-one to O-three, consistent prog…”
Insight
Noam Brown: Models need baseline capabilities to benefit from test-time reasoning
“One thing that I think is underappreciated is that the models, the pre-trained models need a certain level of capability in order to really benefit from this, like, extra thinking.”
Assertion Not checkable as stated
Brown: OpenAI's o3 Gets 'Not Very Far' Playing Pokémon Unharnessed
“How far does O three get without any harness? How far does it get playing Pokemon? And the answer is like, not very far, you know?”
Insight
Brown: Data for reinforcement fine-tuning survives future model scaling
“I think the difference is that like for reinforcement fine tuning, you're collecting data that's going to be useful As the models improve as well. So if we come out with, like, future models that are even more capable, you could still fine tune them on your da…”
Assertion Not checkable as stated
Noam Brown: OpenAI succeeded early by betting on scaling over small experiments
“One of OpenAI's big success was betting on the scaling paradigm. It is just kind of odd because, you know, they were not the biggest lab, you know, it was, like, difficult for them to scale. Back then, it was much more common to do, like, a lot of small experi…”
Disclosure
Noam Brown: OpenAI o3 has basically replaced Google Search for me
“Like I've been using it day to day. It's basically replaced Google search for me. Like I just use it all the time.”
Prediction Held up
OpenAI's technology will surpass o3 within six months
“I think that Oh, three is not where the technology will be in six months.”
Disclosure
Brown: OpenAI team is scaling test-time compute to hours and days
“The team, in many ways, is actually a misnomer, because we're working on more than just multi-agent. Multi-agent is one of the things we're working on. Some other things we're working on is just like being able to scale up test time compute by a ton. So how, y…”
Prediction Not checkable as stated
Noam Brown: A Superhuman Magic: The Gathering AI Is Feasible Today
“And my guess is that if somebody put in the effort, they could probably make a superhuman bot for Magic the Gathering now.”
Insight
Brown: Debugging game AI requires deep mastery to spot novel brilliance
“When you work on these games, you kind of have to understand the game well enough to like be able to debug your bot because If the bot does something that's, like, really radical and, like, that humans typically wouldn't do, you're not sure if that's, like, a …”
What-if
Noam Brown: Reasoning paradigms would have failed on GPT-2
“If you try to do the reasoning paradigm on top of GPT-II, I don't think it would have gotten you almost anything.”