The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 30 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Opinion
Brown: LLMs implicitly develop world models through scale alone
“I think it's pretty clear that as these models get bigger, they have a world model, and that world model becomes better with scale. So they are implicitly developing a world model, and I don't think it's something that you need to explicitly model.”
Noam Brown Jun 19, 2025 ▶ 52:30 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Opinion
Brown: AI models implicitly develop theory of mind through scale
“If these models become smart enough, they develop things like theory of mind. They develop an understanding that there are other agents that like can take actions and have motives and all this stuff. And these models just develop that implicitly with scale and…”
Noam Brown Jun 19, 2025 ▶ 53:31 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Assertion Not checkable as stated
Brown: GPT-4o and o3 are passing the Turing test
“So at this point, like, you know, the truth is, you know, GPT-IV-O and like O-III, these models are like passing the Turing test.”
Noam Brown Jun 19, 2025 ▶ 3:28 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Insight
Brown: Deep Research proves reasoning models work in unverifiable domains
“And that is very clearly a domain where you don't have an easily verifiable metric for success. It's very like, what is the best research report that you could generate? And yet these models are doing extremely well at this domain. So I think that's like an ex…”
Noam Brown Jun 19, 2025 ▶ 7:32 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Insight
Noam Brown: The Ideal AI Agent Harness Is No Harness
“The ideal harness is no harness. Right. I think harnesses are like a crutch that eventually we're going to be able to move beyond.”
Noam Brown Jun 19, 2025 ▶ 14:03 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Prediction Not checkable as stated
Brown: Model routers will become obsolete as unified models emerge
“We've said pretty openly that we want to move to a world where there is a single unified model. And in that world, you shouldn't need a router on top of the model. So I think that the router issue Will eventually be solved also.”
Noam Brown Jun 19, 2025 ▶ 19:00 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Prediction Not checkable as stated
Brown: Pre-training scaling will hit economic limits before superintelligence without reasoning
“Like, we're gonna scale it, sure, we're gonna scale these things up by a few more orders of magnitude, they're gonna become more capable, but we're not gonna see superintelligence from just that. And like, yes, if we had a quadrillion dollars to train these mo…”
Noam Brown Jun 19, 2025 ▶ 23:20 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Insight
Noam Brown: Aligned AI will outperform human virtual assistants on effort
“And so if you have an AI model that's, like, actually really aligned, To you and your preferences, then that could end up doing a way better job than a human could. Well, not, not that it's doing a better job than a human could, but like it's doing a better jo…”
Noam Brown Jun 19, 2025 ▶ 40:44 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Prediction Not checkable as stated
Brown: Multi-agent AI civilizations will far surpass current AI capabilities
“And I think that if you're able to have them cooperate and compete with billions of AIs over a long period of time and build up a civilization essentially, the things that they would be able to Produce and answer would be far beyond what is possible today with…”
Noam Brown Jun 19, 2025 ▶ 43:36 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Opinion
Brown: Previous multi-agent research was heuristic and ignored Bitter Lesson
“I think that a lot of the approaches that have been taken have been very heuristic and haven't really been following like the bitter lesson approach to scaling and research.”
Noam Brown Jun 19, 2025 ▶ 44:58 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Insight
Brown: Game Theory Optimal fails in collaborative games like Diplomacy
“Basically, when you're playing, like, the zero-sum games, like, poker, Game Theory Optimal works really well. When you're playing a game like Diplomacy, where there's, like, you need to collaborate and compete, and you need, there's room for collaboration, the…”
Noam Brown Jun 19, 2025 ▶ 49:43 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Opinion
Brown: Scaling self-play beyond zero-sum games will not be as easy as AlphaGo
“My point is that, like, this is where the AlphaGo analogy breaks down. And, not necessarily breaks down, but, like, it's not going to be as easy as self-play was in AlphaGo.”
Noam Brown Jun 19, 2025 ▶ 58:29 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Insight
Brown: Reasoning models improve primarily through compute efficiency rather than longer thinking duration
“These models are becoming more efficient in the way they're thinking, as they're able to do more with the same amount of test time compute, and I think that's a very underappreciated point, that it's not just that we're getting these models to think for longer…”
Noam Brown Jun 19, 2025 ▶ 1:08:26 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Prediction Not checkable as stated
Brown: Reasoning models will progress rapidly into agentic behavior
“I think that we're going to continue to see, as I said before, that we're going to see this paradigm continue to progress rapidly. And I think that that's true even today, that we saw that with like going from O-one preview to O-one to O-three, consistent prog…”
Noam Brown Jun 19, 2025 ▶ 6:07 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Insight
Noam Brown: Models need baseline capabilities to benefit from test-time reasoning
“One thing that I think is underappreciated is that the models, the pre-trained models need a certain level of capability in order to really benefit from this, like, extra thinking.”
Noam Brown Jun 19, 2025 ▶ 9:22 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Assertion Not checkable as stated
Brown: OpenAI's o3 Gets 'Not Very Far' Playing Pokémon Unharnessed
“How far does O three get without any harness? How far does it get playing Pokemon? And the answer is like, not very far, you know?”
Noam Brown Jun 19, 2025 ▶ 14:33 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Insight
Brown: Data for reinforcement fine-tuning survives future model scaling
“I think the difference is that like for reinforcement fine tuning, you're collecting data that's going to be useful As the models improve as well. So if we come out with, like, future models that are even more capable, you could still fine tune them on your da…”
Noam Brown Jun 19, 2025 ▶ 21:26 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Assertion Not checkable as stated
Noam Brown: OpenAI succeeded early by betting on scaling over small experiments
“One of OpenAI's big success was betting on the scaling paradigm. It is just kind of odd because, you know, they were not the biggest lab, you know, it was, like, difficult for them to scale. Back then, it was much more common to do, like, a lot of small experi…”
Noam Brown Jun 19, 2025 ▶ 32:07 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Disclosure
Noam Brown: OpenAI o3 has basically replaced Google Search for me
“Like I've been using it day to day. It's basically replaced Google search for me. Like I just use it all the time.”
Noam Brown Jun 19, 2025 ▶ 35:46 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Prediction Held up
OpenAI's technology will surpass o3 within six months
“I think that Oh, three is not where the technology will be in six months.”
Noam Brown Jun 19, 2025 ▶ 38:35 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Disclosure
Brown: OpenAI team is scaling test-time compute to hours and days
“The team, in many ways, is actually a misnomer, because we're working on more than just multi-agent. Multi-agent is one of the things we're working on. Some other things we're working on is just like being able to scale up test time compute by a ton. So how, y…”
Noam Brown Jun 19, 2025 ▶ 41:59 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Prediction Not checkable as stated
Noam Brown: A Superhuman Magic: The Gathering AI Is Feasible Today
“And my guess is that if somebody put in the effort, they could probably make a superhuman bot for Magic the Gathering now.”
Noam Brown Jun 19, 2025 ▶ 1:16:47 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Insight
Brown: Debugging game AI requires deep mastery to spot novel brilliance
“When you work on these games, you kind of have to understand the game well enough to like be able to debug your bot because If the bot does something that's, like, really radical and, like, that humans typically wouldn't do, you're not sure if that's, like, a …”
Noam Brown Jun 19, 2025 ▶ 0:53 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
What-if
Noam Brown: Reasoning paradigms would have failed on GPT-2
“If you try to do the reasoning paradigm on top of GPT-II, I don't think it would have gotten you almost anything.”
Noam Brown Jun 19, 2025 ▶ 9:35 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Assertion Supported
Noam Brown: GPT-4.5 makes Tic-Tac-Toe mistakes without System 2 reasoning
“With Tic-Tac-Toe, we see that, like, GPD-Four .5 falls over. You know, it plays decently well. I shouldn't say it falls over. It does reasonably well. You can draw the board. It can make legal moves, but it will make mistakes sometimes, and if you really need …”
Noam Brown Jun 19, 2025 ▶ 12:06 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Assertion Not checkable as stated
Brown: OpenAI saw conclusive proof of its reasoning paradigm in late 2023
“I think it was around, like, November, twenty-twenty-three, or October, twenty-twenty-three, when I think I was convinced that we had, like, very conclusive signs of life, that, like, oh, this was going to be, this is the paradigm, and it's going to be a big d…”
Noam Brown Jun 19, 2025 ▶ 25:17 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Opinion
Noam Brown: Closing the human data efficiency gap is a top unsolved problem
“I think it's a fair statement to say that these models are less data efficient than humans. And I think that that's an unsolved research question and probably one of the most important unsolved research questions.”
Noam Brown Jun 19, 2025 ▶ 30:21 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Assertion Supported
Brown: Modern poker AIs stick to static GTO without player adaptation
“The way the Poker AI's work today, they're just kind of like sticking to their precomputed GTO strategy. And they're not adapting to the other players at the table.”
Noam Brown Jun 19, 2025 ▶ 51:37 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Assertion Supported
Brown won the 2025 World Diplomacy Championship
“When we released Cicero, we announced it in, like, late twenty-twenty-two, I still found the game, like, really fascinating, and so I, like, kept up with it, I, like, continued to play, and that led to me winning the championship in the World Championship in t…”
Noam Brown Jun 19, 2025 ▶ 1:32 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Disclosure
Brown: OpenAI models undergo mid-training and post-training before release
“For open AI models, like, they go through a mid-training step, and then they go through a post-training step, and then they're released, and they're a lot more useful. Like, frankly, if you interacted with the only pre-trained model, it would be super difficul…”
Noam Brown Jun 19, 2025 ▶ 1:11:38 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.