The Ledger, every show

Every statement that passed quotation and attribution checks, across all 44 shows. Pick shows below, then mix any filter with any other.

shows every show 44 of 44
every show
clear all ✕
NO PRIORS Prediction Not checkable as stated
Overnight AI intelligence explosion unlikely due to test-time compute bottlenecks
“And I don't think we're headed to that world largely because of the fact that the models rely so much on large scale test time compute. In order to achieve their greatest intelligence. If you, if it requires so much test time compute to unlock the full capabil…”
Noam Brown Jun 26, 2026 ▶ 26:21 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
NO PRIORS Assertion Open · timeframe Jun 2027
Brown: Modern AI models can reason for weeks before plateauing
“What we're seeing today with the modern models is that 5.5 and other models can think for, if you scaffold them reasonably well, can think for weeks even before having performance plateau on some of these benchmarks.”
Noam Brown Jun 26, 2026 ▶ 3:34 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
NO PRIORS Opinion
Noam Brown: AI Model Outputs Are Arguably More Trustworthy Than Humans
“I use it day to day for a lot of this kind of stuff, and I think they're at a point now where They've actually been at a point for a while now where I feel like I can just trust the outputs, arguably more than I could trust the output from a human.”
Noam Brown Jun 26, 2026 ▶ 31:36 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
Brown: LLMs implicitly develop world models through scale alone
“I think it's pretty clear that as these models get bigger, they have a world model, and that world model becomes better with scale. So they are implicitly developing a world model, and I don't think it's something that you need to explicitly model.”
Noam Brown Jun 19, 2025 ▶ 52:30 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Brown: AI models implicitly develop theory of mind through scale
“If these models become smart enough, they develop things like theory of mind. They develop an understanding that there are other agents that like can take actions and have motives and all this stuff. And these models just develop that implicitly with scale and…”
Noam Brown Jun 19, 2025 ▶ 53:31 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
NO PRIORS Insight
Inference-time compute is the missing scaling dimension for AI reasoning
“This is why I'm interested in the reasoning direction, because I think there's this whole other dimension. That people are not scaling right now, which is the amount of compute at inference time.”
Noam Brown Apr 25, 2023 ▶ 17:12 No Priors Ep. 1 | With Noam Brown, Research Scientist at Meta
NO PRIORS Insight
Brown: AI benchmarks must control for test-time compute
“And so I think the proper way to, and so my claim is the proper way to evaluate the models now is you either have some kind of budget for the benchmark, whether it's tokens or cost or time or whatever, or you plot the performance as a function of the amount of…”
Noam Brown Jun 26, 2026 ▶ 4:01 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
NO PRIORS Insight
Brown: Scaffolding Easily Inflates AI Benchmark Scores Without Real Gains
“It's really easy to show you can do much better than previous benchmarks or previous, previous models on benchmarks by just, for example, scaffolding a bunch of models together. So if you say, okay, well, we're going to, instead of just running this model once…”
Noam Brown Jun 26, 2026 ▶ 7:03 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
NO PRIORS Prediction Not checkable as stated
Brown predicts AI will zero-shot his entire PhD thesis within one year
“And I wouldn't be surprised if, you know, six months or a year from now, the model is able to do zero shot an entire poker solver, basically my entire PhD thesis in one go.”
Noam Brown Jun 26, 2026 ▶ 11:17 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
NO PRIORS Insight
Current AI safety frameworks fail to account for test-time compute scaling
“The preparedness frameworks and responsible scaling policies, they don't really account for the amount of tests I'm computed. They just say, okay, well, what's the capability of the model? The problem is we're in a world now where the capability of the model i…”
Noam Brown Jun 26, 2026 ▶ 12:55 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
NO PRIORS Assertion Supported
OpenAI internal model reportedly disproved the Erdős unit distance conjecture
“We used an internal model at OpenAI a few weeks ago to disprove the unit Erdos unit distance conjecture.”
Noam Brown Jun 26, 2026 ▶ 17:20 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
NO PRIORS Insight
Complex task compute costs fall 10x to 100x per model release
“The model release cycle is every, every couple months we put out a new model that's even more powerful, and so the cost of disproving the Erdos unit distance gesture drops by, like, 10 or a hundred x with every model release cycle. Probably, in some cases, mor…”
Noam Brown Jun 26, 2026 ▶ 19:28 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
NO PRIORS Disclosure
OpenAI discourages researchers from using current models on open math problems
“We are trying to encourage people to not spend all their time just, like, Going through all the mathematical open problems, physics problems, and just seeing, pushing the models to their limits to see what they can prove or disprove. Because we really think th…”
Noam Brown Jun 26, 2026 ▶ 20:14 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
NO PRIORS Assertion Not checkable as stated
Brown: AI cannot invent novel algorithms better than existing research
“Go ahead and like look at all the published work and synthesize that and then try to come up with something novel and it's not able to do it. And I can give it a lot of time and it's still not able to do it.”
Noam Brown Jun 26, 2026 ▶ 24:30 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
NO PRIORS Insight
Noam Brown: AI community stuck in bad equilibrium publishing static benchmark grids
“I would talk to researchers about we, it makes sense to show the benchmarks with an x-axis, whether it's tokens or cost or time, there should be an x-axis, and everybody would say, like, yeah, that makes sense, we should do that, but. Well, really, their respo…”
Noam Brown Jun 26, 2026 ▶ 32:27 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
LATENT SPACE Assertion Not checkable as stated
Brown: GPT-4o and o3 are passing the Turing test
“So at this point, like, you know, the truth is, you know, GPT-IV-O and like O-III, these models are like passing the Turing test.”
Noam Brown Jun 19, 2025 ▶ 3:28 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Brown: Deep Research proves reasoning models work in unverifiable domains
“And that is very clearly a domain where you don't have an easily verifiable metric for success. It's very like, what is the best research report that you could generate? And yet these models are doing extremely well at this domain. So I think that's like an ex…”
Noam Brown Jun 19, 2025 ▶ 7:32 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Noam Brown: The Ideal AI Agent Harness Is No Harness
“The ideal harness is no harness. Right. I think harnesses are like a crutch that eventually we're going to be able to move beyond.”
Noam Brown Jun 19, 2025 ▶ 14:03 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
LATENT SPACE Prediction Not checkable as stated
Brown: Model routers will become obsolete as unified models emerge
“We've said pretty openly that we want to move to a world where there is a single unified model. And in that world, you shouldn't need a router on top of the model. So I think that the router issue Will eventually be solved also.”
Noam Brown Jun 19, 2025 ▶ 19:00 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
LATENT SPACE Prediction Not checkable as stated
Brown: Pre-training scaling will hit economic limits before superintelligence without reasoning
“Like, we're gonna scale it, sure, we're gonna scale these things up by a few more orders of magnitude, they're gonna become more capable, but we're not gonna see superintelligence from just that. And like, yes, if we had a quadrillion dollars to train these mo…”
Noam Brown Jun 19, 2025 ▶ 23:20 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Noam Brown: Aligned AI will outperform human virtual assistants on effort
“And so if you have an AI model that's, like, actually really aligned, To you and your preferences, then that could end up doing a way better job than a human could. Well, not, not that it's doing a better job than a human could, but like it's doing a better jo…”
Noam Brown Jun 19, 2025 ▶ 40:44 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
LATENT SPACE Prediction Not checkable as stated
Brown: Multi-agent AI civilizations will far surpass current AI capabilities
“And I think that if you're able to have them cooperate and compete with billions of AIs over a long period of time and build up a civilization essentially, the things that they would be able to Produce and answer would be far beyond what is possible today with…”
Noam Brown Jun 19, 2025 ▶ 43:36 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Brown: Previous multi-agent research was heuristic and ignored Bitter Lesson
“I think that a lot of the approaches that have been taken have been very heuristic and haven't really been following like the bitter lesson approach to scaling and research.”
Noam Brown Jun 19, 2025 ▶ 44:58 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Brown: Game Theory Optimal fails in collaborative games like Diplomacy
“Basically, when you're playing, like, the zero-sum games, like, poker, Game Theory Optimal works really well. When you're playing a game like Diplomacy, where there's, like, you need to collaborate and compete, and you need, there's room for collaboration, the…”
Noam Brown Jun 19, 2025 ▶ 49:43 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Brown: Scaling self-play beyond zero-sum games will not be as easy as AlphaGo
“My point is that, like, this is where the AlphaGo analogy breaks down. And, not necessarily breaks down, but, like, it's not going to be as easy as self-play was in AlphaGo.”
Noam Brown Jun 19, 2025 ▶ 58:29 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Brown: Reasoning models improve primarily through compute efficiency rather than longer thinking duration
“These models are becoming more efficient in the way they're thinking, as they're able to do more with the same amount of test time compute, and I think that's a very underappreciated point, that it's not just that we're getting these models to think for longer…”
Noam Brown Jun 19, 2025 ▶ 1:08:26 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
NO PRIORS Opinion
The Turing test is no longer a useful measure for AI
“I think the Turing test is no longer really a useful measure the way it was intended to be. Certainly Just because we have bots that can, I wouldn't say they can pass the Turing test, but I mean, like they're getting close enough that it's no longer that usefu…”
Noam Brown Apr 25, 2023 ▶ 12:18 No Priors Ep. 1 | With Noam Brown, Research Scientist at Meta
NO PRIORS Opinion
Data availability is not the true bottleneck for AI scaling
“It's not clear that data really is the bottleneck on performance here. And I've talked to AI researchers about this, and I think there isn't as much of a worry about this as people might think. Probably that's because there's a lot more data that's out there t…”
Noam Brown Apr 25, 2023 ▶ 15:54 No Priors Ep. 1 | With Noam Brown, Research Scientist at Meta
NO PRIORS Assertion Supported
Supervised learning on human games fails to produce expert players
“Like we also found in chess and go, we actually ran this experiment. If you do. Just pure supervised learning on a giant data set of human chess and go games. The bot that you get out from that is not an expert chess or go player. Even if it's like conditioned…”
Noam Brown Apr 25, 2023 ▶ 20:08 No Priors Ep. 1 | With Noam Brown, Research Scientist at Meta
NO PRIORS Insight
Reinforcement learning struggles in trading because financial markets are non-stationary
“I think the major challenge with Using things like reinforcement learning for trading is that it's a non-stationary environment. So you can have all this historical data, but it's not a stationary system and it's gonna like the markets respond to world events,…”
Noam Brown Apr 25, 2023 ▶ 30:54 No Priors Ep. 1 | With Noam Brown, Research Scientist at Meta
NO PRIORS Prediction Not checkable as stated
AI models could likely beat humans today in constrained business negotiations
“I think if you were to look at constrained domains certain negotiation tasks, I think that AIs could probably do better than humans in that today. I mean, I'm trying to think of like specific examples, but things like you know, if you wanted to negotiate over …”
Noam Brown Apr 25, 2023 ▶ 32:34 No Priors Ep. 1 | With Noam Brown, Research Scientist at Meta
NO PRIORS Prediction Not checkable as stated
An AI model could prove the Riemann hypothesis by 2028
“You know, it doesn't seem crazy to me that you could have a model that can prove the Riemann hypothesis within the next five years. If you can solve the reasoning problem in a truly general way.”
Noam Brown Apr 25, 2023 ▶ 35:45 No Priors Ep. 1 | With Noam Brown, Research Scientist at Meta
NO PRIORS Assertion Supported
Brown: GPT-5.5 is far more compute-efficient than GPT-5.4
“It turned out that 5.5 is just much more efficient with its thinking. If you run it at max settings, 5.4 is thinking for a lot longer. It takes longer to get back a response than 5.5. And once you control for the amount of thinking time, actually you can see t…”
Noam Brown Jun 26, 2026 ▶ 2:43 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
NO PRIORS Assertion Supported
Brown: AISI evals show AI cyber capabilities improve past 100M tokens
“Actually the AISI in their evaluations has shown that the models continue to improve at A hundred million tokens. You know, if you run them for a hundred million tokens, they're still improving at beyond that point.”
Noam Brown Jun 26, 2026 ▶ 4:41 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
NO PRIORS Insight
Brown: Long AI Deliberation Time Is Impractical for Real Workflows
“This idea that the models, you just let them think for a week or whatever, and then they respond, it's, it sounds nice, and yes, the benchmarks look great, but it's not very practical when working because like, okay, you ask the model a question, and then you …”
Noam Brown Jun 26, 2026 ▶ 6:13 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
NO PRIORS Assertion Supported
Brown: Modern AI Models Can Run Scaffolded Experiments for Months
“We're seeing now with the most recent models that you can actually scaffold, for example, 5.5 into doing a series of experiments that can run for weeks, for months.”
Noam Brown Jun 26, 2026 ▶ 15:03 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
NO PRIORS Insight
Brown: Rapid AI Release Cycles Obscure True Model Capability Ceilings
“The model release cycle is, look, we're releasing new models, like, every two or three months at this point, and so a model comes out, it takes two or three months to push it to its limits, and then you have another model come out, and so nobody actually knows…”
Noam Brown Jun 26, 2026 ▶ 16:10 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
NO PRIORS Assertion Supported
Brown: GPT-5.5 can derive Erdős disproof with proper scaffolding
“After we announced the results, A bunch of people found that you could get the answer out of 5.5 as well. If, now, it's not as simple as just asking 5.5, hey, here's the Irish unit distance conjecture. What's the disproof? You had to scaffold it a bit. You had…”
Noam Brown Jun 26, 2026 ▶ 18:01 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
NO PRIORS Insight
Brown: Extra test-time compute does not improve factual retrieval in AI
“There are some benchmarks where the models will just not improve if they have more inference budget. So I think a lot of factual factual retrieval kind of questions fall into this category of if you ask a person when was Abraham Lincoln born and they don't kno…”
Noam Brown Jun 26, 2026 ▶ 21:43 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
NO PRIORS Insight
Noam Brown: AI Models Cannot Organically Accumulate Shared Knowledge Today
“We're not seeing that with AI models today. They kind of, they're born into a world for, and they exist for a very short context window, and then they just, like, disappear. And yeah, there are things that you can kind of do to, like, continue them, but it's v…”
Noam Brown Jun 26, 2026 ▶ 28:26 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
LATENT SPACE Prediction Not checkable as stated
Brown: Reasoning models will progress rapidly into agentic behavior
“I think that we're going to continue to see, as I said before, that we're going to see this paradigm continue to progress rapidly. And I think that that's true even today, that we saw that with like going from O-one preview to O-one to O-three, consistent prog…”
Noam Brown Jun 19, 2025 ▶ 6:07 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Noam Brown: Models need baseline capabilities to benefit from test-time reasoning
“One thing that I think is underappreciated is that the models, the pre-trained models need a certain level of capability in order to really benefit from this, like, extra thinking.”
Noam Brown Jun 19, 2025 ▶ 9:22 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
LATENT SPACE Assertion Not checkable as stated
Brown: OpenAI's o3 Gets 'Not Very Far' Playing Pokémon Unharnessed
“How far does O three get without any harness? How far does it get playing Pokemon? And the answer is like, not very far, you know?”
Noam Brown Jun 19, 2025 ▶ 14:33 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Brown: Data for reinforcement fine-tuning survives future model scaling
“I think the difference is that like for reinforcement fine tuning, you're collecting data that's going to be useful As the models improve as well. So if we come out with, like, future models that are even more capable, you could still fine tune them on your da…”
Noam Brown Jun 19, 2025 ▶ 21:26 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
LATENT SPACE Assertion Not checkable as stated
Noam Brown: OpenAI succeeded early by betting on scaling over small experiments
“One of OpenAI's big success was betting on the scaling paradigm. It is just kind of odd because, you know, they were not the biggest lab, you know, it was, like, difficult for them to scale. Back then, it was much more common to do, like, a lot of small experi…”
Noam Brown Jun 19, 2025 ▶ 32:07 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
LATENT SPACE Disclosure
Noam Brown: OpenAI o3 has basically replaced Google Search for me
“Like I've been using it day to day. It's basically replaced Google search for me. Like I just use it all the time.”
Noam Brown Jun 19, 2025 ▶ 35:46 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
LATENT SPACE Prediction Held up
OpenAI's technology will surpass o3 within six months
“I think that Oh, three is not where the technology will be in six months.”
Noam Brown Jun 19, 2025 ▶ 38:35 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
LATENT SPACE Disclosure
Brown: OpenAI team is scaling test-time compute to hours and days
“The team, in many ways, is actually a misnomer, because we're working on more than just multi-agent. Multi-agent is one of the things we're working on. Some other things we're working on is just like being able to scale up test time compute by a ton. So how, y…”
Noam Brown Jun 19, 2025 ▶ 41:59 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.