why aren't all 52 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 1 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Prediction Not checkable as stated
Overnight AI intelligence explosion unlikely due to test-time compute bottlenecks
“And I don't think we're headed to that world largely because of the fact that the models rely so much on large scale test time compute. In order to achieve their greatest intelligence. If you, if it requires so much test time compute to unlock the full capabil…”
Assertion Open · timeframe Jun 2027
Brown: Modern AI models can reason for weeks before plateauing
“What we're seeing today with the modern models is that 5.5 and other models can think for, if you scaffold them reasonably well, can think for weeks even before having performance plateau on some of these benchmarks.”
Opinion
Noam Brown: AI Model Outputs Are Arguably More Trustworthy Than Humans
“I use it day to day for a lot of this kind of stuff, and I think they're at a point now where They've actually been at a point for a while now where I feel like I can just trust the outputs, arguably more than I could trust the output from a human.”
Insight
Inference-time compute is the missing scaling dimension for AI reasoning
“This is why I'm interested in the reasoning direction, because I think there's this whole other dimension. That people are not scaling right now, which is the amount of compute at inference time.”
Insight
Brown: AI benchmarks must control for test-time compute
“And so I think the proper way to, and so my claim is the proper way to evaluate the models now is you either have some kind of budget for the benchmark, whether it's tokens or cost or time or whatever, or you plot the performance as a function of the amount of…”
Insight
Brown: Scaffolding Easily Inflates AI Benchmark Scores Without Real Gains
“It's really easy to show you can do much better than previous benchmarks or previous, previous models on benchmarks by just, for example, scaffolding a bunch of models together. So if you say, okay, well, we're going to, instead of just running this model once…”
Prediction Not checkable as stated
Brown predicts AI will zero-shot his entire PhD thesis within one year
“And I wouldn't be surprised if, you know, six months or a year from now, the model is able to do zero shot an entire poker solver, basically my entire PhD thesis in one go.”
Insight
Current AI safety frameworks fail to account for test-time compute scaling
“The preparedness frameworks and responsible scaling policies, they don't really account for the amount of tests I'm computed. They just say, okay, well, what's the capability of the model? The problem is we're in a world now where the capability of the model i…”
Assertion Supported
OpenAI internal model reportedly disproved the Erdős unit distance conjecture
“We used an internal model at OpenAI a few weeks ago to disprove the unit Erdos unit distance conjecture.”
Insight
Complex task compute costs fall 10x to 100x per model release
“The model release cycle is every, every couple months we put out a new model that's even more powerful, and so the cost of disproving the Erdos unit distance gesture drops by, like, 10 or a hundred x with every model release cycle. Probably, in some cases, mor…”
Disclosure
OpenAI discourages researchers from using current models on open math problems
“We are trying to encourage people to not spend all their time just, like, Going through all the mathematical open problems, physics problems, and just seeing, pushing the models to their limits to see what they can prove or disprove. Because we really think th…”
Assertion Not checkable as stated
Brown: AI cannot invent novel algorithms better than existing research
“Go ahead and like look at all the published work and synthesize that and then try to come up with something novel and it's not able to do it. And I can give it a lot of time and it's still not able to do it.”
Insight
Noam Brown: AI community stuck in bad equilibrium publishing static benchmark grids
“I would talk to researchers about we, it makes sense to show the benchmarks with an x-axis, whether it's tokens or cost or time, there should be an x-axis, and everybody would say, like, yeah, that makes sense, we should do that, but. Well, really, their respo…”
Opinion
The Turing test is no longer a useful measure for AI
“I think the Turing test is no longer really a useful measure the way it was intended to be. Certainly Just because we have bots that can, I wouldn't say they can pass the Turing test, but I mean, like they're getting close enough that it's no longer that usefu…”
Opinion
Data availability is not the true bottleneck for AI scaling
“It's not clear that data really is the bottleneck on performance here. And I've talked to AI researchers about this, and I think there isn't as much of a worry about this as people might think. Probably that's because there's a lot more data that's out there t…”
Assertion Supported
Supervised learning on human games fails to produce expert players
“Like we also found in chess and go, we actually ran this experiment. If you do. Just pure supervised learning on a giant data set of human chess and go games. The bot that you get out from that is not an expert chess or go player. Even if it's like conditioned…”
Insight
Reinforcement learning struggles in trading because financial markets are non-stationary
“I think the major challenge with Using things like reinforcement learning for trading is that it's a non-stationary environment. So you can have all this historical data, but it's not a stationary system and it's gonna like the markets respond to world events,…”
Prediction Not checkable as stated
AI models could likely beat humans today in constrained business negotiations
“I think if you were to look at constrained domains certain negotiation tasks, I think that AIs could probably do better than humans in that today. I mean, I'm trying to think of like specific examples, but things like you know, if you wanted to negotiate over …”
Prediction Not checkable as stated
An AI model could prove the Riemann hypothesis by 2028
“You know, it doesn't seem crazy to me that you could have a model that can prove the Riemann hypothesis within the next five years. If you can solve the reasoning problem in a truly general way.”
Assertion Supported
Brown: GPT-5.5 is far more compute-efficient than GPT-5.4
“It turned out that 5.5 is just much more efficient with its thinking. If you run it at max settings, 5.4 is thinking for a lot longer. It takes longer to get back a response than 5.5. And once you control for the amount of thinking time, actually you can see t…”
Assertion Supported
Brown: AISI evals show AI cyber capabilities improve past 100M tokens
“Actually the AISI in their evaluations has shown that the models continue to improve at A hundred million tokens. You know, if you run them for a hundred million tokens, they're still improving at beyond that point.”
Insight
Brown: Long AI Deliberation Time Is Impractical for Real Workflows
“This idea that the models, you just let them think for a week or whatever, and then they respond, it's, it sounds nice, and yes, the benchmarks look great, but it's not very practical when working because like, okay, you ask the model a question, and then you …”
Assertion Supported
Brown: Modern AI Models Can Run Scaffolded Experiments for Months
“We're seeing now with the most recent models that you can actually scaffold, for example, 5.5 into doing a series of experiments that can run for weeks, for months.”
Insight
Brown: Rapid AI Release Cycles Obscure True Model Capability Ceilings
“The model release cycle is, look, we're releasing new models, like, every two or three months at this point, and so a model comes out, it takes two or three months to push it to its limits, and then you have another model come out, and so nobody actually knows…”
Assertion Supported
Brown: GPT-5.5 can derive Erdős disproof with proper scaffolding
“After we announced the results, A bunch of people found that you could get the answer out of 5.5 as well. If, now, it's not as simple as just asking 5.5, hey, here's the Irish unit distance conjecture. What's the disproof? You had to scaffold it a bit. You had…”
Insight
Brown: Extra test-time compute does not improve factual retrieval in AI
“There are some benchmarks where the models will just not improve if they have more inference budget. So I think a lot of factual factual retrieval kind of questions fall into this category of if you ask a person when was Abraham Lincoln born and they don't kno…”
Insight
Noam Brown: AI Models Cannot Organically Accumulate Shared Knowledge Today
“We're not seeing that with AI models today. They kind of, they're born into a world for, and they exist for a very short context window, and then they just, like, disappear. And yeah, there are things that you can kind of do to, like, continue them, but it's v…”
Opinion
A five-person team could solve Settlers of Catan in a year
“To then, like, go to a game like Sotos of Catan, it just felt, like, too easy. Like, you could just take a team of five people, spend a year on that, and you'd have it cracked.”
Insight
People assume weird online text is human error before suspecting AI
“If somebody's saying something a little weird, because the bot does say weird things every once in a while, their first instinct is not going to be like, oh, I'm talking to a bot. Their first instinct is going to be like, oh, this person is like dumb or distra…”
Prediction Didn’t hold up
A $500 million AI model will likely be trained by 2025
“You can probably easily 10 X that, you know, I wouldn't be surprised if there's a five hundred million dollar model that's trained in the next year or two.”
Insight
Self-play negotiation bots trained without human data invent unintelligible languages
“If you train that bot from scratch with no human data it's going to, it could learn to negotiate, but it could learn to negotiate in a language that's not English. It could learn to negotiate in some like gibberish robot language. And then when you stick it in…”
Prediction Not checkable as stated
An AI-generated novel rivaling Harry Potter could arrive by 2028
“I don't think you can get an AI to output like the next Harry Potter just yet. That might not be that far off. Maybe it's like five years away or something. But I don't think it's happening just yet.”
Assertion Supported
Humans require orders of magnitude less data than AI to achieve mastery
“Like how many games does it take for an AI, for a human to become a good chess player or a good diplomacy player or a good artist? The answer is orders of magnitude less than it takes for an AI.”
Prediction Not checkable as stated
Next-token prediction will not replace big-company software engineers
“Like next, next token prediction is going to, is getting you surprisingly far. But I don't think it's gonna get you all the way there to like replacing, you know engineers at big companies.”
Assertion Contradicted
Inference-time search improved Noam Brown's poker AI performance by 100,000x
“If we were to add this search, this planning algorithm that would come up with a better strategy when it's actually in the hand, how much better could it do? And the answer was it improved the performance by about a 100,000 X. It was the equivalent of scaling …”
What-if
The multiplayer poker AI breakthrough was algorithms, not just scaling compute
“This wasn't just a matter of scaling compute. It really was an algorithmic breakthrough, and this kind of result would have been doable 20 years ago if people knew the approach to dig.”
Insight
Brown: Test-time AI performance scales along a continuous, projectable slope
“You also do see that like the performance is, is it's not just like a discontinuous jump. It's actually like, you can see the slope of improvement over those hundred million tokens. And so you could probably do some kind of evaluation up to a certain budget an…”
Insight
Brown: Poker bot creation is a superior AI reasoning evaluation
“I think it's a nice eval because there is very little open source code for making poker bots. And there's a lot of published essays, there's a lot of published papers on it, but you really have to reason through everything.”
Assertion Supported
Brown: GPT-3 Capabilities Could Not Scale With Test-Time Compute Budget
“Like, with GPT-III, you couldn't scale test time compute. Like, if you gave it a budget of ten million dollars and said, okay, well, let's see what GPT-III can do, it really can't do that much, more than what you could do with, like, 10 dollars or one dollar.”
Assertion Not checkable as stated
Brown says AI models optimized his PhD poker algorithms by 1,000x
“I was really impressed with the model's ability to optimize the algorithms that I had developed in my PhD. It was honestly, it was shocking to see how inefficient I was in retrospect, and they were able to make it like, you know, 1000 x faster.”
Insight
Brown: Benchmark Gains From Routing May Fail in Real-World Use
“One issue you could run into is that you could optimize for certain benchmarks with the routing and then show like, oh yeah, we see this big improvement on these benchmarks. But in real world use cases, it actually ends up not being a significant improvement.”
Insight
Computer science iterates faster than economics because building requires no permission
“If you come up with an idea, you have to get it passed through legislation and it's a very long process. Computer science is much more exciting in that way because you can just build something. You don't really need permission to do it.”
Assertion Not checkable as stated
AI was widely considered a dead field when Noam Brown began studying
“The idea of AGI was really science fiction. There were some people that were, you know serious about it, but very few, the majority opinion was that AI was, if anything, it was kind of a dead field.”
Insight
IBM's Deep Blue proved that scaling search works in AI
“We learned that scale really does work. And in that case, it wasn't scaling, you know, training and neural nets, it was scaling search.”
Assertion Supported
Meta's Cicero played 40 Diplomacy games without detection as a bot
“But surprisingly, we managed to go, like, the full 40 games without being detected as a bot.”
Assertion Supported
AlphaGo's raw neural network performs substantially below top human players
“If you take out the planning that's being done in AlphaGo and just use the raw Policy network, the raw neural network, it's actually substantially below top human performance.”
Assertion Supported
Monte Carlo Tree Search fails in imperfect-information games like poker
“And that planning algorithm that's used in AlphaGo, Monte Carlo Tree Search, is very domain specific. I think people don't appreciate just how domain specific it is because it works in chess, it works in Go, and these have been like the classic domains that pe…”
Assertion Not checkable as stated
Cicero is the first major game AI breakthrough involving cooperation
“What's really interesting about diplomacy, aside from just the natural language component, is that it really is the first major game AI breakthrough in a game that involves cooperation.”