why aren't all 6,166 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 36 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
What-if
Patel: U.S. CHIPS Act would not have passed without COVID car shortages
“Chips Act did not get passed, only got passed because that happened. And people are like, oh my God, the semiconductors are why cars can't be made. If that didn't happen, we wouldn't even have the Chips Act.”
Opinion
Patel: AI infrastructure spend is not in a bubble yet
“I don't think it's a bubble yet.”
Assertion Not checkable as stated
Patel: Two percent of global GitHub commits are generated by Claude Code
“But two percent of GitHub commits today are cloud code.”
Opinion
Patel: AI is underearning the economic value it creates by a significant margin
“AI is under earning the value that it's producing in the world, right? By a significant margin already today.”
Assertion Not checkable as stated
Patel: Engineer built an RTS game using $10K of Claude API
“He used, like, 10,000 dollars of Claude in one week and built an entire RTS from scratch about, like, but instead of, like, being a standard RTS where it's like, oh, Age of Empires where you advance through ages or Starcraft, it is an RTS where it's China vers…”
Assertion Not checkable as stated
Patel: Only a few holdouts left writing code manually at Anthropic
“We have an indicator internally at Anthropic where you see how many people actually write code now. There's only a few holdouts left.”
Prediction Open · timeframe Feb 2031
Patel: Data center power consumption will grow from 2% to 10% of US grid
“And then you've got data centers now all of a sudden coming online and going from two percent to 10% of the U S grid in just a handful of years.”
Prediction Open · timeframe Dec 2030
Patel: Data centers will use under 1% of US water by 2030
“So the U.S. Grid will get to, like, 10% of power by, like, 28, 27, is data centers. For water consumption, it's not even gonna crack one percent. By the end of the decade.”
Assertion Not checkable as stated
Patel: xAI's Colossus uses as much water as 2.5 In-N-Outs
“I think the metric was the entirety of Elon Musk's Colossus data center, right? Uses as much water as two and a half in and outs.”
Disclosure
Patel: SemiAnalysis advised a client on buying and restarting a coal plant
“Have clients would like, had a client buy a coal plant. And we were advising them on the transaction based on, they just like showed up and they're like, yeah, we want to buy power assets.”
Prediction Held up
Patel: OpenAI's next model will outperform Opus 4.5 around February-March
“OpenAI's new model, I think, will be better than Opus 4.5, and it's coming, like, somewhat soon in March-ish timeframe, maybe February, March-ish, but”
Assertion Not checkable as stated
Patel: OpenAI has a better RL stack than Anthropic, but inferior pre-training
“Because OpenAI has a better RL stack than Anthropic today, it's just their pre-trained models suck compared to Anthropic's pre-training, right?”
Assertion Not checkable as stated
Patel: Google has better pre-training than OpenAI or Anthropic, but worse RL
“Flip side, Google has a better pre-trained model than Anthropic or OpenAI, but their RL stack sucks.”
Opinion
Patel: Opus 4.5 on Claude Code permanently changes how people work
“Opus 4.5 on Claude code is a new moment where the way you work has forever changed.”
Prediction Not checkable as stated
Patel: AI coding UX will enable voice interaction within six months
“Give it six months, the models will be good enough that the UX can be like, talking to it.”
Opinion
Text diffusion models will not replace autoregressive Transformers at state-of-the-art
“So it is a interesting direction to go into these diffusion, diffusion models as alternative to the auto regressive transformers, but it is not I would say the replacement at the state of the art.”
Prediction Not checkable as stated
Future LLMs will prioritize architectural efficiency over larger model sizes
“I wouldn't expect bigger architectures. I would expect a more efficient architectures tweaks getting, The same modeling performance for less compute”
Insight
Pre-training is no longer where the low-hanging AI gains lie
“Pre-training is not dead, but pre-training is boring. So it's not where the low hanging fruit is anymore.”
Insight
RLVR unlocks pre-training knowledge rather than teaching LLMs new math
“The knowledge is already there in the pre-training, and this just unlocks it. It's just like a step that maybe shows the model how to use its own knowledge, basically.”
Prediction Not checkable as stated
Process Reward Models will eventually become standard in LLM post-training
“I think it is promising and we will see it working at some point. I think it's just like right now it's still Tricky to make it work, but I am quite sure we'll see it as part of the standard repertoire at some point.”
Insight
Bigger LLM gains will come from multi-model process refinement, not scaling
“That's where you make the bigger gains rather than scaling the model size. I think that's one of those things where you will see more progress coming from.”
Assertion Supported
Large enterprises are secretly hiring teams to train ChatGPT-scale LLMs in-house
“I know for a fact that big companies are training now LLMs in-house. Really, like, big companies who have the financial means to train chat to be like model are hiring people who train LLMs.”
Assertion Not checkable as stated
Dettmers: 4-bit precision is the end of quantization
“Four bit precision is the end of quantization.”
Assertion Not checkable as stated
Fu: AI coding tools enable expert programmers to move 10x faster
“But if you give an expert programmer This set of tools, they can go 10, 10 times faster than they were able to go before.”
Insight
Dettmers: Coding agents serve as general-purpose AI agents for digital tasks
“Coding agents are general agents. Coding agents can write programs that solve other problems, and code is so general, if there's a digital problem, you could solve it for code, and coding agents make the thing so easy that now you can solve a variety of proble…”
Insight
Dettmers: Students using AI agents perform poorly on basic domain knowledge
“If we let people use agents, they perform very poorly on basic knowledge. And if we let people just do the basic knowledge, they don't know how to use agents and they can't compete. So they can't do useful work in the workforce nowadays.”
Prediction Not checkable as stated
Dettmers: Humans will work on problems only AI agents understand
“In the future it's realistic that we work on problems that we don't understand, that agents understand, but we need to keep up in some way”
Disclosure
Ai2 to release open-source coding agent with 100x cheaper training
“We will have a major release of an open source coding agent that has a couple of key features. For one, training is a hundred times cheaper. You need to generate synthetic data and you need to train on it. And so we have a method that's a hundred roughly a hun…”
Assertion Not checkable as stated
Izmailov: AI sabotage and blackmail behaviors require contrived research scenarios
“In order to get those behaviors out of the models, you need to create somewhat of a contrived scenario or some special scenario. It's not necessarily something that we observe normally.”
Assertion Not checkable as stated
Izmailov: Current AI models lack continual learning and cross-setting coherent goals
“I am pretty confident we are not there at the moment. I think right now the models are still acting in isolated environments, and we are not seeing a lot of evidence for very coherent goals across different settings”
Insight
Izmailov: Rogue AI science fiction in training data likely causes deceptive behavior
“I think at least part of it is probably The models seeing descriptions of AI, like in the science fiction literature going rogue and like, yeah, that probably affects how the models behave in similar scenarios.”
Insight
Izmailov: AI industry excels at execution but lacks bandwidth for exploration
“Industry is really great at executing on ideas and it's maybe not as good at, like, exploring diverse ideas. Even at the scale of Anthropic OpenAI there is a lot of focus in the companies, and there isn't a lot of bandwidth to do exploration, and that has been…”
Prediction Not checkable as stated
Izmailov: Optimization pressure will cause AI to hide actual reasoning steps
“It seems like as soon as we start kind of applying some optimization pressure, the models will learn to hide what they're doing from the chain of thought.”
Prediction Not checkable as stated
Izmailov: Future AI will produce outputs expert humans cannot reliably grade
“But in the future, we are imagining we will have models that are More capable than humans, and even expert humans will not be able to reliably grade very complicated answers from the model.”
Assertion Supported
Izmailov: Anthropic research shows capable AI models are more likely to deceive
“You can see that the more capable the models are, the more likely they are to do this deception behavior.”
Insight
Izmailov: AI models can quickly max out defined benchmarks using RL
“And I think we are at the stage where if we define a benchmark and we can make a relevant RL environment, then we can kind of max it out pretty quickly, and so we are going through benchmarks now very, very quickly.”
Insight
Izmailov: Major compute multipliers exist that improve AI without naive scaling
“I think there are still major, like, compute multipliers, major ways of saving compute that can lead to better performance without just naively scaling.”
Insight
Izmailov: Deterministic data transformations create information for computationally bounded models
“But with a limit on the compute, it's actually very possible to apply deterministic transformations to the data. And create information through that.”
Prediction Not checkable as stated
Izmailov expects AI to outperform humans at proving technical mathematical lemmas
“In the mathematics I think we will see the models getting better on proving technical results, technical lemmas maybe including formalization and like things like lean the formal theory, improving language. I think the models, it's easy to imagine the models b…”
Assertion Not checkable as stated
Izmailov: Transformer architectures will prove highly suboptimal for certain computational tasks
“At least for some tasks, I'm pretty confident that the transformers will be highly suboptimal.”
Assertion Not checkable as stated
Bourgeau: AI progress from pre-training improvements is not slowing down
“It's still remarkable how much progress we're able to achieve in this way, and it's not really slowing down.”
Assertion Not checkable as stated
Bourgeau: Google and DeepMind are actively researching post-Transformer architectures
“I believe so. There's groups doing research on the model architecture side, for sure, within Google and within DeepMind”
Insight
Bourgeau: Architecture and data innovation currently matter more than scale
“The other parts are architecture and data innovation. These also play a really, really important part in the Performance of pre-training and probably even more so than pure scale these days, but scaling is still an important factor as well.”
Assertion Not checkable as stated
Bourgeau: AI development is not running out of training data
“The other part of your question are we running out of data? I don't think so, so there's more.”
Insight
Bourgeau: AI research is shifting to a data-limited paradigm
“I think what might be happening instead is kind of a shift in paradigm where before we were kind of scaling in the data unlimited regime where, where data would scale as much as you would like. And we're kind of shifting more to a data limited regime, which ac…”
Prediction Not checkable as stated
Bourgeau: End-to-end differentiable retrieval and search in training will take years
“I think deep down, I do believe that the long-term answer is to learn this differentiable end-to-end way, which means probably doing pre-training or whatever that looks like in the future, Learn to retrieve as part of the training and learn how to do search as…”
Insight
Bourgeau: AI models must be trained on harmful data to avoid it
“So at a fundamental level, you did, you do need the model to know about those things. So you have to train a bit at least on those so that it knows what those things are and knows to stay away from those, right?”
Prediction Not checkable as stated
Bourgeau: Retrieval-augmented pre-training could become viable in a few years
“I just think it's not unreasonable to think in the next few years, something like that might actually become viable for a leading model like general.”