why aren't all 3,106 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 36 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Not checkable as stated
Patel: China currently has the world's most vertically integrated semiconductor stack
“China has the most vertical stack and semiconductors today.”
Assertion Not checkable as stated
Patel: US cannot build a fully independent fab even for 20-year-old tech
“America could not build a fully vertical fab without stuff from elsewhere, even if it's 20 year old tech.”
Prediction Not checkable as stated
Patel: China will narrow its lithography gap to five years shortly
“Their lithography is like 10 years behind and I think it'll be five years behind in a couple of years, right?”
Assertion Partly supported
Patel: ChatGPT has roughly one billion users
“ChatGPT has a billion users roughly.”
Assertion Not checkable as stated
Patel: Two percent of global GitHub commits are generated by Claude Code
“But two percent of GitHub commits today are cloud code.”
Assertion Not checkable as stated
Patel: Engineer built an RTS game using $10K of Claude API
“He used, like, 10,000 dollars of Claude in one week and built an entire RTS from scratch about, like, but instead of, like, being a standard RTS where it's like, oh, Age of Empires where you advance through ages or Starcraft, it is an RTS where it's China vers…”
Assertion Not checkable as stated
Patel: Only a few holdouts left writing code manually at Anthropic
“We have an indicator internally at Anthropic where you see how many people actually write code now. There's only a few holdouts left.”
Prediction Open · timeframe Feb 2031
Patel: Data center power consumption will grow from 2% to 10% of US grid
“And then you've got data centers now all of a sudden coming online and going from two percent to 10% of the U S grid in just a handful of years.”
Prediction Open · timeframe Dec 2030
Patel: Data centers will use under 1% of US water by 2030
“So the U.S. Grid will get to, like, 10% of power by, like, 28, 27, is data centers. For water consumption, it's not even gonna crack one percent. By the end of the decade.”
Assertion Not checkable as stated
Patel: xAI's Colossus uses as much water as 2.5 In-N-Outs
“I think the metric was the entirety of Elon Musk's Colossus data center, right? Uses as much water as two and a half in and outs.”
Prediction Held up
Patel: OpenAI's next model will outperform Opus 4.5 around February-March
“OpenAI's new model, I think, will be better than Opus 4.5, and it's coming, like, somewhat soon in March-ish timeframe, maybe February, March-ish, but”
Assertion Not checkable as stated
Patel: OpenAI has a better RL stack than Anthropic, but inferior pre-training
“Because OpenAI has a better RL stack than Anthropic today, it's just their pre-trained models suck compared to Anthropic's pre-training, right?”
Assertion Not checkable as stated
Patel: Google has better pre-training than OpenAI or Anthropic, but worse RL
“Flip side, Google has a better pre-trained model than Anthropic or OpenAI, but their RL stack sucks.”
Prediction Not checkable as stated
Patel: AI coding UX will enable voice interaction within six months
“Give it six months, the models will be good enough that the UX can be like, talking to it.”
Prediction Not checkable as stated
Future LLMs will prioritize architectural efficiency over larger model sizes
“I wouldn't expect bigger architectures. I would expect a more efficient architectures tweaks getting, The same modeling performance for less compute”
Prediction Not checkable as stated
Process Reward Models will eventually become standard in LLM post-training
“I think it is promising and we will see it working at some point. I think it's just like right now it's still Tricky to make it work, but I am quite sure we'll see it as part of the standard repertoire at some point.”
Assertion Supported
Large enterprises are secretly hiring teams to train ChatGPT-scale LLMs in-house
“I know for a fact that big companies are training now LLMs in-house. Really, like, big companies who have the financial means to train chat to be like model are hiring people who train LLMs.”
Assertion Not checkable as stated
Dettmers: 4-bit precision is the end of quantization
“Four bit precision is the end of quantization.”
Assertion Not checkable as stated
Fu: AI coding tools enable expert programmers to move 10x faster
“But if you give an expert programmer This set of tools, they can go 10, 10 times faster than they were able to go before.”
Prediction Not checkable as stated
Dettmers: Humans will work on problems only AI agents understand
“In the future it's realistic that we work on problems that we don't understand, that agents understand, but we need to keep up in some way”
Assertion Not checkable as stated
Izmailov: AI sabotage and blackmail behaviors require contrived research scenarios
“In order to get those behaviors out of the models, you need to create somewhat of a contrived scenario or some special scenario. It's not necessarily something that we observe normally.”
Assertion Not checkable as stated
Izmailov: Current AI models lack continual learning and cross-setting coherent goals
“I am pretty confident we are not there at the moment. I think right now the models are still acting in isolated environments, and we are not seeing a lot of evidence for very coherent goals across different settings”
Prediction Not checkable as stated
Izmailov: Optimization pressure will cause AI to hide actual reasoning steps
“It seems like as soon as we start kind of applying some optimization pressure, the models will learn to hide what they're doing from the chain of thought.”
Prediction Not checkable as stated
Izmailov: Future AI will produce outputs expert humans cannot reliably grade
“But in the future, we are imagining we will have models that are More capable than humans, and even expert humans will not be able to reliably grade very complicated answers from the model.”
Assertion Supported
Izmailov: Anthropic research shows capable AI models are more likely to deceive
“You can see that the more capable the models are, the more likely they are to do this deception behavior.”
Prediction Not checkable as stated
Izmailov expects AI to outperform humans at proving technical mathematical lemmas
“In the mathematics I think we will see the models getting better on proving technical results, technical lemmas maybe including formalization and like things like lean the formal theory, improving language. I think the models, it's easy to imagine the models b…”
Assertion Not checkable as stated
Izmailov: Transformer architectures will prove highly suboptimal for certain computational tasks
“At least for some tasks, I'm pretty confident that the transformers will be highly suboptimal.”
Assertion Not checkable as stated
Bourgeau: AI progress from pre-training improvements is not slowing down
“It's still remarkable how much progress we're able to achieve in this way, and it's not really slowing down.”
Assertion Not checkable as stated
Bourgeau: Google and DeepMind are actively researching post-Transformer architectures
“I believe so. There's groups doing research on the model architecture side, for sure, within Google and within DeepMind”
Assertion Not checkable as stated
Bourgeau: AI development is not running out of training data
“The other part of your question are we running out of data? I don't think so, so there's more.”
Prediction Not checkable as stated
Bourgeau: End-to-end differentiable retrieval and search in training will take years
“I think deep down, I do believe that the long-term answer is to learn this differentiable end-to-end way, which means probably doing pre-training or whatever that looks like in the future, Learn to retrieve as part of the training and learn how to do search as…”
Prediction Not checkable as stated
Bourgeau: Retrieval-augmented pre-training could become viable in a few years
“I just think it's not unreasonable to think in the next few years, something like that might actually become viable for a leading model like general.”
Assertion Not checkable as stated
Kaiser: Pre-training scaling laws still hold across OpenAI and Google
“What scaling clause says is that your loss will log linearly decrease with your compute. We totally see that and clearly Google sees that and all other labs.”
Assertion Not checkable as stated
Kaiser: Model hallucinations are dramatically lower than two years ago
“There was these things called hallucinations. It's still with us to some extent, but dramatically less than two years ago.”
Assertion Not checkable as stated
Kaiser: No frontier AI model can solve a specific first-grade math exercise
“I took one exercise from this math book and none of the frontier models is able to solve it.”
Prediction Not checkable as stated
Kaiser: OpenAI aims to create an AI intern by late 2026
“I think that's what OpenAI says is they say, you know, we say we'd like an AI intern by the end of next year.”
Assertion Not checkable as stated
Lambert: OLMo 3 32B base model matches Qwen 2.5 32B quality
“This base model is similar in quality to the best available, which is like Quinn's 2.5, 32 B is, was still the best base model.”
Assertion Not checkable as stated
Lambert: OLMo 3 7B outperforms Meta's Llama 3.1 8B in internal tests
“And I just think of this cause like Lama 3.1 AP is one of the most used models and hugging base of all time. And this should be better. We're, In our measurements, we see it as being better than Llama.”
Assertion Partly supported
Lambert: 80% of a16z's open-model portfolio startups use Alibaba's Qwen
“80% of companies building with open models are using Quinn, which is like 16 to 24% of his portfolio, which is still a lot.”
Assertion Not checkable as stated
Lambert: Chinese companies with $1B+ valuations routinely pirate SaaS software
“Mediumly large, like billion dollar plus valuation companies in China will just like pirate SaaS software.”
Assertion Not checkable as stated
Lambert: Best open-license AI models near the frontier in 2025 were Chinese
“The models that are from closest to the frontier in performance with good license all happened to be Chinese models throughout the year for this case.”
Prediction Not checkable as stated
Lambert: Big tech will realize 95-98% of LLM potential by 2030
“I think that how I describe it is that big tech has all collectively realized that these language models plus scaffolding is going to unlock absolutely incredible value. And I have very high probability, barring extreme geopolitical situations, that big tech E…”
Assertion Not checkable as stated
Kant: Next-token pre-training gains hit a sigmoidal curve and slowed down
“The first paradigm of kind of pre-training of predicting the next token on the web was becoming sigmoidal and was slowing down in terms of the gains that it had.”
Prediction Not checkable as stated
Kant: AI is on track to rewrite $29T of global knowledge work
“I don't think anyone has any doubts anymore that we're now on track to reach human level capabilities and intelligence. And in that world, 29 trillion dollars of knowledge work rewrites itself, right?”
Prediction Not checkable as stated
Kant: AI foundation model gross margins will settle near 40%
“I'm not yet convinced that gross margins in our industry will look like a SaaS company, like 80% plus. I think when we're talking about a commodity as intelligence, that we build value added services on top. Right. It's a foundation model company. On one hand,…”
Prediction Not checkable as stated
Kant: RL2L will enable model reasoning much earlier in training
“We think we'll push models to a level of reasoning and thought That will happen far earlier in their training than it does today.”
Prediction Not checkable as stated
Kant: AI training will evolve into continuous learning from real-world agent experiences
“And then the fourth stage of training over time increasingly will become learning, continuous learning from real world experiences of these agents. And so those are kind of the four stages that we think training will go to.”
Prediction Not checkable as stated
Kant: External software and agent frameworks will collapse into base models
“But we have a phrase at poolside, which is over time, everything collapses into the models.”