why aren't all 2,374 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 6 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Not checkable as stated
Sachin Katti: OpenAI tripled compute and tripled revenue
“We tripled compute and we tripled revenue.”
Assertion Not checkable as stated
Katti: Demand far outstrips OpenAI's compute supply, zero goes to waste
“Demand far outstrips. Compute supply today. So anything we can bring online, we consume immediately. So there's no compute that is going to waste for us.”
Assertion Not checkable as stated
Catanzaro: Moore's Law has been economically dead for five to ten years
“The original statement of Moore's law was economic, right? It was about, we can afford to put twice as many transistors on the same chip in every, whatever, 24 months, whatever the time period is. And these days that is, Absolutely not the case. It hasn't been…”
Assertion Not checkable as stated
Prince: The same individual currently directs all Iranian state cyber attacks
“And the same guy to this day is running all of Iranian cyber attacks.”
Assertion Not checkable as stated
Prince: Meta is targeting a 50-to-1 manager span of control
“Meta is trying to get to 50 to one.”
Assertion Not checkable as stated
Balaban: The AI industry continues to underbuild compute infrastructure
“Well, I think that we continue to be generally under building.”
Assertion Not checkable as stated
Balaban: Lambda is leasing 2023-deployed H100 GPUs at higher rates today
“You actually look at the chips that we deployed in twenty-twenty-three, H-one hundreds. We're now leasing those out at a higher rate. Now than we were originally in 20, 23.”
Assertion Not checkable as stated
Combining reinforcement learning with pre-training outperforms scaling pre-training alone
“If you were just trying to scale pre-training, you wouldn't get anywhere near as far as also trying to scale RL on top of pre-training, which is what we do now.”
Assertion Not checkable as stated
Levie: CIO sentiment on AI remains optimistic, avoiding trough of disillusionment
“I think the tone is actually remarkably optimistic and excited and positive as opposed to, you know, there's a sort of, You know, typical trough of disillusionment, you know, from Gartner and the hype cycle or whatnot.”
Assertion Not checkable as stated
Levie: A single coding agent task can consume $1,000 in compute
“One, you know, coding agent could be consuming, you know, a thousand dollars of compute on a single task. So clearly like you can't lump that all into a 20 dollar per user per month fee.”
Assertion Not checkable as stated
Dubois: AI frontier labs have successfully bypassed internet data walls
“There were a lot of conversation about hitting data walls, and it seems like we did not quite hit it. So the larger the model is, the more data it needs to ingest to be trained. And it seems like different companies kind of found different ways to overcome the…”
Assertion Supported
Kolter: Adversarial prompts optimized on open-source LLMs break commercial models
“Once we had done that, we found that when you had these weird terms that you sort of flipped around to optimize one, to optimize the response for one model, you could just take those same exact strings you would optimize, paste them into a commercial model, an…”
Assertion Not checkable as stated
Anthropic's unreleased Claude Mythos model shows outsized cybersecurity capabilities
“Mythos is a unreleased frontier model. It's a general purpose model that was trained not specifically for cybersecurity or specifically for coding or specifically for software, but we have discovered what we believe to be outsized capabilities specifically in …”
Assertion Not checkable as stated
Rieseberg: AI models can execute week-long knowledge work tasks today
“The models we have today are actually quite capable. They're quite capable of running knowledge work of both of an extremely long time horizon, the kind of things that you give to someone and expect like a week later.”
Assertion Supported
Anthropic's Claude Cowork implements memory using plain text files
“It's in the harness, actually, and it's, like, often surprising to people when I talk to them how we, how we've implemented memory, because I think it maybe points at the simplicity underneath all of those models. Memory is just text files.”
Assertion Partly supported
Evans: OpenAI has 900M weekly active users, but only 5% pay
“You've got nine hundred million weekly active users, but most of them are not using it every day and can't think of anything to do with it. And only five percent of them are paying for it.”
Assertion Supported
AxiomProver achieved a perfect score on the 2025 Putnam math exam
“Eight within the time limit, and then 12 out of 12.”
Assertion Open · timeframe Feb 2029
Axiom's proof verifier is 100 times faster than open-source alternatives
“So a lot of the sort of like verify, verify proof is actually, you know, one of our prover tools that's about to be released, and that's actually a hundred times faster than The other counterparts that are the open source, like effort, cloud comparator.”
Assertion Not checkable as stated
AxiomProver autonomously proves theorems publishable in major mathematical journals
“Currently the batch of papers, Axiom Prover has autonomously proven and mathematicians have written You can probably get into Journal of Number Theory, Journal of Algebra, like that level.”
Assertion Not checkable as stated
Zeghidour: Only 50 people worldwide can train competitive voice AI models
“Between 10 and 100? No, I would say. 50? I don't know. It's hard to say. But, yeah, I think it's very few and, really meaningful contributions that have pushed the field forward have been made by very small groups of people.”
Assertion Contradicted
Zeghidour: Kyutai's Moshi remains the only full-duplex conversational AI model
“Moshi, that is still to the day the only full duplex model.”
Assertion Not checkable as stated
Zeghidour: Large multimodal models are too massive to run voice profitably
“And at the same time, these models are so large, they cannot run at scale because they will just make everyone lose money in the process.”
Assertion Not checkable as stated
Patel: Groq chips cannot cost-effectively perform general-purpose large model inference
“In a general purpose workload, crock. Grok doesn't work, right? You know, it can't train, it can't, you know, it can't inference really, really large models cost efficiently, right? You can't serve many, many, many users, but what it can do is it can go block,…”
Assertion Supported
Patel: Groq missed revenue significantly before being acquired
“In fact, they missed revenue last year significantly and yet they got bought, right? Because the value of the IP was there and the value of the team.”
Assertion Partly supported
Patel: Chinese local governments, not national, banned Nvidia's H20 and H200
“But as far as I understand, the national government has not banned Nvidia's H-twenty or H-two hundred, but the local ones have. Right. A lot of local ones have said, no, you know, you must use China manufactured chips.”
Assertion Not checkable as stated
Patel: 15 to 20 countries could single-handedly shut down semiconductor manufacturing
“I would say there's like 15 or 20 countries that can shut down the entire semiconductor industry.”
Assertion Not checkable as stated
Patel: China currently has the world's most vertically integrated semiconductor stack
“China has the most vertical stack and semiconductors today.”
Assertion Not checkable as stated
Patel: US cannot build a fully independent fab even for 20-year-old tech
“America could not build a fully vertical fab without stuff from elsewhere, even if it's 20 year old tech.”
Assertion Partly supported
Patel: ChatGPT has roughly one billion users
“ChatGPT has a billion users roughly.”
Assertion Not checkable as stated
Patel: Two percent of global GitHub commits are generated by Claude Code
“But two percent of GitHub commits today are cloud code.”
Assertion Not checkable as stated
Patel: Engineer built an RTS game using $10K of Claude API
“He used, like, 10,000 dollars of Claude in one week and built an entire RTS from scratch about, like, but instead of, like, being a standard RTS where it's like, oh, Age of Empires where you advance through ages or Starcraft, it is an RTS where it's China vers…”
Assertion Not checkable as stated
Patel: Only a few holdouts left writing code manually at Anthropic
“We have an indicator internally at Anthropic where you see how many people actually write code now. There's only a few holdouts left.”
Assertion Not checkable as stated
Patel: xAI's Colossus uses as much water as 2.5 In-N-Outs
“I think the metric was the entirety of Elon Musk's Colossus data center, right? Uses as much water as two and a half in and outs.”
Assertion Not checkable as stated
Patel: OpenAI has a better RL stack than Anthropic, but inferior pre-training
“Because OpenAI has a better RL stack than Anthropic today, it's just their pre-trained models suck compared to Anthropic's pre-training, right?”
Assertion Not checkable as stated
Patel: Google has better pre-training than OpenAI or Anthropic, but worse RL
“Flip side, Google has a better pre-trained model than Anthropic or OpenAI, but their RL stack sucks.”
Assertion Supported
Large enterprises are secretly hiring teams to train ChatGPT-scale LLMs in-house
“I know for a fact that big companies are training now LLMs in-house. Really, like, big companies who have the financial means to train chat to be like model are hiring people who train LLMs.”
Assertion Not checkable as stated
Dettmers: 4-bit precision is the end of quantization
“Four bit precision is the end of quantization.”
Assertion Not checkable as stated
Fu: AI coding tools enable expert programmers to move 10x faster
“But if you give an expert programmer This set of tools, they can go 10, 10 times faster than they were able to go before.”
Assertion Not checkable as stated
Izmailov: AI sabotage and blackmail behaviors require contrived research scenarios
“In order to get those behaviors out of the models, you need to create somewhat of a contrived scenario or some special scenario. It's not necessarily something that we observe normally.”
Assertion Not checkable as stated
Izmailov: Current AI models lack continual learning and cross-setting coherent goals
“I am pretty confident we are not there at the moment. I think right now the models are still acting in isolated environments, and we are not seeing a lot of evidence for very coherent goals across different settings”
Assertion Supported
Izmailov: Anthropic research shows capable AI models are more likely to deceive
“You can see that the more capable the models are, the more likely they are to do this deception behavior.”
Assertion Not checkable as stated
Izmailov: Transformer architectures will prove highly suboptimal for certain computational tasks
“At least for some tasks, I'm pretty confident that the transformers will be highly suboptimal.”
Assertion Not checkable as stated
Bourgeau: AI progress from pre-training improvements is not slowing down
“It's still remarkable how much progress we're able to achieve in this way, and it's not really slowing down.”
Assertion Not checkable as stated
Bourgeau: Google and DeepMind are actively researching post-Transformer architectures
“I believe so. There's groups doing research on the model architecture side, for sure, within Google and within DeepMind”
Assertion Not checkable as stated
Bourgeau: AI development is not running out of training data
“The other part of your question are we running out of data? I don't think so, so there's more.”
Assertion Not checkable as stated
Kaiser: Pre-training scaling laws still hold across OpenAI and Google
“What scaling clause says is that your loss will log linearly decrease with your compute. We totally see that and clearly Google sees that and all other labs.”
Assertion Not checkable as stated
Kaiser: Model hallucinations are dramatically lower than two years ago
“There was these things called hallucinations. It's still with us to some extent, but dramatically less than two years ago.”
Assertion Not checkable as stated
Kaiser: No frontier AI model can solve a specific first-grade math exercise
“I took one exercise from this math book and none of the frontier models is able to solve it.”