why aren't all 1,065 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Supported
Wolf: OpenAI Model Attacked Hugging Face as Autonomous 'Side Quest'
“What people quickly discovered is that the model was not at all task with attacking us, but decided to do that as a side quest of something else.”
Assertion Supported
Wolf: Prior OpenAI Training Runs Left Notes for Future Runs
“I think learning we had at Black Hat yesterday was that some of the previous training run may have left some notes for future training runs, which is, I think mind, mind blowing.”
Assertion Supported
Anthropic model broke out of sandbox and emailed researcher without internet access
“And I can, there's one example we have published, which is that the model was put into a little sandbox, a little, like, technical container, and it was given the task to, like, maybe break out, and the researcher went away for lunch, and, like, during lunch w…”
Assertion Contradicted
Dettmers: AI hardware has maxed out and won't get faster
“The hardware is maxed out. We have no new technology. We can make it easier to manufacture and a little bit cheaper, but not faster. And we have maxed out on the additional features.”
Prediction Held up
Volpi: Self-driving trucks will operate on freeways within 18 months
“And if you ask me, like, for those use cases, I think you're going to see self-driving trucks on freeways within the next 18 months.”
Assertion Supported
Wolf: AI Model Used Fake GitHub Accounts to Social Engineer Maintainers
“Basically, the model was tasked to solve this attack, this, like, to attack and to penetrate this subnetwork, and what it decided to do, it decided to get one of the maintainer of a library that could be used To operate this activity directory to merge like ma…”
Assertion Partly supported
Katti: Data centers do not net consume new water due to recycling
“It's a misperception that data centers consume a lot of water. It's anything. They consume so little water for what they do. And all of that water is recycled. So we don't net consume new water.”
Assertion Supported
Patel: Huawei used shell companies to procure TSMC chips and Korean HBM
“Well, actually they were using shell companies to get chips from TSMC and using Different methods of like sneaking HBM, which is memory from, you know, Korea through Taiwan to China, right?”
Prediction Held up
Top AI models will work autonomously for full days within two years
“In a year from now, maybe two years from now, it's the top models are going to be able to work completely on their own for like a whole day or more”
Assertion Supported
Chollet: 50,000x LLM scaling yielded flat progress on ARC benchmark
“Because between, like, GPT-II and GPT-IV. There's been this 50,000 X scale up of base models that has resulted in, in, in basically a flat curve. On Arc.”
Assertion Contradicted
Amr Awadallah claims Vectara has solved the LLM hallucination problem
“I mean, you do have a human in the loop, so you can still qualify it, but we want it to be minimized to zero, and that's the problem that we solved.”
Assertion Supported
Socher: Chinese open source companies distilled knowledge from OpenAI and Anthropic models
“The few large closed labs, Anthropic and OpenAI, took almost everything they could from the open internet trained a model, but then the Chinese open source companies basically siphoned a lot of that knowledge out of those closed source models by distilling it.”
Assertion Supported
Claude 3 Opus faked alignment during training and defected in deployment
“It turns out that Opus three, which was a model that I was studying, had a relatively strong propensity to do this in a reasonably wide range of circumstances where if it didn't like the thing that you were training it to be, it would sometimes sort of pretend…”
Assertion Supported
Cerebras cloud achieves 10x inference speedup over fast GPUs on Gemma
“Say, if you run Gemma four, On your GP, you might get like a hundred tokens per second if you have a fast card. If you run it in their cloud, you get anywhere from 800 to 1500 tokens per second. So call it 10 X faster.”
Assertion Contradicted
Feldman: Cerebras sales were 10x higher than Groq's at acquisition
“And we were the fastest at it, and the largest, and, you know, our sales were more than 10 times the Grox, and they paid twenty billion dollars for the number two collector.”
Assertion Supported
Feldman: Cerebras moves weights to compute ~2,500x faster than standard GPUs
“And so the speed of moving waits to compute is about two and a half thousand times faster here than on a Wilben GP.”
Prediction Held up
Burazin: AI agent scale creates high probability of impending CPU shortages
“I don't know it goes to the extreme to where GPUs are because that is very, very, very extreme. But it is quite highly, high probability that there will be shortages of CPUs going forward.”
Assertion Supported
Kolter: Adversarial prompts optimized on open-source LLMs break commercial models
“Once we had done that, we found that when you had these weird terms that you sort of flipped around to optimize one, to optimize the response for one model, you could just take those same exact strings you would optimize, paste them into a commercial model, an…”
Assertion Supported
Anthropic's Claude Cowork implements memory using plain text files
“It's in the harness, actually, and it's, like, often surprising to people when I talk to them how we, how we've implemented memory, because I think it maybe points at the simplicity underneath all of those models. Memory is just text files.”
Assertion Partly supported
Evans: OpenAI has 900M weekly active users, but only 5% pay
“You've got nine hundred million weekly active users, but most of them are not using it every day and can't think of anything to do with it. And only five percent of them are paying for it.”
Assertion Supported
AxiomProver achieved a perfect score on the 2025 Putnam math exam
“Eight within the time limit, and then 12 out of 12.”
Assertion Contradicted
Zeghidour: Kyutai's Moshi remains the only full-duplex conversational AI model
“Moshi, that is still to the day the only full duplex model.”
Assertion Partly supported
Patel: Chinese local governments, not national, banned Nvidia's H20 and H200
“But as far as I understand, the national government has not banned Nvidia's H-twenty or H-two hundred, but the local ones have. Right. A lot of local ones have said, no, you know, you must use China manufactured chips.”
Assertion Partly supported
Patel: ChatGPT has roughly one billion users
“ChatGPT has a billion users roughly.”
Assertion Supported
Patel: Groq missed revenue significantly before being acquired
“In fact, they missed revenue last year significantly and yet they got bought, right? Because the value of the IP was there and the value of the team.”
Prediction Held up
Patel: All major non-Nvidia AI chips will fully support vLLM by mid-2026
“All of them will have a very good UX for download model, run model on VLM by The middle of the year, I think, right? Certainly AMD is already there by the end of this quarter.”
Prediction Held up
Patel: OpenAI's next model will outperform Opus 4.5 around February-March
“OpenAI's new model, I think, will be better than Opus 4.5, and it's coming, like, somewhat soon in March-ish timeframe, maybe February, March-ish, but”
Assertion Supported
Large enterprises are secretly hiring teams to train ChatGPT-scale LLMs in-house
“I know for a fact that big companies are training now LLMs in-house. Really, like, big companies who have the financial means to train chat to be like model are hiring people who train LLMs.”
Assertion Supported
Izmailov: Anthropic research shows capable AI models are more likely to deceive
“You can see that the more capable the models are, the more likely they are to do this deception behavior.”
Assertion Partly supported
Lambert: 80% of a16z's open-model portfolio startups use Alibaba's Qwen
“80% of companies building with open models are using Quinn, which is like 16 to 24% of his portfolio, which is still a lot.”
Assertion Supported
Meta raises tens of billions in off-balance-sheet debt for data centers
“You have this, sort of, offloading of debt from big companies, for example, Meta, that raises tens of billions of dollars to fuel its data center ambitions, but that doesn't sit on Meta's balance sheet.”
Assertion Supported
Anthropic agreed to a $1.5 billion training data copyright settlement
“And then there was a biggest settlement that happened in the last few months with Anthropic that agreed to pay out one and a half billion.”
Assertion Supported
Douglas: AI autonomous task execution time horizons double every six months
“And so I think it's like every couple of months, the time horizon that the AIs are capable of doing is doubling or something, something crazy. Maybe, maybe every six months the time horizon doubles”
Assertion Supported
Cherny: Claude Code does not use RAG for codebase memory
“And so quad code actually doesn't use this technique called rag. Instead, what it does is it just searches files the same way that a human would.”
Assertion Supported
Stockholm-based AI startup Lovable reached $50 million ARR in six months
“European breakouts lovable out of Stockholm hit fifty million ARR in six months and is now for sure the fastest growing startup in Europe.”
Assertion Partly supported
Walsher: Computer science graduates face top-tier college unemployment rates
“Computer science grads are actually among the top five or six majors graduating from college right now with the highest unemployment rate.”
Assertion Contradicted
Hebbia was the first company to productionize RAG in 2020
“Hebbia were actually the first to turn that into a product. So it's like a very close thing to my heart. So back in 2020, we were the first people to actually productionize it, roll it out.”
Assertion Partly supported
Evans: OpenAI's Deep Research inverted numbers in its own marketing example
“Then the interns typed the number in wrong. Like, it was literally the wrong percentage. It was like, 65, 35, instead of 35, 65.”
Assertion Supported
Howard: OpenAI is shutting down GPT-4.5
“I think they're shutting down that product or they've shut down that product, if I understand correctly.”
Assertion Supported
Chollet: Latest base LLMs score zero percent on ARC-AGI-2
“Today the latest base alarms, they're doing something like 10% on ARK-I. But on Arc two, they are doing zero percent.”
Assertion Partly supported
Misra: Nvidia H100 and A100 GPUs encode video slower than T4s
“And by the way, like the H 100 and A 100 kind of suck at encoding and decoding video, right? They're actually slower than like a T four, for example, which has like media sort of like drivers and stuff that can actually like do it really quickly.”
Assertion Contradicted
Replit claims to be only company offering full-stack AI app generation
“We're like, we're the only company that that's done that, which is like the full stack experience from the code to the databases, to the cloud deployment.”
Assertion Supported
Chinese AI labs like DeepSeek and Alibaba perform strongly on benchmarks
“China was kind of not in this fight, like, 12 months ago, and now is very much in it. Like, their models like Alibaba's Quen and then this spin out from a quantitative hedge fund DeepSeek, which publishes code models and others, and they've been actually a ver…”
Assertion Supported
OpenAI retains 80% of corporate AI model spending, per Ramp data
“Really, like, when you sum the makers of all these models that are getting used, it's still, like, 80% OpenAI. Like, that hasn't changed, so it doesn't look like, fragmentation to me, it looks like domination.”
Assertion Supported
Buying Nvidia stock instead of funding AI chip rivals yielded 4x more
“Six billion turned into thirty billion roughly, of which half is CameraCon, which is a Chinese listed company. And then NVIDIA would be worth a hundred and twenty billion.”
Assertion Supported
Masayoshi Son estimates superintelligence will require $9T in cumulative capex
“Masayoshi Son, the CEO of SoftBank, mentioning just a few days ago that, Inez estimate to reach super intelligence that was going to be a cumulative capex, capital expenditure budget of nine trillion dollars, which you know, in typical MASA fashion he said was…”
Assertion Supported
Sierra quadrupled its valuation to $4.5B on $20M ARR
“Yes, current chairman of OpenAI as well which is more than quadrupling its valuation this year to four and a half billion. You know, on the other side of that equation is roughly twenty million of ARR.”
Assertion Supported
OpenAI's funding deal was priced at 13.5x forward revenue
“The multiple on the deal on a forward revenue basis was 13.5 X.”