why aren't all 2,374 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 6 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Not checkable as stated
Lambert: OLMo 3 32B base model matches Qwen 2.5 32B quality
“This base model is similar in quality to the best available, which is like Quinn's 2.5, 32 B is, was still the best base model.”
Assertion Not checkable as stated
Lambert: OLMo 3 7B outperforms Meta's Llama 3.1 8B in internal tests
“And I just think of this cause like Lama 3.1 AP is one of the most used models and hugging base of all time. And this should be better. We're, In our measurements, we see it as being better than Llama.”
Assertion Partly supported
Lambert: 80% of a16z's open-model portfolio startups use Alibaba's Qwen
“80% of companies building with open models are using Quinn, which is like 16 to 24% of his portfolio, which is still a lot.”
Assertion Not checkable as stated
Lambert: Chinese companies with $1B+ valuations routinely pirate SaaS software
“Mediumly large, like billion dollar plus valuation companies in China will just like pirate SaaS software.”
Assertion Not checkable as stated
Lambert: Best open-license AI models near the frontier in 2025 were Chinese
“The models that are from closest to the frontier in performance with good license all happened to be Chinese models throughout the year for this case.”
Assertion Not checkable as stated
Kant: Next-token pre-training gains hit a sigmoidal curve and slowed down
“The first paradigm of kind of pre-training of predicting the next token on the web was becoming sigmoidal and was slowing down in terms of the gains that it had.”
Assertion Not checkable as stated
Top AI serving companies achieve 70% to 90% gross margins
“There are companies here that are making very, very good margins on serving their AI systems, like, 7080, sometimes 90%, depending on the modality and so, like with everything, the average number sucks but, like, when you look at the best companies, it's reall…”
Assertion Supported
Meta raises tens of billions in off-balance-sheet debt for data centers
“You have this, sort of, offloading of debt from big companies, for example, Meta, that raises tens of billions of dollars to fuel its data center ambitions, but that doesn't sit on Meta's balance sheet.”
Assertion Not checkable as stated
Major tech companies abandon green commitments to secure AI power
“Well, a year or two ago, big companies did make commitments to be green as of, you know, as of, you know, as soon as they started inking deals with you know, nuclear companies and, Various energy providers for data centers, all those commitments basically got,…”
Assertion Supported
Anthropic agreed to a $1.5 billion training data copyright settlement
“And then there was a biggest settlement that happened in the last few months with Anthropic that agreed to pay out one and a half billion.”
Assertion Not checkable as stated
Google Search survives because ChatGPT relies heavily on referencing it
“People say, oh, Google search is dead. I think that's like probably completely wrong because ChatGPT references Google a ton.”
Assertion Not checkable as stated
AI autonomous task duration doubles every three to four months
“We are seeing this very consistent improvement over many, many years where every say like, you know, three, four months is able to like do a task that is twice as long as before completely on its own.”
Assertion Not checkable as stated
Tworek: Coding agents are the first successful agentic AI products
“Like, coding agents are at the moment the first, like, pretty successful agentic products built on top of AI.”
Assertion Not checkable as stated
Tworek: OpenAI's o1 release caught US AI labs unprepared for RL
“As far as I know, like our O-one release mostly caught a lot of us labs by surprise. They didn't have like similarly advanced RL research program to my knowledge, basically no one.”
Assertion Not checkable as stated
Douglas: An Anthropic AI agent operated autonomously for 30 hours building apps
“We asked it to build something that looks roughly like a chat app, you know, something like Slack or, you know. And it was, it, the model just worked for 30 hours. Like, it was just spinning there on a computer for 30 hours, and came out with a really good wor…”
Assertion Supported
Douglas: AI autonomous task execution time horizons double every six months
“And so I think it's like every couple of months, the time horizon that the AIs are capable of doing is doubling or something, something crazy. Maybe, maybe every six months the time horizon doubles”
Assertion Not checkable as stated
Douglas: AI coding interventions stem from taste, not raw programming capability
“Right now you need to intervene quite frequently, but it's usually on questions of taste rather than it is questions of, like, raw programming ability.”
Assertion Not checkable as stated
Douglas: Robotic locomotion is essentially solved using basic reinforcement learning
“Locomotion's kind of solved, to be honest, with basic RL.”
Assertion Supported
Cherny: Claude Code does not use RAG for codebase memory
“And so quad code actually doesn't use this technique called rag. Instead, what it does is it just searches files the same way that a human would.”
Assertion Not checkable as stated
Laskin: Reflection AI regularly beats OpenAI, Anthropic, and DeepMind for talent
“We win over candidates over OpenAI and Anthropic Meta, DeepMind regularly.”
Assertion Not checkable as stated
AI startups reach $30M ARR three times faster than fast-growing SaaS startups
“Those that already hit thirty million in annualized revenue got there in about a year and a half. For comparison, you know, many of us were around five years ago, like the fastest growing SaaS startups on Stripe took, you know, five and a half years to hit tha…”
Assertion Supported
Stockholm-based AI startup Lovable reached $50 million ARR in six months
“European breakouts lovable out of Stockholm hit fifty million ARR in six months and is now for sure the fastest growing startup in Europe.”
Assertion Not checkable as stated
Walsher: Cursor announced reaching $500M in ARR
“Cursor, maybe three weeks ago, announced they're at five hundred million dollars of ARR.”
Assertion Partly supported
Walsher: Computer science graduates face top-tier college unemployment rates
“Computer science grads are actually among the top five or six majors graduating from college right now with the highest unemployment rate.”
Assertion Not checkable as stated
Rauch: Vercel's v0 has positive and improving gross margins
“Yeah, BZero has positive gross margins. It's a healthy business. The margins are improving, the business is improving”
Assertion Not checkable as stated
Gomez: Modern Transformers look strikingly similar to the original 2017 architecture
“And so one of the big shocks is how over the past eight years, how little things have changed. Like it, it's really surprising to me. That the Transformers we train today looks so similar to what was back then.”
Assertion Not checkable as stated
Gomez: Google failed to lean into language modeling early, unlike OpenAI
“To say they didn't lean hard enough into language modeling, like just pure Sequence modeling of text on the internet. That's, I think the accurate statement. That's what OpenAI did early and uniquely well.”
Assertion Not checkable as stated
Gomez: AI agents cut financial research tasks from a month to hours
“So we can take something that used to be a month. And bring it down to, you know, four hours, eight hours.”
Assertion Not checkable as stated
Hebbia pioneered inference-time compute scaling two years ago
“And actually the whole scaling at inference paradigm was pioneered at Hebbia. So our early matrix product two years ago, we're one of the first people to say, hey, you get way better accuracy from using more large language model calls at runtime.”
Assertion Contradicted
Hebbia was the first company to productionize RAG in 2020
“Hebbia were actually the first to turn that into a product. So it's like a very close thing to my heart. So back in 2020, we were the first people to actually productionize it, roll it out.”
Assertion Not checkable as stated
Hebbia developed an undefeated re-ranker architecture that it does not use
“And we came up with a novel re-ranker architecture, which four years later, academia and industry have not beat, and we do not use it.”
Assertion Not checkable as stated
Evans: AI frontier models are commodities spread across half a dozen organizations
“The thing that's become very clear, if it wasn't clear a year ago, is that the models themselves are sort of commodities in that, you know, there's half a dozen people who have a state-of-the-art model.”
Assertion Partly supported
Evans: OpenAI's Deep Research inverted numbers in its own marketing example
“Then the interns typed the number in wrong. Like, it was literally the wrong percentage. It was like, 65, 35, instead of 35, 65.”
Assertion Not checkable as stated
Evans: No standalone breakout consumer apps exist wrapping ChatGPT API
“We don't have a breakout, there's no standalone breakout consumer app. There's no one, there's no, there's all these, there's, there are all these enterprise SaaS stuff. There is not really a consumer equivalent. There aren't hundreds of consumer apps using th…”
Assertion Not checkable as stated
Evans: Public discourse around AI doomerism has effectively vanished
“Yes, all of the dumerism's gone away.”
Assertion Supported
Howard: OpenAI is shutting down GPT-4.5
“I think they're shutting down that product or they've shut down that product, if I understand correctly.”
Assertion Not checkable as stated
Arvind Jain: Enterprise customers rarely run AI agents fully unsupervised
“We barely see any application where customers are running agents in a fully unattended unsupervised, you know, setting.”
Assertion Not checkable as stated
Levie: AI coding tools have reached mainstream enterprise adoption
“AI coding is very clearly in the pragmatists. You know, get up co-pilot now, I don't know, three years old cursor's flying off your shelves, wind service flying off the shelf, you know, replit, all these guys are kind of being, you know, used everywhere. So yo…”
Assertion Not checkable as stated
Ramaswamy: Snowflake's early AI efforts suffered from an infrastructure-first mentality
“And so part of the difficulty that Snowflake had with machine learning and AI was they brought a similar mentality. You're going to write a design doc. It's going to take us 18 months. It's going to be great. Opposite of what you need to succeed in iterative e…”
Assertion Not checkable as stated
Chollet: GPT-4 lacks fluid intelligence, but OpenAI's o3 model has it
“GPT-IV does not have fluid intelligence, for instance, but O-III does.”
Assertion Supported
Chollet: Latest base LLMs score zero percent on ARC-AGI-2
“Today the latest base alarms, they're doing something like 10% on ARK-I. But on Arc two, they are doing zero percent.”
Assertion Not checkable as stated
Kiela: DeepSeek's total development cost was at least 100x its $6M training
“So I would guess that they spent at least a hundred X The amount of that, that single training run, right?”
Assertion Not checkable as stated
Misra: AI-generated marketing videos outperform human actors via infinite variation testing
“And very quickly for us, it became, wait, this actually performs better than if we were to, you know, hire someone to like make this video basically. Right. And the reason is because we can customize it to an unlimited ability, right? Like we can get all the v…”
Assertion Partly supported
Misra: Nvidia H100 and A100 GPUs encode video slower than T4s
“And by the way, like the H 100 and A 100 kind of suck at encoding and decoding video, right? They're actually slower than like a T four, for example, which has like media sort of like drivers and stuff that can actually like do it really quickly.”
Assertion Not checkable as stated
Misra: Captions AI videos get tens of millions of views without viewers noticing AI
“We run these things in ads like on scale. We also run these things on social media. People, big creators use, you know, these models and they get tens of millions, sometimes many, many tens of millions of views on their videos and nobody can tell that it was g…”
Assertion Not checkable as stated
Ada's AI agents resolve up to 85% of customer service conversations autonomously
“Our agents are now autonomously resolving. We have a new high watermark as of last week, week before, 85% of all conversations that, up to 85% that they're seeing. That's without human involvement”
Assertion Not checkable as stated
Ada measures customer conversation quality as accurately as humans across 100% of interactions
“And we can say with confidence right now that yes, that is exactly true. We can measure the quality of a conversation transcript as well as a human annotator. And we can do it with a hundred percent coverage.”
Assertion Contradicted
Replit claims to be only company offering full-stack AI app generation
“We're like, we're the only company that that's done that, which is like the full stack experience from the code to the databases, to the cloud deployment.”