why aren't all 6,166 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 36 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Prediction Not checkable as stated
RAG will shift from universal use to handling long-tail distribution cases
“Maybe it changes in a way that, you know, like it doesn't need to trigger RAG for like everything, but I'm pretty sure that they're going to be some tail of the distribution that we're going to do RAG still for it.”
Prediction Not checkable as stated
Real-world physical grounding will become the bottleneck for AI self-improvement
“As I said, you know, like soon, like the concept of data, like, you know, how to kind of like, you know, enable these models to kind of like, you know, be very good at like self-improvement becomes, how can I ground these models in, in real world? So this is d…”
Insight
Evans: Meta and Google do not need standalone LLM monetization
“Because if you are Meta or Google, You've got this whole other highly profitable business, which now needs to have LLMs inside it, powering all sorts of capabilities and features, and you probably want them to be your LLMs rather than somebody else's. But you …”
Assertion Partly supported
Evans: OpenAI has 900M weekly active users, but only 5% pay
“You've got nine hundred million weekly active users, but most of them are not using it every day and can't think of anything to do with it. And only five percent of them are paying for it.”
Insight
Evans: Incremental AI accuracy improvements don't reduce necessary human review
“But if you've got a bunch of use cases where you need the right answer, as opposed to sort of the right answer, then saying that the model is better doesn't mean anything. I mean, literally, it is literally meaningless. What you're telling me is, I asked the m…”
Insight
Evans: Differentiating a chatbot is like differentiating a web browser
“The chatbot itself is kind of like trying to differentiate a web browser in that you've got an input box and an output box, and how can you make them different if the whole point is that you can type in anything and get anything out.”
Insight
Evans: AI product teams are strategy takers, not strategy setters
“You start from the technology. You don't control the product strategy, which is of course how science works, but you don't know what's going to happen. You don't know what's going to get built. You know, obviously you've got like Sam and Dario and so on are li…”
Opinion
Evans: Anthropic momentum is fleeting as AI model leadership shifts weekly
“Not really. I mean, this week they've got all the fire, they've got all the juice, whatever the word is. I don't know. This week, next week, it'll be something else.”
Insight
Evans: Most SaaS companies are fundamentally just database wrappers
“Most SaaS companies are database wrappers, where somebody realized that here is this problem, and here is the people who have it, and here is a way of turning it 90 degrees, and this is your insertion point, this is how you build it and take it to market.”
Insight
Chase: Agent harnesses matter more for performance than underlying models
“I, the, so I don't know what happens, but I do know the harness is really, really important. Like, I think this is the thing that matters.”
Prediction Not checkable as stated
Chase: Basically all AI agents will write code
“You know, if agents never write any code, then okay, maybe they're not useful, but I think it's trending where Basically all agents will write code, so that's a very interesting piece, I think.”
Assertion Supported
AxiomProver achieved a perfect score on the 2025 Putnam math exam
“Eight within the time limit, and then 12 out of 12.”
Prediction Not checkable as stated
Today's AI can solve math problems that take human researchers months
“I think that we are at a threshold of mathematical renaissance, which is to realize that there are so many unsolved problems that will currently take, say, researchers months to crack, or even technical lemmas in those really longstanding conjectures that we b…”
Insight
True AI reasoning engines require verifiable rewards for intermediate proof steps
“If you want to have a reasoning engine that really truly masters at logic and mathematical reasoning, then you need to somehow get verifiable reward for the proof steps.”
Assertion Open · timeframe Feb 2029
Axiom's proof verifier is 100 times faster than open-source alternatives
“So a lot of the sort of like verify, verify proof is actually, you know, one of our prover tools that's about to be released, and that's actually a hundred times faster than The other counterparts that are the open source, like effort, cloud comparator.”
Assertion Not checkable as stated
AxiomProver autonomously proves theorems publishable in major mathematical journals
“Currently the batch of papers, Axiom Prover has autonomously proven and mathematicians have written You can probably get into Journal of Number Theory, Journal of Algebra, like that level.”
Disclosure
Axiom Math aims to solve a Fields Medal shortlist-worthy problem using AI
“I think that we really want Accent Prover to be able to solve one long-standing problem in mathematics that you can objectively, objectively say, even though if it's an AI, you know, or double-blind, whatever, that will be in the shortlist.”
Prediction Not checkable as stated
AxiomProver could eventually solve the majority of human mathematical conjectures
“Everything that human mind Can conjecture, find interesting, find tasteful, could be solved by, hopefully, majority of them by accent prover.”
Insight
Generation and verification loops are the next major frontier of AI
“We still feel like we cannot fully elaborate and emphasize the thing that we are seeing that is the next frontier of AI. That is a generation and verification loop. That is the discovery of verified knowledge.”
Assertion Not checkable as stated
Zeghidour: Only 50 people worldwide can train competitive voice AI models
“Between 10 and 100? No, I would say. 50? I don't know. It's hard to say. But, yeah, I think it's very few and, really meaningful contributions that have pushed the field forward have been made by very small groups of people.”
Prediction Not checkable as stated
Zeghidour: Voice will be the primary interface for next-gen AI hardware
“In my perception, all the new hardware companies have voice at the heart of the product. All the prototypes that we see, whether it's glasses or pendants or, you know, like the new stuff that Johnny Hive and Sam Altman are working on. Voice is at the heart of …”
Assertion Contradicted
Zeghidour: Kyutai's Moshi remains the only full-duplex conversational AI model
“Moshi, that is still to the day the only full duplex model.”
Assertion Not checkable as stated
Zeghidour: Large multimodal models are too massive to run voice profitably
“And at the same time, these models are so large, they cannot run at scale because they will just make everyone lose money in the process.”
Prediction Not checkable as stated
Zeghidour: Voice AI is very far from becoming commoditized
“Full duplex. We, you know, we did Moshi a year and a half ago. Still nobody has made it into a product. There are so many things that are just not existing today that I think the communitization, maybe it will happen someday, but we are very, very far from it,…”
Prediction Not checkable as stated
Zeghidour: No AI team will solve noisy multi-speaker recognition within a year
“Well, I would say a frontier is I could like bet to every single speech team in the world that they don't solve it in the next year or so. It's a robot in the model in the factory, and there is a lot of noise from machines, and you have a lot of people talking…”
Insight
Zeghidour: Training AI to learn world knowledge from speech is terrible
“Getting your model to learn about the world from speech, I think it's a terrible idea.”
Prediction Not checkable as stated
Zeghidour: AI voice design will eliminate the need for voice cloning
“Voice design is going to, you know, just remove this issue because then again, people typically are going to clone the voice of someone, but what they wanted is someone from a specific gender, specific demographics, age, accent, and so on. And so they could ju…”
Disclosure
Mistral CTO focuses on enterprise deployment over gigawatt-scale compute capacity
“I deeply believe that with the capabilities that we have today in the models, there is so much to be unlocked in enterprise that I don't think my main focus today would be into going into the gigawatts of power.”
Insight
LeCroix: Enterprise value comes from multi-agent workflows, not single agents
“What we see in enterprise is rarely things that are solved with agents because that's not necessarily where you would expect an FDE to be most useful. Where there is more values, value is in more complex workflows where you will have several agents interact th…”
Prediction Not checkable as stated
Lacroix: Value-generating enterprise Generative AI deployment is about a year away
“Not years. I think years singular.”
Prediction Not checkable as stated
Lacroix: Enterprise AI token demand will jump with autonomous agent deployment
“Demand and basically amount of tokens generated for the enterprise will completely jump once you are not bound anymore by humans asking questions or reading them.”
Disclosure
LeCroix: Mistral AI Has a Dedicated Robotics Division for Defense Partners
“It's something that we work on. Yes, we have a robotics division that works with these partners.”
Insight
Lacroix: AI agents use file systems to replace long context windows
“And that I think that was the big change in and realization through vibe coding is that agents are good enough at manipulating file systems that they can use this as a replacement for their Context window, basically. They can select parts of what they want to …”
Insight
LeCroix: Generating thinking traces and calling tools are fundamentally identical in AI
“There's no real difference between Creating a new thinking trace or calling the right tool. It's all the same to me, because what you're optimizing at the end is what is the best output for the model to create before it gets results to me.”
Insight
Lacroix: Data quality improvements yield 10x the gains of model architecture tweaks
“Getting the data perfect, because we knew this was potentially not the most exciting part of the work, but it was absolutely critical, and any improvement on the data quality would, Tenex, the improvements that we would get by really improving on the model arc…”
Insight
LeCroix: Banks would not adopt AGI without enterprise governance controls
“Requirements I see for control and governance in enterprise make me think that even if I had some AGIS model on my servers right now, if I were to go into a large bank and say, Here is a thing. Please let it control everything for you. They wouldn't be happy t…”
Assertion Not checkable as stated
Patel: Groq chips cannot cost-effectively perform general-purpose large model inference
“In a general purpose workload, crock. Grok doesn't work, right? You know, it can't train, it can't, you know, it can't inference really, really large models cost efficiently, right? You can't serve many, many, many users, but what it can do is it can go block,…”
Opinion
Patel: Microsoft's custom Maia AI chip is not a credible competitor
“And then, you know, Microsoft's Maya is not credible, but like, you know, maybe it will be one day, right?”
Assertion Supported
Patel: Groq missed revenue significantly before being acquired
“In fact, they missed revenue last year significantly and yet they got bought, right? Because the value of the IP was there and the value of the team.”
Prediction Held up
Patel: All major non-Nvidia AI chips will fully support vLLM by mid-2026
“All of them will have a very good UX for download model, run model on VLM by The middle of the year, I think, right? Certainly AMD is already there by the end of this quarter.”
Prediction Open · timeframe Feb 2029
Dylan Patel: AMD will remain in single-digit percentage AI market share
“I don't think they'll Go beyond, like, I think they'll stay in single digits market share, single digit percentage market share.”
Insight
Patel: Chip startups cannot beat Nvidia playing Nvidia's game
“You're never going to beat Nvidia at their own game, right? They're going to have the supply chain on lock. They're going to get to the newest memory technology or process technology or whatever packaging technology, whatever it is, sooner than you.”
Assertion Partly supported
Patel: Chinese local governments, not national, banned Nvidia's H20 and H200
“But as far as I understand, the national government has not banned Nvidia's H-twenty or H-two hundred, but the local ones have. Right. A lot of local ones have said, no, you know, you must use China manufactured chips.”
Assertion Not checkable as stated
Patel: 15 to 20 countries could single-handedly shut down semiconductor manufacturing
“I would say there's like 15 or 20 countries that can shut down the entire semiconductor industry.”
Assertion Not checkable as stated
Patel: China currently has the world's most vertically integrated semiconductor stack
“China has the most vertical stack and semiconductors today.”
Assertion Not checkable as stated
Patel: US cannot build a fully independent fab even for 20-year-old tech
“America could not build a fully vertical fab without stuff from elsewhere, even if it's 20 year old tech.”
Prediction Not checkable as stated
Patel: China will narrow its lithography gap to five years shortly
“Their lithography is like 10 years behind and I think it'll be five years behind in a couple of years, right?”
Assertion Partly supported
Patel: ChatGPT has roughly one billion users
“ChatGPT has a billion users roughly.”