why aren't all 2,445 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 100 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Prediction Not checkable as stated
Morcos: Proper training curricula could reduce model training costs by 10x
“And getting a curriculum right could literally make the difference between, you know, spending 10 times as much on a model training, you know, hundreds of millions of dollars potentially.”
Assertion Not checkable as stated
Morcos: Qwen is much easier to align than Llama due to pre-training
“It's much easier to RL Quen than it is to do Lama. Likely that has to do with the fact that Quen put a lot of synthetic reasoning traces into their training data.”
Prediction Not checkable as stated
Morcos: AI inference costs will skyrocket, penalizing oversized models
“The inference costs are going to skyrocket with these models. And if you use a general purpose model, then you constrain to say, hey, this model knows about everything, but now only do this one thing. That model is going to have a ton of parameters that do not…”
Prediction Not checkable as stated
Morcos: Most AI models used in three years will be under 10B parameters
“Most of the models that the vast majority of people will be using in say three years will be single digit B or smaller.”
Assertion Not checkable as stated
Morcos: Yann LeCun was never defining Meta's AI strategy
“I don't think he was ever you know, or at least not since the beginning in a role where he was defining AI strategy for Meta. I don't think that's the role he wanted at any point. You know, I think he really wanted to be doing that research, and I think, so I …”
Prediction Not checkable as stated
Huber: LLMs will largely replace purpose-built re-rankers
“I think that, like, this is going to be the dominant paradigm. I actually think that, like, probably purpose-built re-rankers will go away, and the same way that, like, purpose-built, they'll still exist, right? Like, if you're at extreme scale, extreme cost, …”
Prediction Not checkable as stated
Huber: Future retrieval systems will operate entirely within latent space
“I think, like, there's a few things that I think might be true about retrieval systems in the future. So, like, number one, they just stay in latent space, they don't go back to natural language.”
Prediction Didn’t hold up
Sohmers: NVIDIA Blackwell memory bandwidth efficiency will be lower than Hopper
“All indications are, even though they, you know, more than doubled the theoretical memory bandwidth going from Hopper to Blackwell, the actual percentage of theoretical that you can achieve is, again, going to be less than the previous generation”
Assertion Open · timeframe Aug 2026
Sohmers: Positron hardware achieves 70% higher performance than NVIDIA at lower power
“So, you know, what that actually results in is like today, we're you know, able to achieve about you know, 70% higher performance than NVIDIA with the cards that we're shipping today. Significantly lower power and price point.”
Assertion Supported
Sohmers: Positron AI requires zero compilers to run Hugging Face models
“So rather than having like, we don't have a compiler whatsoever. There's no compiler. There's no translator, no tooling that's involved in actually taking those and getting that to, you know, for your common, you know, Huggy Face Transform models to be able to…”
Prediction Open · timeframe Dec 2027
Agrawal: Positron ASIC will lead all silicon in memory capacity by 2027
“So we are going to be coming out with our ASIC and then like later, it will have more memory capacity than any other silicon in late 2026 or in 27, actually.”
Assertion Contradicted
Sohmers: Google Veo and Imagen 3 are pure autoregressive transformers, not diffusion
“A lot of things have actually been moving away from diffusion to being pure autoregressive transformers for image and video generation. So like the latest, yeah, there's a VO three and since image and three on, on Google side have been pure autoregressive movi…”
Assertion Not checkable as stated
Brockman: AI models are reaching parameter counts comparable to human synapses
“It's a hundred T synapses, which kind of corresponds to the weights of the neural net. And so there's some sort of equivalence there. And so we're starting to get to the right numbers. Let me just say that.”
Assertion Not checkable as stated
Brockman: Physicists say GPT-5 re-derived research insights taking months of work
“We've seen physicists starting to kick the tires on GPT-V and say that, like, hey, this thing was able to get, this model was able to re-derive an insight that took me many months worth of research to produce.”
Prediction Not checkable as stated
Brockman: Post-AGI humans will survive without work, but compute will differentiate capability
“And so I think that the question of exactly how, you know, if you don't do work, do you survive?
I think the answer will be yes.
You'll have plenty of material, your material needs met.
But I think the question of
Can you do more?
Can you have not just generat…”
Assertion Partly supported
The Information: OpenAI hit $12B ARR as burn rose to $8B
“We had a story yesterday about open AI and how, like, I think they've reached about twelve billion ARR and yeah, but their burn went from like They projected, like, one billion to, like, eight billion or something.”
Assertion Supported
Palazzolo: Claude Code leads stayed at Cursor only two weeks
“We know that they went there, they were there for, I think, about two weeks, and they came back.”
Assertion Not checkable as stated
Palazzolo: VCs are growing frustrated with unconventional AI acqui-hire deals
“Talking to investors, I think they're starting to, I think they're kind of over these types of deals. Like, they're like, hey, I mean, like, this is, like, fine, but, like, we don't invest in companies so they can get, like, a weird acquihire situation a coupl…”
Prediction Not checkable as stated
Dax Reed predicts Claude's $200/month pricing is an unsustainable growth strategy.
“I think it's pretty easy to, like, use more than 200 dollars worth, even by accident. So I would also lean towards that the Claude Max plans are a growth strategy, not like any long term pricing thing that can work, at least at the current, given the current s…”
Prediction Not checkable as stated
Dax Reed predicts OpenCode will dominate when a competitor beats Claude Sonnet.
“What would change things is if there's a day where either another LM lab or like, you know, an open source model drops that is competitive with Sonnet, maybe even better than Sonnet on that day, open code is going to be the only way to do this kind of thing. C…”
Assertion Supported
Ermon: Diffusion LLMs Pareto-dominate autoregressive models on inference efficiency
“On the inference side, what we're seeing is that diffusion models are much more efficient. We're actually able to Pareto dominate autoregressive models. If you think about the typical trade-off between throughput versus latency, which you kind of like cannot, …”
Assertion Partly supported
Inception generalist model matches Claude Haiku quality at 5-10x speed
“We had our generalist model evaluated by artificial analysis and the intelligence score from AA artificial analysis around 40. So it's comparable to GPT, 4.1 nano, cloud haiku, kind of like Close source speed optimized models. It's roughly comparable in terms …”
Prediction Not checkable as stated
Ermon: Power constraints will drive diffusion models to replace frontier LLMs
“If it happens, it's gonna be driven by efficiency. Like we're all constrained by essentially power. And if you have, I mean, at the end of the day, it's all an inference game, right? Okay. Training is expensive, but then the thing that matters is being able to…”
Assertion Supported
Lambert: Tulu 3 matches or beats Meta Llama 3.1 on core evals
“On, like, core evals for our Suite of models from, I think, eight, seven D and four or five B is based on llama at the time. It's like it matches or beats meta on these core valves.”
Prediction Not checkable as stated
Lambert: Hybrid reasoners may be phased out except for niche uses
“I think in plenty of ways, like hybrid reasoners might just be aged out except for niche applications because quality is so much more important than having a hundred X less inference tokens. It's like you just pay for it and compute and that'll get better.”
Assertion Not checkable as stated
Swix: OpenAI Deep Research was built by three people as an o3 wrapper
“As far as I know, it's three people did it. It was Isa and like the two other collaborators that she had. I don't know if they did that much on top of all three, like every indication I've had from over the eye is that deep research is more or less a thin wrap…”
Assertion Partly supported
Lambert: OLMo 32B roughly matches original GPT-4 level while fully open
“Like Olmo-Thirty-Tube is if you squint like original GPT-IV level and fully open.”
Assertion Supported
Fortuna: New reasoning models show no big leap on medical coding tasks
“I do know that when you kind of plot out base model performance on some medical tasks like ICD-X coding between like, you know, previous generations and new reasoning generations, there's actually not like a big leap.”
Assertion Not checkable as stated
Mohan: Codeium quality matches Copilot and drives user churn
“The product is actually one of those products where even use Copilot and use us, it's hard to tell the difference actually. And a lot of our users have actually churned off of Copilot.”
Assertion Not checkable as stated
Hou: Codeium inference costs 1/100th of competitors by avoiding third-party APIs
“It's that idea that our computation is one 100th of the cost of the competitors. We are not using APIs, and as a result, our customers and our users actually get 100 X the amount of compute that they would on another product.”
Prediction Open · timeframe Jul 2030
Scott Wu: There will be way more software engineers than ever
“And so, you know, I think software engineering, the job that we call software engineering is going to change, but I think practically, like, there's actually going to be way more software engineers than ever, you know, and I think there's a lot of precedent fo…”
Prediction Not checkable as stated
Scott Wu: AI will make engineers 5-10x more effective
“I think, you know, our demand for software to be built is actually probably a lot more than 10 X what we're currently getting, and so, you know, I think what happens is we get to open up the power of software engineering to a lot more people, and every single …”
Prediction Not checkable as stated
Wu: AI automation will expose junior engineers to core architecture earlier
“You know, I think what happens, honestly, is I think that demand is going to just keep rising with supply. And I think the training process is going to change a little bit, but, you know, I think a lot of these core fundamentals of, you know, if you think of s…”
Prediction Not checkable as stated
Mohan: Explicit user prompting will soon become an anti-pattern in AI coding
“I actually think asking people to do things explicitly is probably going to be more of an anti-pattern if we can actually go and passively suggest the entire change for the user.”
Assertion Not checkable as stated
Wu: Autonomous coding agent capability currently doubles every 70 days
“What you see in general is that that doubling time is about every seven months, which already is pretty crazy, actually, but in code, it's actually even faster. It's every 70 days, which is two or three months, and so, you know, if you look at various software…”
Prediction Not checkable as stated
Wu predicts AI coding agents will advance 16x to 64x in 12 months
“And I think that, you know, we're gonna see another 16 to 64 X over the next 12 months as well.”
Assertion Not checkable as stated
Hou: Windsurf's SWE-1 achieves near-frontier model results at lower cost
“And we've been able to achieve near frontier model results at the fraction of the cost, and with a significantly smaller team.”
Assertion Not checkable as stated
Hou: Users choose SWE-1 at higher frequencies than Claude 3.7 and 3.5
“People are choosing SWE-ONE because it recognizes how they do work, not necessarily how to generate code. And it's contributing, actually, an even higher frequency than models like 3.7 and 3.5.”
Assertion Not checkable as stated
Scott Wu: No customer or training data was shared in Windsurf transaction
“There was no, no, no information that was given out, for example, in terms of customer data, training data, any things like that. And so, so, you know, all of that is, is, is strictly proprietary and then remains, you know, our exclusive access.”
Assertion Supported
OpenAI's IMO performance was not officially verified by the IMO
“It turns out, like, OpenAI actually didn't involve officially with IMO. They just, like, use the problems, but, and then just, like, use their model to test the results, and ask, like, three previous IMO analysts to review them.”
Prediction Not checkable as stated
Scaling math AI becomes purely compute and data once auto-evaluation works
“And then my guess is, I believe in IL, so if for each category, we can figure out the A way to auto-evaluate the results, then after that, it will just be compute and data.”
Prediction Open · timeframe Jul 2028
Formalizing Fermat's Last Theorem in Lean is doable in 2-3 years
“It's I think definitely possible. Yeah. Like he, so the professor is Kevin buzzard and he got like a grant and now he just like focused on writing the proof for Ling. Like he's hoping to finish that in like two or three years. And then basically like if Ling i…”
Assertion Not checkable as stated
McCloy: Gray-Hat SEO Tactics Currently Work in AI Search
“In terms of like some of the more gray area SEO tactics like that, you know, Penguin was addressing with Google back in the day. I mean, I think the reality is a lot of those things do work today. Like I can't tell you they don't work, but I would say as like,…”
Prediction Partly held up
McCloy: ChatGPT Search Bans for Prompt Injection Are Coming
“I think it works until it stops working.
Right.
And I would say like, there's not a lot of stories of people getting banned for like Chatsby D search so far, but it's coming.”
Assertion Supported
McCloy: ChatGPT Does Not Index or Retrieve llms.txt by Default
“I knew that there's debate about this, but I'd say the evidence is like ChatTriPT is not indexing and it's not retrieving content from LMS.txt by default.”
Prediction Not checkable as stated
Kamradt Predicts AGI Will Be Declared Via An Interactive Benchmark
“My hypothesis is that when AGI is declared, it will happen via an interactive benchmark. We're not going to know that AGI is here just via a static benchmark.”
Prediction Not checkable as stated
Kamradt Predicts AGI Will Require Heavy Scaffolding and Cooperating Components
“Now I know that sounds kind of like a weird question, but my current hypothesis that AGI will be heavily scaffolded. And why do you have that? Well, my hypothesis is that you're going to need different components that are working together in order to get the e…”
Prediction Didn’t hold up
Kamradt Predicts ARC-AGI-2 Will Not Be Beaten For 12 Months
“My guess is it's not going to be beat for the next 12 months.”