why aren't all 2,445 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 100 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Supported
Agarwal: Filtered 9B Synthetic Data Outperforms 27B Self-Generated Data
“One thing we found consistently, so here what we had two models, nine Gemma, nine B and Gemma, 27 B, and we found consistently that actually generating data from nine B in a compute match setting is always better, even better for distilling or actually improvi…”
Prediction Not checkable as stated
Agarwal: Logit Distillation Can Match Giant Teacher Models on Reasoning
“My hunch is that the logic-based distillation can go even further, and you might be able to even close the gap with the biggest of the teachers you have, because I don't think you need a huge number of parameters, because the reasoning process is very, very, l…”
Assertion Supported
Cursor Composer solves Convex benchmarks but fails on alternative backends
“We did notice that I mean, with convex, it pretty much autonomously just solves the first two tasks. It has a few round trips on like some errors that are only show up and playing with the front end. And then it's able to complete this files task and kind of g…”
Assertion Not checkable as stated
Claude 3.7 performed worse than Claude 3.5 on Convex evals
“For example, we just tried clod three seven and it performs worse than clod three five on convex evals with the same prompting.”
Assertion Supported
AI models struggle debugging Supabase RLS recursion compared to procedural code
“The particular example was like RLS rules and Supabase where debugging like an infinite loop for infinite recursion for the RLS rules was something that the models just really struggled with in a way that we didn't see for procedural code.”
Prediction Not checkable as stated
Kozlov: Wave of 'agent-first' businesses will emerge without traditional UIs
“I think that we're going to see a wave of, you know, going back to, again, Sunil, your point about mom and pop shops, these businesses turn up that are agent first, and that's kind of the only interface to using them as opposed to, you know, UIs and APIs and o…”
Prediction Not checkable as stated
Solving Autonomous Coding Is the Direct Path to AGI
“Our core belief is that if you solve this problem, you solve the autonomous coding problem and build a super intelligent coding agent, that that thing will lead to super intelligence more broadly.”
Prediction Not checkable as stated
Superintelligence Cannot Be Trained Entirely From Scratch
“In the era of language models, I don't think you'll be able to train superintelligence from scratch.”
Prediction Not checkable as stated
Future UIs Will Be Built as Programmatic Interfaces for AI Models
“Over the coming years, there'll be more kind of AI friendly or language model friendly UIs. And what's friendly to a language model is, is code. So the way a model will be doing work, not just for coding and software engineering, Is by basically making functio…”
Prediction Not checkable as stated
Frontier Labs May Hoard Superintelligent Models and Release Nerfed Versions
“You can imagine you know, the world converging on a few companies have really powerful coding models. They basically release a nerfed version of that to the public at large, and basically have a competitive advantage by having, you know, a super intelligent co…”
Prediction Not checkable as stated
Klein: AI agents will use existing web interfaces instead of rebuilding APIs
“The thing that's hard is like reinventing the internet for agents. We don't want to rebuild the internet. That's an impossible task. And I think people often say like, well, we'll have this second layer of APIs built for agents. I'm like, we will for the top u…”
Assertion Not checkable as stated
Klein: Browserbase delivers 90% of computer-use functionality at 10% of OS cost
“BrowserBase can run browsers for way cheaper than you can if you're running a full-fledged OS with a GUI, you know, operating system. And I think that's just an advantage of the browser. It is like, browsers are like little OSs, and you can run them very effic…”
Prediction Open · timeframe Feb 2030
Klein: Browserbase will be a billion-dollar company within five years
“I can predict that Browserbase will be a billion dollar company one day. So let's check back in five years.”
Prediction Not checkable as stated
Next billions of AI users will onboard via SMS and phone
“I don't think those people are going to come in through, you know, some front end website somewhere. Like those people are going to come in through audio from a telephone to texting to email. Like that's just the most obvious outcome.”
Prediction Not checkable as stated
Sutin: Traditional speech-to-text ASR will soon be obsolete and uninvestable
“Cause it's very clear that like all ASR, all speech to text is going to be pretty obsolete pretty soon. So like investing into that is probably kind of a dead end cause it's just going to be obsolete.”
Prediction Held up
Roucher: AI agents will reach a 90% GAIA score by 2026
“So I think if we solve Gaia, that's like 90% score. That means mostly we double productivity of every task done in front of a computer. And if you take the trend line of the scores so far this should be crossed in 2026 or something.”
Prediction Not checkable as stated
Bret Taylor: Open source will broadly win in AI developer tooling
“Nowadays the tools are changing so rapidly that I'm like not totally skeptical of tool makers, but I just think that open source will broadly win.”
Prediction Not checkable as stated
Colvin: Gen AI observability will merge into general-purpose observability platforms
“Web observability stopped being a thing, not because the web stopped being a thing, but because all observability had to do web. If you were talking to people in 2010 or 2012, they would have talked about cloud observability. Now that's not a term because all …”
Assertion Not checkable as stated
Agarwal: 90% of production LLM use cases do not use automatic routing
“In fact, I would say in production, 90% of the use cases do not use automatic routing. What they want is deterministic flows. As long as the gateway manages authentication authorization for them, it's perfectly fine. The request hitting A specific model that t…”
Assertion Not checkable as stated
Nguyen: Stanford HELM benchmark under-reported Claude performance due to improper prompting
“This has happened with, like, Stanford, I remember, like, when Stanford had lists also, like, they were, like, running benchmarks. Yeah, Helm. And somehow, like, Claude was, like, always, like, not performing well, and that's because, like, the way they prompt…”
Prediction Not checkable as stated
Nguyen: Website clicks will drop as internet access shifts to AI models
“In my opinion, like, people in, like, few years will click On, like, websites way less. I want to see the plot of, like, website clicks over time, but then my prediction is, like, it will go down and, like, people's access to the internet will be through the m…”
Prediction Not checkable as stated
Autonomous AI programmers will work effectively within the next two years
“I think that we will have autonomous AI programmers working really well for us, you know, within the next year or two.”
Assertion Supported
DeepSeek-R1 researchers found MCTS and Process Reward Models were not useful
“R-one specifically said, yes, we tried MCTS. Yes, we tried PRMs. And none of that is useful.”
Assertion Supported
Zhang: XGrammar outperforms Outlines and is integrated into TensorRT-LLM
“And I think Xgrammar's performance is better than the outline's, and also in the TensorFlow RTLM, the latest release, TensorFlow RTLM also integrates Xgrammar as the backend for the constructed coding.”
Assertion Not checkable as stated
Automatic1111's inefficient SDXL implementation drove ComfyUI's viral user adoption
“The big, one point zero release happened, and wow, Confu UI was the only way a lot of people could actually run it on their computers, because it just, like, automatic was so, like, inefficient and bad that most people couldn't act, like, it just wouldn't work…”
Prediction Open · timeframe Jan 2028
Swyx predicts OpenAI will launch a $2,000 per month ChatGPT tier
“I think that 2000 dollars ChatGPT will come.”
Assertion Not checkable as stated
Swyx: VC appetite for GPU-rich early-stage startups is completely gone
“The appetite for GPU rich startups, like the, you know, the funding plan is we will raise sixty million and we'll give 50 of that to Nvidia. That is gone, right? Like no one's pitching that. This was literally the plan, the exact plan of like, I can name like …”
Prediction Not checkable as stated
Fanelli: 2025 will be the first year AI sets job skill floors
“And I think the skill floor more and more, I think, 20, 25 will be the first year where the AI sets the skill floor of a role.”
Assertion Supported
Neubig: SWE-bench Scores Are Inflated By Training Data Contamination
“Sweebench is on popular open source repos and all of these popular open source repos were included in the training data for all of the language models. And so, the language models already know these repos. In some cases, the language models already know the in…”
Prediction Not checkable as stated
Ben Allal: Properly curated synthetic data prevents model collapse
“And I think there's a lot of concerns about model collapse, and I'm going to talk about that later, but we'll see that like, if we use synthetic data properly and we curate it carefully that shouldn't happen.”
Assertion Supported
Ben Allal: Hugging Face SmolLM2-1.7B outperforms Llama 3.2 models
“So it's a series of three models, which are the best in class in each model size. For example, our 1.7 B model outperforms Lama one B and also .2.”
Prediction Not checkable as stated
Ben Allal: AI industry will shift to fine-tuning over prompt engineering
“And I think we're going back to fine tuning where we realize these models are really cosplay. It's better to use just a small model. We try to specialize it. So I think it's a little bit of a cycle and we're going to start to see like more of fine tuning and l…”
Prediction Open · timeframe Dec 2029
Fu: Real-time long-context video generation cannot use quadratic attention
“You're certainly not going to do a giant quadratic attention computation to try to run that.”
Assertion Supported
Reddy: Chai Discovery's open-source Chai-1 model outperforms Google's AlphaFold 3
“We're lucky to work with the folks at Chai Discovery who just released Chai One, which is open source model that outperforms Alpha Fold Three.”
Prediction Not checkable as stated
Guo: Cheaper AI code generation will increase software volume, not replace developers
“If we take the cost of software and high quality software down two orders of magnitude, we're just gonna end up with more software in the world. We're not gonna end up with fewer people doing development.”
Assertion Not checkable as stated
Mohan: Codeium generated dynamic PNGs due to VS Code API limitations
“On VS Code, actually, the problem for us wasn't actually being able to implement the feature. We had the feature for a while. Problem was actually even to show the feature, VS Code would not expose an API for us to do this. So what we actually ended up doing w…”
Assertion Supported
Cerebras WSE-3 runs Llama inference 70x faster than NVIDIA GPUs
“Cerebris came out that the wafer scale engine three can serve llama 70 B at 2.1 thousand sorry, 202,100 tokens per second and serves llama four or five B at nearly 1000 tokens per second. So this, you know, to give you an understanding, like this is about 70 t…”
Prediction Not checkable as stated
Friedman: Dedicated AI agents will prevail over general computer-use models in enterprise
“We're seeing it for a while, and I think it will stay like that despite the computer use, et cetera, that supposedly can just replace us, and it could, you can just, like, prompt it to be, hey, now be a QA, you know, or be a QA person or developer. I still thi…”
Assertion Contradicted
Friedman: GitHub Copilot user retention in enterprise is 38% to 50%
“Between 38 to 50%
Retention for users using Copilot and Enterprise.”
Assertion Supported
Friedman: AlphaCodium boosts OpenAI o1, proving o1 lacks true System 2
“We took their all one preview with Alpha Codium and did better. Like it just shows like, and there is a big difference between the preview and the IOI. It shows, like, that these models are not still system two thinkers, and there's a big difference.”
Prediction Not checkable as stated
Crivello: A Few Very Big Horizontal AI Agent Platforms Will Dominate
“I think AI is going to completely penetrate every category of software, but then I also think there are going to be a few very, very, very big horizontal agents that serve a lot of functions for people.”
Prediction Not checkable as stated
Crivello: AI Agent Software Will Create Trillions Replacing Human Labor
“It's going to sound like an exaggeration, but it is a fact that it's going to create trillions of dollars of value in a few years, right? It's going to, for the first time, we're actually having software directly replace human labor.”
Prediction Not checkable as stated
Crivello: Infinite and cheap context windows will arrive within 18 months
“Now we just assume that infinite context windows are going to be here in a year or something, a year and a half and infinitely cheap as well. And dynamic compute is going to be here. Like we just assume all of these things are going to happen.”
Assertion Not checkable as stated
Crivello: Marc Andreessen Blocked Him on Twitter Over AI Safety Views
“Like at some point, Marc Andreessen blocked me on Twitter and I, it hurt, frankly, I really look up to Marc Andreessen and I knew he would block me.”
Prediction Not checkable as stated
Polu: Post-hyper-growth tech companies may increasingly eliminate traditional SaaS
“So it's interesting that we might see kind of a bad time for SaaS in post-hyper-growth tech companies. So it's still a big market, but it's not that big, because if you're not a tech company, You don't have the capabilities to reduce desk cost. If you're a hig…”
Prediction Not checkable as stated
Polu: The next generation will see billion-dollar companies with 20 engineers
“All generations of company might be the first billion dollar companies with engineering teams of 20 people. That would be so exciting as well. That would be so great. You know, you don't have the management hurdle. You're just 20 focused people with a lot of a…”
Assertion Supported
SWE-Bench public test splits enable trivial cheating via runtime PR retrieval
“The entire test split here is public. So you can do things like just overfit to the patches in the test set. You can do things like, let me add at runtime, pull the PR and just get the answer and just use it.”
Prediction Not checkable as stated
Houston: Fully autonomous AI knowledge workers will take a long time
“People sort of skip to like level five full autonomy, or we're going to have like an autonomous knowledge worker that's just going to take, that's going to, and then we won't need humans anymore kind of Projection that that's going to take a long time, but the…”