why aren't all 6,166 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 36 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Prediction Not checkable as stated
Levie: A $5B startup will be built around AI compute ERP
“You're going to need, you know, new, new pieces of software. Probably there's probably a, you know, a five billion dollar startup waiting to happen just in like ERP for your AI compute.”
Prediction Open · timeframe May 2031
Levie: Average enterprise will use half a dozen AI models
“I think what's going to happen is you're going to have a mosaic of models in the enterprise. I think the average enterprise will certainly be using, you know, half a dozen models in their organization.”
Insight
Levie: Rapid AI breakthroughs paradoxically slow corporate AI adoption
“The technology is getting so advanced that it makes obsolete the prior thing that you implemented, which actually means that the rollout takes longer because we have no stable, there's no stable environment to roll things out in.”
Prediction Not checkable as stated
Levie: Headless AI queries will be 100x larger than interface work
“So I think it's just going to be this sort of dual, dual model with the one nuance being probably by like you know, database queries, headless will just be a hundred times larger than the interface driven, you know, way of doing work.”
Prediction Not checkable as stated
Levie: Enterprise software will combine seat and consumption models within three years
“So I think any enterprise software company in three years from now, that sort of, that, that gets through this AI transformation period, It will have a seat business model, assuming it has an end user component, and it'll have a consumption business model.”
Prediction Not checkable as stated
Levie: AI agents will drive hiring as executives capture new value
“You will actually see, interestingly, if you're like an executive and you start to do this, you'll see lots of areas actually where you should hire more people because you're like, oh my God, this thing is spitting out, you know, incredible goldmine of value, …”
Insight
Levie: AI labs won't displace vertical startups without massive specialized headcount
“Unless the labs build out literally the equivalent of hundreds or thousands of people for every single vertical and every single line of business, that means that there's actually a lot of opportunity in that kind of bridge area of the work.”
Insight
Dubois: AI progress feels discontinuous because OpenAI crossed a reliability threshold
“Even though the, in my mind, everything, the progress is actually pretty continuous, you need to reach this level of reliability. To really make any of these AI tools very useful, and I think we just crossed that probably December last year, at least at OpenAI…”
Insight
Dubois: Test-time compute scaling exhibits logarithmic, diminishing returns
“We, we've seen again and again, the longer the model think for the better answers we will get. The problem is that this, these curves that we're talking about are not, are definitely not linear, and like they, there's some plateauing effect, and they kind of l…”
Insight
Dubois: Larger AI models achieve higher efficiency by thinking through weights
“If you have larger models the amount of thinking time, so the amount of tokens they will think for will usually decrease. And the way that you can think about it is that metaphorically, the model already thinks through its weights when it generates a certain t…”
Assertion Not checkable as stated
Dubois: AI frontier labs have successfully bypassed internet data walls
“There were a lot of conversation about hitting data walls, and it seems like we did not quite hit it. So the larger the model is, the more data it needs to ingest to be trained. And it seems like different companies kind of found different ways to overcome the…”
Prediction Not checkable as stated
Dubois: Simulations will never fully eliminate the need for real-world AI training
“The problem is simulations are always going to be really hard and are not going to be truthful. So I think there will always need to be a certain, a little bit of training that will need to happen in the real world to make sure that the model realizes kind of …”
Insight
Dubois: RL becomes effective once base models possess strong world priors
“It seems that after crossing a certain scale of models that know basically everything about the world, and what we call, like, good priors about the world, It seems that reinforcement learning just started to work, and this is not only with LMS. Robotics seems…”
Insight
Dubois: AI model knowledge calibration generalizes across all domains
“When you have hallucination of LMs, if a model is really bad at saying that it doesn't know, that usually happens in every single domain. You won't have, like, one domain where the model is extremely calibrated about its knowledge, and another domain where it'…”
Insight
Dubois: Effective reinforcement learning pipelines prevent AI hallucinations caused by SFT
“So, so hallucination at least the intuition that people have is that it can come, for example, from SFT, and it can come from this, like, pursuing pipeline, but if you have good reinforcement in pipeline, that shouldn't happen too often.”
Prediction Not checkable as stated
Dubois: Model capacity does not limit AI performance in legal or medical fields
“But there's nothing, I would say, in the capacity of the model That is constraining the model to be as good at legal and like medical and like other domains.”
Prediction Not checkable as stated
Dubois: AI's coding discontinuity will permeate other verticals within two years
“Now the feeling of discontinuity will happen. It did happen three months ago with coding or four months ago with coding, and I think that will happen now in every other domains. Like most people are not feeling the same way Like the, like kind of the capabilit…”
Insight
Dubois: AI models outperform new employees initially but lack continual learning
“Right now, actually most models at day zero, if you just drop them in a company arguably they are more useful than most new employees. So they start higher at T zero. But then across time they are mostly constant because they don't really learn kind of company…”
Prediction Not checkable as stated
Dubois: General AI agent harnesses designed to endure will not work
“If you try to have, like, a general harness to, that will, like, sustain over time I don't think that will work.”
What-if
Dubois: Current models with optimized harnesses would feel like AGI
“If we froze the models that we have right now, and you really worked on the harness, and, like, maybe, like, we also spend more time, like, training with, like, a great harness I think people would really feel the AGI in every single domain, or could already f…”
Insight
Dubois: Last-mile integration is the main AI bottleneck, not raw intelligence
“I think most of the time, the bottleneck is the last mile.”
Prediction Not checkable as stated
Dubois: Horizontal AI model progress will not stop anytime soon
“Maybe one day when we stop making horizontal progress, which I don't think is anytime soon, maybe we will start focusing on that, but yeah, that's not what we're doing now.”
Prediction Not checkable as stated
Burazin: Every AI agent will require at least one sandbox
“My argument is that every agent will need at least one sandbox, sometimes more”
Prediction Held up
Burazin: AI agent scale creates high probability of impending CPU shortages
“I don't know it goes to the extreme to where GPUs are because that is very, very, very extreme. But it is quite highly, high probability that there will be shortages of CPUs going forward.”
Insight
Kolter: AI models do not get safer automatically by scaling up
“You can't just sort of trust models to get safer by getting bigger. You have to put in the work to actually make them safer. And this is, I think what a lot of AI companies are investing in. This is why we in fact do have models that are improving on these dim…”
Opinion
Kolter: Robotics AI is not yet ready for pure compute scaling
“Certain fields. I think things like robotics is still one. I don't think we're quite at the, let's just scale it up level with robotics yet. Some companies might argue we are. I don't think we are. I think we're still in the let's explore methods to find the r…”
Assertion Supported
Kolter: Adversarial prompts optimized on open-source LLMs break commercial models
“Once we had done that, we found that when you had these weird terms that you sort of flipped around to optimize one, to optimize the response for one model, you could just take those same exact strings you would optimize, paste them into a commercial model, an…”
Opinion
Kolter: AI agent benefits outweigh security risks if deployed with proper guardrails
“Yes, I think so, actually. I think if you run with proper guardrails, you know, we release guardrails for coding agents, for example. If you're on proper guardrails with proper sandboxing, and right now, yes, you probably also take some care to be a little bit…”
Disclosure
Kolter no longer writes code manually, relying entirely on AI agents
“I don't write code anymore. I do all my work now, and I do lots of, you know, I still do some research, right? It's entirely telling Codex what to do.”
Insight
The vast majority of modern AI intelligence comes from self-training
“I don't think people have properly internalized the fact that the vast majority of intelligence comes from self training effectively.”
Prediction Not checkable as stated
Zico Kolter: Current AI trajectory will yield capable systems without breakthroughs
“I think the current trajectory we're on is going to get us, even if there were no more breakthroughs, I think, you know, with the minor additions that we are doing right now, we will get to incredibly capable systems, even if we were to freeze things right now…”
Opinion
Text model scaling is one of humanity's top scientific discoveries
“The discovery That when you train big enough models on lots of text, and then turn, and then a little bit of additional sort of, you know, fine-tuning text, and then turn them loose to generate, that this generates long-form coherent thought. That was probably…”
Assertion Not checkable as stated
Anthropic's unreleased Claude Mythos model shows outsized cybersecurity capabilities
“Mythos is a unreleased frontier model. It's a general purpose model that was trained not specifically for cybersecurity or specifically for coding or specifically for software, but we have discovered what we believe to be outsized capabilities specifically in …”
Insight
AI product design overhang is now larger than model capability overhang
“As we get more and more powerful, I actually think the overhang in the product is bigger than in the model.”
Assertion Not checkable as stated
Rieseberg: AI models can execute week-long knowledge work tasks today
“The models we have today are actually quite capable. They're quite capable of running knowledge work of both of an extremely long time horizon, the kind of things that you give to someone and expect like a week later.”
Assertion Supported
Anthropic's Claude Cowork implements memory using plain text files
“It's in the harness, actually, and it's, like, often surprising to people when I talk to them how we, how we've implemented memory, because I think it maybe points at the simplicity underneath all of those models. Memory is just text files.”
Insight
Modern AI UI features exist for human reassurance, not model execution
“Most of the buttons you add and most of the product services you build are probably more for the human than they are for the model.”
Insight
Future AI product winners will be decided by UX, not superior models
“If someone beats me, Felix, at like building very good products, I suspect it's going to be not because they built a better model, but likely because they figured out a better user experience.”
Insight
Building hyper-specialized AI products is risky as models dynamically build infrastructure
“As the models get more and more capable, what I'm noticing inside my products and inside my work is that we're sort of like pulling back the edge cases we account for. And I mentioned earlier that memory is just a text file. If Claude needs a database, it will…”
Prediction Not checkable as stated
Rieseberg: Software creation skills will shift from code to human language
“My prediction is going to be That we are going to have a lot more software. That software is probably going to be slightly more specialized. I don't think everyone is going to build their own software. I think people will still build things and, like, share th…”
Prediction Not checkable as stated
Felix Rieseberg predicts AI progress is accelerating into larger capability gains
“We have reasons to believe the journey is accelerating so that the steps are going to get bigger and bigger.”
Insight
Current AI products are at the Nokia 3320 stage, not the iPhone
“I think, a thing I tell a lot of my colleagues is that we're really in the silly times of mobile phones, and then if we get really lucky, maybe what we're currently working on is like the Nokia 33 20, like a good phone, but it's not yet the smartphone. It's no…”
Prediction Not checkable as stated
Fully automated AI self-improvement will eliminate human bottlenecks and trigger breakthroughs
“The moment that we had this full automation, I would say we can close the loop of self-improvement and then it becomes the Like, you know, the problems become like, you know, mostly providing compute for these models to actually do what they want to do. And as…”
Prediction Not checkable as stated
AI model progress will alternate between pre-training and post-training breakthroughs
“We're going to be having a bit of a swing back and forth between pre-training and post-training.”
Insight
Post-training techniques cannot compensate for a weak base AI model
“Pre-training is still the foundation and like, you can never post-train your way out of a week-based model.”
Insight
Video data conveys physical world knowledge to AI more efficiently than text
“So because of that, like picking up a lot of knowledge about the word through language is just not really efficient. I don't want to say that it's impossible, but it's not efficient, you know, like to learn about gravity. If you kind of like, you know, have yo…”
Insight
Demonstrating that image training lowers text perplexity remains extremely difficult
“So it turned out to be a really, really good model, but it was like really hard to see that. Wow. You know, I train on images and then like Text perplexity goes down. That was hard to see. You know, like the fact that, you know, you train in native model and i…”
Insight
Jagged intelligence in AI is a structural flaw, not a patchable bug
“Not easy to pinpoint like specific things, but again, like, you know, this is just like my personal opinion and maybe I have colleagues and like the other people like sharing this with me, but I think we're underestimating how hard like jagged intelligence is …”