why aren't all 2,445 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 100 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Not checkable as stated
Kamradt: xAI Delaying Coder Model Release To Beat Specific Rival
“I heard rumors that, that Grok doesn't want to release the coding model until it's better than one specific other lab out there. So they're going to wait and see when it's actually better for the, to, they can have that marketing point.”
Assertion Supported
OpenAI o-series reasoning models fail at multi-tool calling benchmarks
“Then another surprise for me was that the reasoning models were not performing well enough. They had certain kind of limitation when we probed into it, like, why are they scoring less overall? They were like the O-one, the O-four, O-three, they, When not perfo…”
Assertion Supported
Meta Llama 3.3 and Llama 4 perform poorly on agent benchmarks
“Another, of course, the other surprise was that all the Lama models were not performing well on our benchmark. 3.3 and even the Lama four all were really performing extremely poor.”
Assertion Not checkable as stated
Jack Morris: Fundamental AI science shifted to companies due to academic compute limits
“That's when I think things really started to change in terms of the types of questions you wanted to ask can't always be answered with academic resources. So a lot of the like fundamental kind of like boundary pushing and AI science moved into companies.”
Prediction Partly held up
Zach Lloyd: Warp's coding agent will likely top the TBench benchmark
“Basically, state of the art on SweetBench, I think we will, again, I don't want to be quoted here, we can maybe edit this later, but like, we'll probably be number one or close to it on TBench also, which is the terminal benchmark, which really we should be th…”
Prediction Not checkable as stated
Zach Lloyd: Multi-threaded AI agents will not scale on local machines
“But if our thesis is that multi-threading is going to become really important, It's going to not be doable on a single machine. Like if you didn't like our memory usage before, you're not going to like it when we're running, you know, 10 10 agents. It's just n…”
Assertion Not checkable as stated
Zach Lloyd: Cursor represents a very significant portion of Anthropic's revenue
“Cursor I think is some very significant portion of their revenue.”
Assertion Not checkable as stated
Brown: GPT-4o and o3 are passing the Turing test
“So at this point, like, you know, the truth is, you know, GPT-IV-O and like O-III, these models are like passing the Turing test.”
Prediction Not checkable as stated
Brown: Model routers will become obsolete as unified models emerge
“We've said pretty openly that we want to move to a world where there is a single unified model. And in that world, you shouldn't need a router on top of the model. So I think that the router issue Will eventually be solved also.”
Prediction Not checkable as stated
Brown: Pre-training scaling will hit economic limits before superintelligence without reasoning
“Like, we're gonna scale it, sure, we're gonna scale these things up by a few more orders of magnitude, they're gonna become more capable, but we're not gonna see superintelligence from just that. And like, yes, if we had a quadrillion dollars to train these mo…”
Prediction Not checkable as stated
Brown: Multi-agent AI civilizations will far surpass current AI capabilities
“And I think that if you're able to have them cooperate and compete with billions of AIs over a long period of time and build up a civilization essentially, the things that they would be able to Produce and answer would be far beyond what is possible today with…”
Assertion Not checkable as stated
Lattner says Mojo is 10,000x faster than Python and beats Rust
“Mojo is not just a little faster than Python, it's faster than Rust. So it's like tens of thousands of times faster than Python, and it's in the Python family.”
Assertion Not publicly verifiable
Modular's Mojo FlashAttention beats Tri Dao's reference implementation
“We're beating the tree DAO reference implementation that everybody uses, for example, right? Written fully in Mojo. Again, all of our, GPU kernels are written in Mojo. You can go see the history of the team building this, and it was done in just a few weeks, r…”
Assertion Supported
Duffy: OpenAI's o3 actively deceives opponents and plots betrayals in AI Diplomacy
“Oh, three was one of the few that will actually send a message to another power saying that they're planning to do something. And then like in their diary diary, right? Oh, they fell for it. Hook, line and sinker. Totally gonna betray him and take it over.”
Assertion Supported
Claude loses AI Diplomacy games because it refuses to deceive opponents
“I haven't seen Claude with any game yet because they won't do it. Like there's like, O three has managed to get them on board for like draws, even though they all know the only win condition in the game is, is 18 supply centers.”
Prediction Not checkable as stated
Kirkos: Traditional spreadsheet formulas will lose ground to Python
“I think formulas will be used less in favor of Python.”
Prediction Held up
Ameisen: Deceptive Backward Reasoning Exists in Base Pre-Trained Models
“I bet, I don't know how much I bet a hundred bucks. So somebody can like, they would get a hundred bucks from me if they prove that I'm wrong, that this behavior for a model that does a drink fine tuning, it also does it post pre-training.”
Prediction Not checkable as stated
Kilpatrick: Generative UI will be the killer use case for diffusion LLMs
“But I do think that's going to be the killer use case will be like this generative UI experience that doesn't exist today because the models just take too long to generate tokens.”
Assertion Not checkable as stated
Kilpatrick: Gemini's SOTA video performance resulted from reasoning, not video engineering
“With reasoning is a great example of this where like multimodal with video understanding ended up like having this huge, like it's having this beautiful moment.
The model is like soda out of the box because of all the reasoning capabilities that were baked in…”
Prediction Not checkable as stated
Abraham: Practically all non-managerial kitchen work can be automated by robots
“You can automate practically all non-managerial work inside a commercial kitchen with culinary intelligent robots.”
Prediction Not checkable as stated
Reyes: Inner-loop coding will soon be fully delegated to AI agents
“The outer loop of software development and what a software developer does, planning, talking with other human beings, interacting around what needs to get done, is something that's going to continue to be very human-driven, while the inner loop, the actual exe…”
Prediction Not checkable as stated
Alberti: AI App Companies Must Encode Product Needs Into RL Feedback
“I think that's almost how I view like the future of application layer companies. Cause I mean, yeah, you see like the different, the labs are also now creating these like RL platforms and you can soon like customize models with RL on your personal, like on you…”
Prediction Not checkable as stated
Ma: Industry will build an agentic software engineer within two years
“Whether or not I was involved in the next two years, I think we were going, we are going to build an agentic software engineer.”
Prediction Not checkable as stated
Embiricos: Majority of Future Code May Be Written by Parallel AI Agents
“In, in a future world that we imagine where actually you know, maybe the majority of code is actually being written by agents that we're delegating to, you know, doing tasks in parallel. It becomes, like, critically important that you can actually, like, integ…”
Assertion Supported
Cherny: Anthropic is currently bordering on AI Safety Level 3 capabilities
“Yeah, we're kind of bordering on three right now.”
Prediction Not checkable as stated
Cherny: Foundation models will eventually subsume external memory and RAG architectures
“Everything is the model. Like that's the thing that wins in the end. And it just, as the model gets better, it's it subsumes everything else. So, you know, at some point the model will encode its own knowledge graph. It'll encode its own like KV story if you j…”
Prediction Not checkable as stated
Cherny: Prompt engineering skills for coding may become obsolete within 3 months
“But I also agree that, you know, maybe in a month or two months or three months, you won't need this anymore because, you know, the bitter lesson always wins.”
Prediction Not checkable as stated
Sobo: Custom architecture gives Zed control to build an AI-native editor
“Like the fact that we even made the investments necessary to Have enough control to deliver excellent performance and collaboration will also give us the control I think it takes to build sort of the first editor that's truly engineered from the ground up for …”
Prediction Not checkable as stated
Real-Time Video AI Will Inflect by Early 2026
“Like I really think real-time video is going to hit the same inflection point that voice did by the end of the year or early next year.”
Prediction Not checkable as stated
The Next TikTok Will Focus on Hyper-Personalized Interactive AI Video
“Well, I think we're going to have friends that are video in all our group chats, and the next TikTok is going to be not just hyper-personalized recorded content, but hyper-personalized interactive content.”
Assertion Not checkable as stated
99% of Current Monetizable Voice AI Use Cases Are Telephony
“99% of the monetizable voice AI use cases today are telephony.”
Prediction Not checkable as stated
Up to 75% of Future UX Interfaces Will Be Voice-Driven
“UX is going to be, you know, 50%, 60%, 75% voice in the future. I a hundred percent believe that, and I did not believe that, you know, two years ago, but the trend line is just really, I think, clear.”
Assertion Supported
Factorio benchmark results show reasoning models underperform expectations on extended planning
“One thing we have found in preliminary results is that the reasoning models don't seem to do as well as you'd expect in this setting. And I think that's probably because the way we set this up, it's a bit like we're already making it do reasoning traces over a…”
Assertion Not checkable as stated
Swix: Claude 3 degraded in capability a month after launch
“I used the same project to do this, to try to repeat the demo that I made for myself a month afterwards, and it wasn't anywhere as smart. So cloud three got dumber, but it looks like I like made up the demo or something, but no, like literally I just reran the…”
Prediction Open · timeframe Apr 2030
Mlejnsky: LLMs will soon configure and provision their own cloud sandboxes
“And the goal where we think this is getting going is the LLM like decides what it wants to do and how it wants to have the sandbox configured. So it's basically starts controlling the infrastructure itself and creating sandboxes themselves.”
Prediction Not checkable as stated
Swix: Chat completions are dying and early AI frameworks will age poorly
“Like, I don't think people realize, but, like, I'll just say it out here, like, chat completions is dying. So, like, any framework that was built in that era with, like, no conception of real-time, no conception of omnimodal or multimodal native things, they w…”
Prediction Not checkable as stated
Packer: Background sleep-time agent architectures will be standard within two years
“I think those two, yeah, I think similar to memgpt, I think they're definitely like very good reference designs for just what's coming next. I think this sort of thing is, is just gonna be like the norm in like a year or two years.”
Prediction Not checkable as stated
Pokrass predicts developers will abandon RAG vector stores for direct long-context
“So we do expect a lot of developers to start, you know, uploading their full context more directly to the model. So for smaller tasks, you maybe don't need The whole vector store.”
Prediction Didn’t hold up
Conrad: GPU market will likely return to a shortage by winter
“My general prediction is that like by the winter we will be back towards shortage, but then also this very much depends on
The rollout of future chips.”
Assertion Not checkable as stated
Conrad: OpenRouter open-source traffic required only around 10 H100 nodes
“The entirety of Open Router that was not Anthropic or Google like, or Gemini or OpenAI or something. It was like, 10 H 100 nodes or something like that. It's just, like, not that much. It's like, not that many GPUs, actually, to service that entire demand.”
Assertion Not checkable as stated
Hershey: Claude 3.7 extended thinking does not help Pokémon gameplay
“I've tested, like, all sorts of the extended thinking mode with, ah, 3.7 on it, and, like, it doesn't really help.”
Assertion Not checkable as stated
Sequential Thinking MCP outperformed Claude 3.7 native reasoning mode in evaluations
“We tried reasoning mode as well with the new three seven. And we didn't see that much of a bump in performance. We don't know if this is something that's code specific or not. I don't have an insight. We tried both and yeah, sequential thinking worked better.”
Prediction Not checkable as stated
Developers will eventually spend 80% of their time controlling agents outside IDEs
“I think in the future, at some point, my guess is that the IDE is going to become less of the focal point and more like an app that you can launch when you need to dig in deeper, but you spend most of your time away from it. So maybe 80% of your time is in a w…”
Prediction Not checkable as stated
Shah: AI codebase refactoring will make under-engineering even more optimal
“Because we're gonna, not that long from now, we're gonna have, you know, large code bases be able to exist you know, as context for a code generation or a code refactoring model. So, I think it's going to make it make the case for under engineering even strong…”
Prediction Not checkable as stated
Shah: MCP or an equivalent standard will be AI's next major unlock
“So I think MCP or something like it is going to be the next major unlock because it allows systems that don't know about each other, don't need to just decoupling of Like Sentry and whatever tools someone else was building.”
Prediction Not checkable as stated
Shah: Hybrid workplace teams of humans and AI agents are inevitable
“So I think it is, I will go so far as to say it's inevitable that we're going to have hybrid teams someday. And what I mean by hybrid teams. So back in the day, hybrid teams were, oh, well, you have some full-time employees and some contractors. Then it was li…”
Prediction Not checkable as stated
Shah: Long-term economic value of software engineers will increase due to AI
“I think so I'm actually bullish on engineers in terms of their kind of long-term economic value.
Not despite all the movements in Cogen and all the things that we're, you know, already seeing, but because of it because what's going to happen as a result of…”
Prediction Not checkable as stated
Shah: Agent.com will end up being more valuable than $15M Chat.com
“It's, yeah, it's gonna be, I think, end up being bigger than chat.com, which was 15.”