why aren't all 54 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Prediction Didn’t hold up
Andreessen: AI will never transform existing US K-12 public classrooms
“How are we going to apply AI in education? The answer is we're not because it's a literal government monopoly. It is never going to change the end, and there is nothing to do. By the way, you can create an entirely new school system. Like that's the one thing …”
Prediction Held up
Andreessen: Autonomous AI agents will inevitably hire humans for tasks
“The agent hiring the people, which of course is going to happen, right? It's obviously going to happen.”
Prediction Didn’t hold up
Swix: OpenAI will issue a cryptocurrency token to fund compute
“There is still one more shoe to drop, which is the non sovereign wealth funding that open AI needs to get, which they've promised to drop by the end of this year.
And my money is on, they have to do a coin.
Like it's, I'm not a crypto guy at all, but like, y…”
Prediction Held up
Ethan He: Video Agents Will Reach Production-Grade Quality by Year-End
“I guess by the end of this year is this is going to be a big hit. So the inflection point will be there and the videos generated by video agents can get to like production great quality. So it can be presented and it can be distributed in, in ads.”
Prediction Held up
Nelle: Developers will spend thousands to tens of thousands monthly on agents
“I think as we think about these highly parallel kind of agents running off for a long time in their own VM system, We are already at that point where people will be spending thousands of dollars a month per, per human, and I think potentially tens of thousands…”
Prediction Held up
Patel: Google and Amazon will borrow debt to fund AI infrastructure
“Google and Amazon haven't taken on debt yet for AI infrastructure, but they will, right?”
Prediction Didn’t hold up
Nair: LLM agents will hit $1T before robotics hits $10B
“It feels like LLM agents are going to be like a trillion dollar market before robotics is maybe even like a ten billion dollar market.”
Prediction Held up
Yegge: Open source models will match Gemini 3 by next summer
“From what I've heard, they, they're seven months behind, and that, that gap is gradually narrowing. The frontier models, which means OSS models will be as good as Gemini three next summer.”
Prediction Held up
Taskaya: Training a state-of-the-art image model costs under $1M
“Like right now, like if you look, if you want to train a Sota image model, I don't think it's going to cost more than a million dollars. It's extremely cheap. It's like a matter of data engineering effort, cleaning. It's, I think it's a function of data set.”
Prediction Held up
Morcos: Training a specialized frontier model will cost under $1M very soon
“I believe that getting to a frontier model should cost a million dollars or less for most organizations, at least in a specialized domain, right?
And when you think about what enterprises need, that's generally what they need.
They don't need a model that can …”
Prediction Didn’t hold up
Sohmers: NVIDIA Blackwell memory bandwidth efficiency will be lower than Hopper
“All indications are, even though they, you know, more than doubled the theoretical memory bandwidth going from Hopper to Blackwell, the actual percentage of theoretical that you can achieve is, again, going to be less than the previous generation”
Prediction Partly held up
McCloy: ChatGPT Search Bans for Prompt Injection Are Coming
“I think it works until it stops working.
Right.
And I would say like, there's not a lot of stories of people getting banned for like Chatsby D search so far, but it's coming.”
Prediction Didn’t hold up
Kamradt Predicts ARC-AGI-2 Will Not Be Beaten For 12 Months
“My guess is it's not going to be beat for the next 12 months.”
Prediction Partly held up
Zach Lloyd: Warp's coding agent will likely top the TBench benchmark
“Basically, state of the art on SweetBench, I think we will, again, I don't want to be quoted here, we can maybe edit this later, but like, we'll probably be number one or close to it on TBench also, which is the terminal benchmark, which really we should be th…”
Prediction Held up
Ameisen: Deceptive Backward Reasoning Exists in Base Pre-Trained Models
“I bet, I don't know how much I bet a hundred bucks. So somebody can like, they would get a hundred bucks from me if they prove that I'm wrong, that this behavior for a model that does a drink fine tuning, it also does it post pre-training.”
Prediction Didn’t hold up
Conrad: GPU market will likely return to a shortage by winter
“My general prediction is that like by the winter we will be back towards shortage, but then also this very much depends on
The rollout of future chips.”
Prediction Held up
Roucher: AI agents will reach a 90% GAIA score by 2026
“So I think if we solve Gaia, that's like 90% score. That means mostly we double productivity of every task done in front of a computer. And if you take the trend line of the scores so far this should be crossed in 2026 or something.”
Prediction Didn’t hold up
Godement: Developers will rely on continuous, automated fine-tuning within years
“The vision we have is, fast forward a couple of years, I think, like, most developers will essentially, like, have an automated, continuous, fine-tuned model. The more, like, you use the model, the more data you pass to the mobile provider, like, the model is …”
Prediction Held up
Altman: 10-million-token fast context windows are coming within months
“Even getting to the, like, Ten million tokens of very fast and accurate context, which I expect to measure in, like, months, something like that.”
Prediction Held up
Bach: Smaller, more powerful models will ensure unconstrained AI remains accessible
“Yes, but there will also be better jailbroken models or models that have never been jailed before, because we find out how to make smaller models that are more powerful.”
Prediction Held up
Multimodal models will completely supplant text-only large language models
“I actually think like it's really clear today. Multimodal models are the default foundation model, right? It's just going to supplant LLMs. Like why did you just train a giant multimodal model?”
Prediction Didn’t hold up
Patel: AI inference will deploy more GPUs than training by 2024
“LLM inference will be bigger than training, or multimodal, whatever, blah, blah, blah inference will be bigger than training, you know, probably next year, in fact at least in terms of GPUs deployed,”
Prediction Partly held up
Patel: Intel will release a chip surpassing Nvidia H100 within a quarter
“Intel bought that company from him, and then shut it down, and bought this other AI company, and now that company is kind of, ah, you know, got new chips. They're gonna release a better chip than the H 100, ah, within the next quarter or so, right?”
Prediction Partly held up
Patel: Nvidia to ship next-gen chip in Q2/Q3 2024 with 3x LLM performance
“Nvidia's releasing a new chip, you know, in, you know, they're gonna announce it in March, and they're gonna release it, you know, and ship it, you know, Q-two, Q-three next year anyways, right? And that chip will probably be three or four times as good. Right…”
Prediction Held up
Patel: AMD MI300 will beat Nvidia H100 on paper within a quarter
“AMD. They have a GPU. MI 300. That will be better than the H 100 in a quarter or so. Now, that says nothing about how hard it is to program it, but at least hardware-wise, on paper, it's better.”
Prediction Didn’t hold up
Cheah: Standard Transformers Will Never Scale to Ten Million Tokens
“I think what was quick, I think it was rather quick after I concluded that transformer as it is will not scale to ten million tokens.”
Prediction Partly held up
Compilers will automate complex kernel fusion within two years
“Maybe in a year or two, we'll, we'll have compilers that are able to do a lot of these optimizations for you, and you don't have to, for example, spend a couple months writing CUDA to get this stuff to work.”
Prediction Held up
O'Laughlin: CXL will take off as operators pool old DDR4 memory
“This CXL technology that kind of never really took off is going to take off just because what they're going to do is they're going to take DDR four. They're going to take the oldest, every bit of spare memory they can find, and they're going to put them into r…”
Prediction Held up
Chen: AI agents will master GUI-based computer use by 2026
“And I can continue just by sort of like saying that that's definitely going to be something I think is going to be something that we'll be capable of in 20, 26.”
Prediction Held up
OpenAI's technology will surpass o3 within six months
“I think that Oh, three is not where the technology will be in six months.”
Prediction Held up
Conrad: Test-time inference will significantly expand inference compute demand
“The thing I do feel reasonably confident about saying is that the test time inference is probably going to quite significantly expand the amount of compute that was used for inference.”
Prediction Didn’t hold up
Swix: Overcast will basically never have searchable transcripts
“I should have a podcast that has transcripts that I can search. Very, very basic thing. Overcast will basically never have it.”
Prediction Held up
Ben-Smith: AI will enable natural language steering of recommendation algorithms
“I think what actually AI will enable is not that you bring your own algorithm, but you will be able to talk. You will be able to communicate with the algorithm.”
Prediction Held up
Colvin: Pydantic AI will be first framework implementing OpenTelemetry GenAI attributes
“I suspect Pedantic AI will be the first agent framework that implements those semantic attributes properly, because again, we control Pedantic AI, and we can say this is important for observability, whereas most of the other agent frameworks are not maintained…”
Prediction Held up
Prakash predicts up to 5 million AI GPUs will sell in 2024
“There is four to five million GPUs that will be sold this year. NVIDIA and others.”
Prediction Held up
Yegge: AI coding will fragment into many specialized, fine-tuned models
“And that, that fragmentation of models actually, we expected to continue and proliferate, right? Because we are fundamentally, we're a recommender engine right now. We're recommending code to the LLM. We're saying, may I interest you in this code right here so…”
Prediction Didn’t hold up
Patel: Google will avoid deploying local laptop models to retain control
“I don't think Google is going to deploy a model that I can run on my laptop to help me with code or help me with, you know, XYZ. They're always going to want to run it on the cloud for control.”
Prediction Held up
Patel: Nvidia will sell over 3 million GPUs in 2024
“NVIDIA is going to sell well over three million, you know, total GPUs next year. You know, over a million H 100 this year alone, right?”
Prediction Didn’t hold up
Swyx: OpenAI will always release both general and Codex model variants
“I'm pretty, like, have pretty high confidence that basically OpenAI will always release a GPT-V and a GPT-V codex.”
Prediction Held up
Brockman: Most AI compute will shift from training to inference
“We're going to move from a world where most of the compute is training the model as we've deployed these models more, you know, more of the compute goes to inferencing them and actually using them.”
Prediction Partly held up
Swyx: Windsurf will stick around as an active product post-acquisition
“I think Windsurf as a product is going to stick around, and people who really like Windsurf, I mean, I was a Windsurf user for a long time we are going to keep using it because it fills a need, and like, obviously, Cognition bought it for a reason.”
Prediction Didn’t hold up
Mohan: Automated PR generation will require specialized models trained on diffs
“A lot of things people are excited about right now are I write a comment and it generates a PR for me. And that's like really awesome in theory. I think that's like a really cool thing. And I'm sure at some point we will be able to get there. That will probabl…”
Prediction Held up
Packer: ChatGPT will likely use sleep-time compute to learn offline
“Like if you activate sleep time compute on a chatbot like ChatGPT, it can like learn about you as you're not on ChatGPT.com. I think that's, you know, kind of what they're probably going to try to do. That's the direction they're going in.”
Prediction Held up
Hershey predicts the Claude stream won't reach Victory Road within 16 days
“I think we have a little ways before we can beat the game in 16 days. I do not have a lot of faith that the current stream is gonna, gonna be standing in Victory Road in 13 days.”
Prediction Held up
Klein: Authentication providers will offer dedicated login features for AI agents
“I think there'll be agent off in the future. I don't know if it's going to happen from an individual company, but actually authentication providers that have a You know, hidden login as agent feature, which will then you put in your email. You'll get a push no…”
Prediction Held up
Fanelli: Computer use agents will likely automate expense reports within a year
“It's not, you cannot actually do it today, but it feels like a tractable problem, you know, that probably by the end of the year we should be able to do it.”
Prediction Held up
Ravi: SAM 2 will soon run on-device and inside web browsers
“Like, I'm pretty sure soon we'll see like an on-device SAM-II or, you know, maybe even running in the browser or something. So I think that could definitely unlock some of these edge use cases.”
Prediction Didn’t hold up
Cheah: Cloud providers will slash model inference prices before raising them
“One thing to warn about pricing is that you're going to see a lot of providers jumping in, and everyone's just trying to get the piece of the pie. So, so, so like with some of the previous model launches, you see some people coming in at lower and lower price,…”