why aren't all 93 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Contradicted
Andreessen: Three-year-old Nvidia chips make more money today than when new
“The current models are getting better faster at such a rate that if you are running an NVIDIA, if you're running an NVIDIA inference chip today that's three years old, you're making more money on it today than you did three years ago. Because the pace of impro…”
Prediction Didn’t hold up
Andreessen: AI will never transform existing US K-12 public classrooms
“How are we going to apply AI in education? The answer is we're not because it's a literal government monopoly. It is never going to change the end, and there is nothing to do. By the way, you can create an entirely new school system. Like that's the one thing …”
Assertion Contradicted
Hill-Smith: Google used unpublished 32-shot CoT to claim Gemini beat GPT-4
“Back when I'm Googled a Gemini one when I ultra and needed a number that would say it was better than GPT four. And Like, constructed I think never published, like, chain of thought examples, 32 of them in every topic in MLU to run it, to get the score.”
Prediction Didn’t hold up
Swix: OpenAI will issue a cryptocurrency token to fund compute
“There is still one more shoe to drop, which is the non sovereign wealth funding that open AI needs to get, which they've promised to drop by the end of this year.
And my money is on, they have to do a coin.
Like it's, I'm not a crypto guy at all, but like, y…”
Assertion Contradicted
Feldman: Cerebras is 20 times faster than Nvidia B200 GPUs
“Really focused on performance, both for training and for inference. You think 20 times faster than Nvidia B 200 GPUs and it's been an amazing run.”
Assertion Contradicted
Bachman: Models claiming 256k+ context use windowed transformers, discarding data
“Anybody who says they're using a transformer
With a context length of, you know, 256,000 or more, they're not using a true transformer.
What they're using is a windowed transformer that essentially throws out a huge amount of its information at various layers …”
Assertion Contradicted
Neural operators are the only AI architecture that works for climate emulation
“This is where the Allen AI Institute has now built climate models based on our neural operator architecture. And that's the only one that works As an AI emulator, right? None of the other architectures work for climate because climate requires us to assume the…”
Prediction Didn’t hold up
Nair: LLM agents will hit $1T before robotics hits $10B
“It feels like LLM agents are going to be like a trillion dollar market before robotics is maybe even like a ten billion dollar market.”
Assertion Contradicted
Johnson: Nvidia Blackwell offers roughly same performance per watt as Hopper
“Like, if you look at the numbers, like, even going from Hopper to Blackwell, like, the performance per watt is about the same. They mostly make the number of transistors go up, and they make the chip size go up, and they make the power usage go up. But even fr…”
Assertion Contradicted
Swix: Every frontier lab now distills dense models into MoEs
“I think like, I think this is the pattern for every frontier lab now.”
Assertion Contradicted
Morcos: DCLM researchers could not predict their own classifiers' filtering decisions above chance
“These are nominally the best experts you could ever hire to do this. These are students who have just spent all of their time looking at NLP data for two years. They could not predict what the DCLM classifiers would say above chance.”
Prediction Didn’t hold up
Sohmers: NVIDIA Blackwell memory bandwidth efficiency will be lower than Hopper
“All indications are, even though they, you know, more than doubled the theoretical memory bandwidth going from Hopper to Blackwell, the actual percentage of theoretical that you can achieve is, again, going to be less than the previous generation”
Assertion Contradicted
Sohmers: Google Veo and Imagen 3 are pure autoregressive transformers, not diffusion
“A lot of things have actually been moving away from diffusion to being pure autoregressive transformers for image and video generation. So like the latest, yeah, there's a VO three and since image and three on, on Google side have been pure autoregressive movi…”
Prediction Didn’t hold up
Kamradt Predicts ARC-AGI-2 Will Not Be Beaten For 12 Months
“My guess is it's not going to be beat for the next 12 months.”
Prediction Didn’t hold up
Conrad: GPU market will likely return to a shortage by winter
“My general prediction is that like by the winter we will be back towards shortage, but then also this very much depends on
The rollout of future chips.”
Assertion Contradicted
Friedman: GitHub Copilot user retention in enterprise is 38% to 50%
“Between 38 to 50%
Retention for users using Copilot and Enterprise.”
Prediction Didn’t hold up
Godement: Developers will rely on continuous, automated fine-tuning within years
“The vision we have is, fast forward a couple of years, I think, like, most developers will essentially, like, have an automated, continuous, fine-tuned model. The more, like, you use the model, the more data you pass to the mobile provider, like, the model is …”
Assertion Contradicted
Bach: AI has casually passed the Turing test in recent years
“At some point in the last few years, we casually skipped the Turing test, right? We broke through it.”
Prediction Didn’t hold up
Patel: AI inference will deploy more GPUs than training by 2024
“LLM inference will be bigger than training, or multimodal, whatever, blah, blah, blah inference will be bigger than training, you know, probably next year, in fact at least in terms of GPUs deployed,”
Prediction Didn’t hold up
Cheah: Standard Transformers Will Never Scale to Ten Million Tokens
“I think what was quick, I think it was rather quick after I concluded that transformer as it is will not scale to ten million tokens.”
Assertion Contradicted
Swyx claims Airtable founder Howie Liu had already sold the company
“I was also mentioned, I was also thinking about Howie Lu. From Airtable. Effectively just did the same thing with Hyperagent, except that he didn't run it in parallel that much. He basically had already sold the company and was just kind of doubleheading for a…”
Assertion Contradicted
Hong: DeepSeek dissolved its formal reasoning team over strategic shift
“And we have since, for example, Deep Seek All right. Like originally having a formal team and then later dissolve that team because of strategic direction change.”
Assertion Contradicted
D'Amico: Decoupling appliances from the grid via batteries is unprecedented
“If you slam a battery into it, you're now not, you, you've decoupled the energy input from the wall with the device's power outputs. You can decouple the user experience from the grid. And that level of, like, approach has not been done in kind of the major ho…”
Assertion Contradicted
Nelle: No one had enabled AI coding agents to run code before Cursor
“Like obviously you need to run the code. And so that I think also is probably not that contrarian of a take, but no one has done that yet.”
Assertion Contradicted
Huber: Frontier AI models are not actually good at agentic search
“We've like sort of stress tested like frontier models and their ability to search. And they are not actually that good at searching.”
Assertion Contradicted
O'Laughlin: Claude Code captured 4% of GitHub commits in two weeks
“I love watching exponential trends. And I've never seen one even remotely at this rate. You would art, you know, four percent in like two weeks.”
Assertion Contradicted
Frontier models are cheaper for agentic tasks because they require fewer turns
“Interestingly, in Tau Tau Two Bench Telecom, it's cheaper to run, you know, on a per token basis, more expensive models, like a GBD five, compared to some smaller open source models, because the some of the GBD five, for instance got to the answer faster. And …”
Assertion Contradicted
All major US AI labs stopped publishing research after OpenAI closed
“Whereas in the United States, since OpenAI closed their doors and stopped publishing, so did all the other labs.”
Assertion Contradicted
OpenAI Spent $7 Billion on Compute, With $5 Billion for R&D
“This year, OpenAI spent seven billion dollars on compute. Only two of that was for all of their inference. The remaining five was R&D. So all of ChatGPT, all eight hundred million users, all of Sora, all of like, all, all the sort of like API volume, two billi…”
Assertion Contradicted
Martin: OpenDeep Research is the top-ranked open-source Deep Research agent
“OpenDeep Research is a deep research agent that I've been working on for about a year, and it's now, according to Deep Research Spence, the best performing Deep Research agent at least on that particular benchmark. So it's pretty good. Listen, it's not as good…”
Assertion Contradicted
Ramachandran: Cascade goes further than any other agentic system
“This allows Cascade to be independent, but Cascade takes it further than any other agentic system. By also generating commands to be run.”
Assertion Contradicted
Rizwan: Cline invented the 'plan and act' developer interaction paradigm
“I'm going to take
The cred for coming up with plan act first.
And then we were, Klein was the first to sort of come up with this concept of having two modes for the developer to engage with.”
Assertion Contradicted
Claude 3.7 remains unbeaten on Galileo Agent Leaderboard
“When we released the leaderboard and just in a week that launched 3.7, And that went straight up, and nobody has beaten it so far.”
Assertion Contradicted
Mallick: Gemini recognizes distinct voices as an unsupported emergent behavior
“This is not officially supported yet. The model just does it.”
Assertion Contradicted
No legitimate open-source million-token context models exist at scale
“Scaling to, like, million token contexts is, like, really, really hard. There, I don't think there are real, like, open source replications, open token context scaling, Beyond, like, tiny, like, academic model sizes.”
Assertion Contradicted
Untrained AI models exhibit a 100x competency gap versus human players
“It took models something like eight hours or so to get to the point where they have a kind of working factory that could make a few things, a few let's say iron gear wheels or electric circuits, or maybe some science and maybe start progressing through the tre…”
Assertion Contradicted
Pai: Durable Objects are the first infrastructure-level JS actor model
“Durable objects have been around in Cloudflare for about four years now, and I think they are the world's first implementation of the actor model in infrastructure, the thing that Erlang Elixir made, like, super popular. Cloudflare got that out for JavaScript …”
Prediction Didn’t hold up
Swix: Overcast will basically never have searchable transcripts
“I should have a podcast that has transcripts that I can search. Very, very basic thing. Overcast will basically never have it.”
Assertion Contradicted
Google Was Firefox's Main Code Contributor Before Launching Chrome
“And then the team that is now the Chrome team believe, and I, my, I don't know this for a fact, but I'm pretty sure Google was the main contributor to Firefox for a long time in terms of code.”
Assertion Contradicted
Yining Zhang: DeepSeek V3 scores 94.6 on GSM8K, outperforming Llama 405B
“Yeah, I think even they use the FP-A to quantization, the benchmark result is very good, such as something like GSM-HK. The score is nearly 94.6. It's so high, you know. I think it's higher than every other open source AIM, even the LAMA 400 zero five billion.”
Assertion Contradicted
No competing AI framework disaggregates model storage from compute like Cerebras
“Like basically not, no one is doing anything close to where you're disaggregating. Model storage from compute. And none of these examples above do that either.”
Assertion Contradicted
Friedman: AlphaCodium reaches 95th percentile Master level on Codeforces
“Alpha Codium is a open source tool. You can go and try it and lets you compete on CodeForce as a website and a competition, and actually reach a master level, level, like, 95 percentile with a click of a button.”
Assertion Contradicted
Cheah: Llama 3.1 405B is first frontier model using pipeline parallelism
“This is the first major model that of this cell class size, right? They're saying, hey, we are doing pipeline parallelism.”
Assertion Contradicted
Firshman: Early 2021 Discord AI bots originated Midjourney's collaborative interface
“It was the start of, it was the start of mid-journey, and, you know, it's where that kind of user interface came from. Like, what's beautiful about the user interface is, like, You could see what other people are doing, and that you could riff off other people…”
Assertion Contradicted
Lambert: GPT-4 Turbo Gap Over Original GPT-4 Exceeds TÜLU 2 to GPT-4 Gap
“So it's like the difference from these, the GPT-IV Turbo to like the GPT-IV that was first released is bigger than the difference from Tulu-II to GPT-IV.”
Assertion Contradicted
Yegge: PostgreSQL matches dedicated graph databases on most graph workloads
“There was some joint study between IBM and some other That basically showed that Postgres was performing as well as most of the graph databases for most graph workloads.”
Prediction Didn’t hold up
Patel: Google will avoid deploying local laptop models to retain control
“I don't think Google is going to deploy a model that I can run on my laptop to help me with code or help me with, you know, XYZ. They're always going to want to run it on the cloud for control.”
Assertion Contradicted
Patel: Large-scale AI training currently requires a single data center
“Everything that we've seen so far is that large-scale training has to happen in an individual data center with very high-speed networking.”