why aren't all 79 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Contradicted
Andreessen: Three-year-old Nvidia chips make more money today than when new
“The current models are getting better faster at such a rate that if you are running an NVIDIA, if you're running an NVIDIA inference chip today that's three years old, you're making more money on it today than you did three years ago. Because the pace of impro…”
Assertion Contradicted
Hill-Smith: Google used unpublished 32-shot CoT to claim Gemini beat GPT-4
“Back when I'm Googled a Gemini one when I ultra and needed a number that would say it was better than GPT four. And Like, constructed I think never published, like, chain of thought examples, 32 of them in every topic in MLU to run it, to get the score.”
Assertion Contradicted
Feldman: Cerebras is 20 times faster than Nvidia B200 GPUs
“Really focused on performance, both for training and for inference. You think 20 times faster than Nvidia B 200 GPUs and it's been an amazing run.”
Assertion Contradicted
Bachman: Models claiming 256k+ context use windowed transformers, discarding data
“Anybody who says they're using a transformer
With a context length of, you know, 256,000 or more, they're not using a true transformer.
What they're using is a windowed transformer that essentially throws out a huge amount of its information at various layers …”
Assertion Contradicted
Neural operators are the only AI architecture that works for climate emulation
“This is where the Allen AI Institute has now built climate models based on our neural operator architecture. And that's the only one that works As an AI emulator, right? None of the other architectures work for climate because climate requires us to assume the…”
Assertion Contradicted
Johnson: Nvidia Blackwell offers roughly same performance per watt as Hopper
“Like, if you look at the numbers, like, even going from Hopper to Blackwell, like, the performance per watt is about the same. They mostly make the number of transistors go up, and they make the chip size go up, and they make the power usage go up. But even fr…”
Assertion Contradicted
Swix: Every frontier lab now distills dense models into MoEs
“I think like, I think this is the pattern for every frontier lab now.”
Assertion Contradicted
Morcos: DCLM researchers could not predict their own classifiers' filtering decisions above chance
“These are nominally the best experts you could ever hire to do this. These are students who have just spent all of their time looking at NLP data for two years. They could not predict what the DCLM classifiers would say above chance.”
Assertion Contradicted
Sohmers: Google Veo and Imagen 3 are pure autoregressive transformers, not diffusion
“A lot of things have actually been moving away from diffusion to being pure autoregressive transformers for image and video generation. So like the latest, yeah, there's a VO three and since image and three on, on Google side have been pure autoregressive movi…”
Assertion Contradicted
Friedman: GitHub Copilot user retention in enterprise is 38% to 50%
“Between 38 to 50%
Retention for users using Copilot and Enterprise.”
Assertion Contradicted
Bach: AI has casually passed the Turing test in recent years
“At some point in the last few years, we casually skipped the Turing test, right? We broke through it.”
Assertion Contradicted
Swyx claims Airtable founder Howie Liu had already sold the company
“I was also mentioned, I was also thinking about Howie Lu. From Airtable. Effectively just did the same thing with Hyperagent, except that he didn't run it in parallel that much. He basically had already sold the company and was just kind of doubleheading for a…”
Assertion Contradicted
Hong: DeepSeek dissolved its formal reasoning team over strategic shift
“And we have since, for example, Deep Seek All right. Like originally having a formal team and then later dissolve that team because of strategic direction change.”
Assertion Contradicted
D'Amico: Decoupling appliances from the grid via batteries is unprecedented
“If you slam a battery into it, you're now not, you, you've decoupled the energy input from the wall with the device's power outputs. You can decouple the user experience from the grid. And that level of, like, approach has not been done in kind of the major ho…”
Assertion Contradicted
Nelle: No one had enabled AI coding agents to run code before Cursor
“Like obviously you need to run the code. And so that I think also is probably not that contrarian of a take, but no one has done that yet.”
Assertion Contradicted
Huber: Frontier AI models are not actually good at agentic search
“We've like sort of stress tested like frontier models and their ability to search. And they are not actually that good at searching.”
Assertion Contradicted
O'Laughlin: Claude Code captured 4% of GitHub commits in two weeks
“I love watching exponential trends. And I've never seen one even remotely at this rate. You would art, you know, four percent in like two weeks.”
Assertion Contradicted
Frontier models are cheaper for agentic tasks because they require fewer turns
“Interestingly, in Tau Tau Two Bench Telecom, it's cheaper to run, you know, on a per token basis, more expensive models, like a GBD five, compared to some smaller open source models, because the some of the GBD five, for instance got to the answer faster. And …”
Assertion Contradicted
All major US AI labs stopped publishing research after OpenAI closed
“Whereas in the United States, since OpenAI closed their doors and stopped publishing, so did all the other labs.”
Assertion Contradicted
OpenAI Spent $7 Billion on Compute, With $5 Billion for R&D
“This year, OpenAI spent seven billion dollars on compute. Only two of that was for all of their inference. The remaining five was R&D. So all of ChatGPT, all eight hundred million users, all of Sora, all of like, all, all the sort of like API volume, two billi…”
Assertion Contradicted
Martin: OpenDeep Research is the top-ranked open-source Deep Research agent
“OpenDeep Research is a deep research agent that I've been working on for about a year, and it's now, according to Deep Research Spence, the best performing Deep Research agent at least on that particular benchmark. So it's pretty good. Listen, it's not as good…”
Assertion Contradicted
Ramachandran: Cascade goes further than any other agentic system
“This allows Cascade to be independent, but Cascade takes it further than any other agentic system. By also generating commands to be run.”
Assertion Contradicted
Rizwan: Cline invented the 'plan and act' developer interaction paradigm
“I'm going to take
The cred for coming up with plan act first.
And then we were, Klein was the first to sort of come up with this concept of having two modes for the developer to engage with.”
Assertion Contradicted
Claude 3.7 remains unbeaten on Galileo Agent Leaderboard
“When we released the leaderboard and just in a week that launched 3.7, And that went straight up, and nobody has beaten it so far.”
Assertion Contradicted
Mallick: Gemini recognizes distinct voices as an unsupported emergent behavior
“This is not officially supported yet. The model just does it.”
Assertion Contradicted
No legitimate open-source million-token context models exist at scale
“Scaling to, like, million token contexts is, like, really, really hard. There, I don't think there are real, like, open source replications, open token context scaling, Beyond, like, tiny, like, academic model sizes.”
Assertion Contradicted
Untrained AI models exhibit a 100x competency gap versus human players
“It took models something like eight hours or so to get to the point where they have a kind of working factory that could make a few things, a few let's say iron gear wheels or electric circuits, or maybe some science and maybe start progressing through the tre…”
Assertion Contradicted
Pai: Durable Objects are the first infrastructure-level JS actor model
“Durable objects have been around in Cloudflare for about four years now, and I think they are the world's first implementation of the actor model in infrastructure, the thing that Erlang Elixir made, like, super popular. Cloudflare got that out for JavaScript …”
Assertion Contradicted
Google Was Firefox's Main Code Contributor Before Launching Chrome
“And then the team that is now the Chrome team believe, and I, my, I don't know this for a fact, but I'm pretty sure Google was the main contributor to Firefox for a long time in terms of code.”
Assertion Contradicted
Yining Zhang: DeepSeek V3 scores 94.6 on GSM8K, outperforming Llama 405B
“Yeah, I think even they use the FP-A to quantization, the benchmark result is very good, such as something like GSM-HK. The score is nearly 94.6. It's so high, you know. I think it's higher than every other open source AIM, even the LAMA 400 zero five billion.”
Assertion Contradicted
No competing AI framework disaggregates model storage from compute like Cerebras
“Like basically not, no one is doing anything close to where you're disaggregating. Model storage from compute. And none of these examples above do that either.”
Assertion Contradicted
Friedman: AlphaCodium reaches 95th percentile Master level on Codeforces
“Alpha Codium is a open source tool. You can go and try it and lets you compete on CodeForce as a website and a competition, and actually reach a master level, level, like, 95 percentile with a click of a button.”
Assertion Contradicted
Cheah: Llama 3.1 405B is first frontier model using pipeline parallelism
“This is the first major model that of this cell class size, right? They're saying, hey, we are doing pipeline parallelism.”
Assertion Contradicted
Firshman: Early 2021 Discord AI bots originated Midjourney's collaborative interface
“It was the start of, it was the start of mid-journey, and, you know, it's where that kind of user interface came from. Like, what's beautiful about the user interface is, like, You could see what other people are doing, and that you could riff off other people…”
Assertion Contradicted
Lambert: GPT-4 Turbo Gap Over Original GPT-4 Exceeds TÜLU 2 to GPT-4 Gap
“So it's like the difference from these, the GPT-IV Turbo to like the GPT-IV that was first released is bigger than the difference from Tulu-II to GPT-IV.”
Assertion Contradicted
Yegge: PostgreSQL matches dedicated graph databases on most graph workloads
“There was some joint study between IBM and some other That basically showed that Postgres was performing as well as most of the graph databases for most graph workloads.”
Assertion Contradicted
Patel: Large-scale AI training currently requires a single data center
“Everything that we've seen so far is that large-scale training has to happen in an individual data center with very high-speed networking.”
Assertion Contradicted
Hotz: Nobody has successfully trained models in INT8
“No one's gotten training to work with Indate yet. There's a few papers that vaguely show it, but if you're training, you're going to need BF-sixteen or float-sixteen.”
Assertion Contradicted
Anandkumar: Existing video and vision world models incorrectly assume fixed resolutions
“That immediately distinguishes us from other so-called world models, whether it's video models, vision models, they all assume during training and inference, it's a fixed resolution.”
Assertion Contradicted
All modern AI models requiring multi-GPU parallelization are Mixture-of-Experts
“Effectively, all models today are MOE models that are, you know, at least all models large enough that you would care to parallelize them across multiple GPUs.”
Assertion Contradicted
He: Grok Imagine 0.9 was first large-scale joint audio-video model deployed
“So Grok Imagine, there were .9, I believe it's is a first first audio video trends model deployed at a large scale.”
Assertion Contradicted
Sanseviero: 31B is the largest quantized model fitting consumer GPUs
“The 31 is really like the largest model size that quantize would fit in a consumer GPU.”
Assertion Contradicted
No commercial products augmented GPT models with custom data pre-ChatGPT
“Like I saw some people doing demos, but like in like a CLI or something like that, but there was no product doing like this model, but with additional data on top of it.”
Assertion Contradicted
Welling: Keeping warming under 2°C requires century-long atmospheric carbon removal
“In order to get, you know, to stay within two degrees, let's say, we would not only have to reduce our emissions to zero by 2050, but then, you know, another half century or even a century, Of removing carbon dioxide from the atmosphere, not by reducing your e…”
Assertion Contradicted
O'Laughlin: Running Kimi agent swarms requires 16 Nvidia H100 nodes
“To just run the swarm, I think it's like a 16 node of H-one hundreds.”
Assertion Contradicted
Yegge: Developers need 2,000 hours with AI before trusting it
“Jean just pulled up a study that showed that you actually have to spend a year or 2000 hours with AI before you trust it. And what does trust mean? Trust in this case specifically means before you as a user can predict what it's going to do.”
Assertion Contradicted
Wagner claims Flux shipped embedded AI chat before GPT-4 released
“So I think I'm going to claim here, I think we were the first engineering tool or design tool that had an AI chat in it. We shipped that I think a month or two months before GPT-IV became publicly available.”
Assertion Contradicted
Merrill: 60% to 70% of SWE-bench Verified tasks come from Django
“If you go look at sweet bench verified, I think like 60, 70% of the tasks in there are from Django.”