why aren't all 93 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Contradicted
Hotz: Nobody has successfully trained models in INT8
“No one's gotten training to work with Indate yet. There's a few papers that vaguely show it, but if you're training, you're going to need BF-sixteen or float-sixteen.”
Assertion Contradicted
Anandkumar: Existing video and vision world models incorrectly assume fixed resolutions
“That immediately distinguishes us from other so-called world models, whether it's video models, vision models, they all assume during training and inference, it's a fixed resolution.”
Assertion Contradicted
All modern AI models requiring multi-GPU parallelization are Mixture-of-Experts
“Effectively, all models today are MOE models that are, you know, at least all models large enough that you would care to parallelize them across multiple GPUs.”
Assertion Contradicted
He: Grok Imagine 0.9 was first large-scale joint audio-video model deployed
“So Grok Imagine, there were .9, I believe it's is a first first audio video trends model deployed at a large scale.”
Assertion Contradicted
Sanseviero: 31B is the largest quantized model fitting consumer GPUs
“The 31 is really like the largest model size that quantize would fit in a consumer GPU.”
Assertion Contradicted
No commercial products augmented GPT models with custom data pre-ChatGPT
“Like I saw some people doing demos, but like in like a CLI or something like that, but there was no product doing like this model, but with additional data on top of it.”
Assertion Contradicted
Welling: Keeping warming under 2°C requires century-long atmospheric carbon removal
“In order to get, you know, to stay within two degrees, let's say, we would not only have to reduce our emissions to zero by 2050, but then, you know, another half century or even a century, Of removing carbon dioxide from the atmosphere, not by reducing your e…”
Assertion Contradicted
O'Laughlin: Running Kimi agent swarms requires 16 Nvidia H100 nodes
“To just run the swarm, I think it's like a 16 node of H-one hundreds.”
Prediction Didn’t hold up
Swyx: OpenAI will always release both general and Codex model variants
“I'm pretty, like, have pretty high confidence that basically OpenAI will always release a GPT-V and a GPT-V codex.”
Assertion Contradicted
Yegge: Developers need 2,000 hours with AI before trusting it
“Jean just pulled up a study that showed that you actually have to spend a year or 2000 hours with AI before you trust it. And what does trust mean? Trust in this case specifically means before you as a user can predict what it's going to do.”
Assertion Contradicted
Wagner claims Flux shipped embedded AI chat before GPT-4 released
“So I think I'm going to claim here, I think we were the first engineering tool or design tool that had an AI chat in it. We shipped that I think a month or two months before GPT-IV became publicly available.”
Assertion Contradicted
Merrill: 60% to 70% of SWE-bench Verified tasks come from Django
“If you go look at sweet bench verified, I think like 60, 70% of the tasks in there are from Django.”
Assertion Contradicted
Lenz: AI21's Jamba is the first hybrid model architecture
“Since then, we've released several models, recent model lines in called Jamba, which I think the fascinating part about it is, is the first hybrid model. It's not just attention.”
Assertion Contradicted
Sohmers: Cray-2 was the last major system with balanced memory-to-compute ratio
“And if you look at sort of traditional big iron compute systems, the last like major compute, compute platform that had that balance of memory to compute ratio was the Cray two supercomputer.”
Assertion Contradicted
Swix: Copilot report estimates 60-70% of AI-generated code is checked in
“There's a report this morning from Copilot where they were estimating the key tabs on amount of code generated by a Copilot that is then left in code repos and checked in. And it's something like 60 to 70%.”
Prediction Didn’t hold up
Mohan: Automated PR generation will require specialized models trained on diffs
“A lot of things people are excited about right now are I write a comment and it generates a PR for me. And that's like really awesome in theory. I think that's like a really cool thing. And I'm sure at some point we will be able to get there. That will probabl…”
Assertion Contradicted
Mohan: Over 80% of software developers are on Windows
“A lot of people, once again, over 80% of developers are on Windows.”
Assertion Contradicted
The Lean theorem proving language has only about one million training tokens
“For Lean there are not enough data for Lean, right? There are, like, probably one million tokens in about Lean. Right now, and I know a lot of, like, professors and PhD researchers are trying to build the biggest lean data set in the future.”
Assertion Contradicted
Morris: Top AI graduate programs do not teach multi-node distributed training
“Oh, to be clear, they don't teach you anything, like anything, like if you see a paper coming out from even, you know, Stanford, they're probably the best school in AI if you had to choose. And it's not like they're learning how to do like multi-node distribut…”
Assertion Contradicted
Mohan: Over 80% of software developers are on Windows
“A lot of people, once again, over 80% of developers are on Windows.”
Assertion Contradicted
Ramachandran: Codeium is the only AI code assistant supporting Eclipse
“Like, we're still the only code assistant that has an extension of Eclipse. That's still true years in, right?”
Assertion Contradicted
Vibhu: Synthetic data research shows LLMs verify better than they generate
“A lot of the synthetic datagen papers, like orca-three, wizard-lm, they show that models are better at verifying output than generating output.”
Assertion Contradicted
Pullen: Most scraped open-source data consists of README and documentation updates
“When you scrape enough of it, most of open source is updating readmes and docs.”
Prediction Didn’t hold up
Cheah: Cloud providers will slash model inference prices before raising them
“One thing to warn about pricing is that you're going to see a lot of providers jumping in, and everyone's just trying to get the piece of the pie. So, so, so like with some of the previous model launches, you see some people coming in at lower and lower price,…”
Assertion Contradicted
Shulman: Bark was the first open-source transformer-based TTS model
“As far as I know there was no other certainly not in the open source text to speech that was kind of transformer based.”
Assertion Contradicted
Doshi Claims Playground AI Was First to Ship LCM Preview Feature
“I think we were the first company to integrate it... I think we were the first company to actually ship a quick LCM thing.”
Assertion Contradicted
Hotz: Consumer AMD GPUs lack peer-to-peer support
“If you have a consumer AMD GPU, they don't support peer to peer.”
Assertion Contradicted
Midha: Stanford Holds America's Second-Largest Longitudinal Patient Dataset Behind the VA
“Stanford is one of the only research facilities in America that has a longitudinal patient data set that's Larger at scale, I think it's at least twelve million patient lives. The only larger data set is the VA, the Veterans Affairs, you know, of America.”
Assertion Contradicted
Haneke: ASIC verification takes 3 to 4 times more resources than design
“My understanding is that the industry standard for design to verification in ASIC ASIC project is like one to three, one to four.”
Assertion Contradicted
Eifrem: Novo Nordisk graph deployment spans over 60 million documents
“Nowhere, nor this is one of the Like public case studies we have here over sixty million documents, you know, billions of notes and relationships use lots of kind of savvy NER and ER.”
Assertion Contradicted
Andreessen: AI Dungeon was the only public way to access GPT-3 for a year
“There was like a year where like the only way for a normal person to use GPT-III was in an AI dungeon.”
Assertion Contradicted
Swix: Sam Altman's Top AI Wish Is a Self-Completing To-Do List
“Do you know this is Sam Altman's number one ask for an AI app? It's the self-completing to-do list.”
Assertion Contradicted
Mirzadegan: Glean is the last incubation Kleiner Perkins has done
“I was working really closely with Arvind at Glean. It's actually the last incubation that we've done here.”
Assertion Contradicted
Ubl: Cognition and Cursor shipped RL fine-tunes of open-source models
“Just yesterday, I think we saw both Cognition, congrats, SWIX, and Cursor to ship RL fine tunes of unnamed open source models.”
Assertion Contradicted
AWS has about 20,000 private equity-backed customers
“And what's interesting is Amazon has about AWS has about 20,000 customers that are backed by PE firms.”
Assertion Contradicted
Brockman: OpenAI's Dota AI used only 300 million parameters
“And by the way, Dota was like a three hundred million parameter neural net. Tiny, tiny little insect brain, right?”
Assertion Contradicted
Google Gemini Flash-Lite ranks in top five on Galileo Agent Leaderboard
“In the, on the original leaderboard, we also have flashlight, which is, nobody talks about, I think, even now, but I think it's got a very decent score, which is, like, in the top five models. It's extremely cheap model, like, it's so dirt cheap that it makes …”
Assertion Contradicted
Ameisen: Golden Gate Claude feature specifically encoded awe of bridge's beauty
“We realized later on that it wasn't really like a Golden Gate Bridge feature. It was like being in awe at the beauty of the majestic Golden Gate Bridge, right?”
Assertion Contradicted
Swyx: Gemini Flash accounts for 50% of OpenRouter requests
“Gemini Flash, according to Open Router, is now 50% of their Open Router requests.”
Assertion Contradicted
Distilling sCM Requires Roughly Twice the Compute of Teacher Training
“One thing that they said in the paper, it's not here, but that that it, they, it takes about two X to compute to train
the
This consistency model from as a as a distillation of whatever they distilled from. So approximately twice the compute.”
Assertion Contradicted
Julien: BFCL v1 was the first dedicated LLM function calling benchmark
“So the first one that came out, like I said, in March was really the first of its kind to Make a function calling leaderboard.”
Assertion Contradicted
Prior video object segmentation models lacked error recovery mechanisms
“That actually is a big limitation of current models, current video object segmentation models. Don't allow any way to recover if the model makes a mistake.”
Assertion Contradicted
Tay: Most first-author papers by Jason Wei reach 1,000 annual citations
“Like, every single, so every single first author paper that, that, like, Jason writes in the last, has like, 1000 citations in one year. Like, no, I mean, not every, but like, most of it that he leads.”
Assertion Contradicted
Haisfield: Dylan Field tested generating Figma inside WebSim
“Dylan field actually posted this recently, like trying Figma in Figma or in web sim.”
Assertion Contradicted
Chintala: PyTorch is around 190,000 lines of code
“PyTorch is like a 190,000 lines of code or something at this point.”