why aren't all 1,786 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 41 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Supported
Schulhoff: LLMs Rely More on Prompt Structure Than Exemplar Labels
“There are a number of papers which have found that the label of the exemplar doesn't really matter, and the model reads the exemplars and cares more about structure than label.”
Assertion Not checkable as stated
Schulhoff: DSPy Beat 20 Hours of Manual Prompt Engineering in 10 Minutes
“And then I spent 20 hours prompt engineering for a task, and Dyspy beat me in 10 minutes, and that's when I changed my mind.”
Assertion Supported
Jamil: Writing in the Margins solves the lost-in-the-middle problem
“It improves the ability of any language model to extract relevant information, so solving the lost in the middle problem”
Assertion Not checkable as stated
Howard: 80% of top unique creators have unconventional or non-mainstream backgrounds
“Like, 80% of the time, I find out the person has a really unusual background. So, like, often they'll have, like, either they, like, came from poverty and, like, didn't get an opportunity to go to good school, or they, like, you know, had dyslexia and, you kno…”
Assertion Not checkable as stated
Howard: Decoder models must be far larger to match DeBERTa
“Now, the interesting thing is, you see, unlike Kaggle competitions, that decoder models still Are at least competitive with things like DiBerta VIII. But they have to be way bigger to be competitive with things like DiBerta VIII. And the only reason they are c…”
Assertion Supported
Scialom: Llama 3 405B is the best open-source model ever released
“At a high level, it's the best open source model ever. It's Better than GPT-IV. I mean, what version? But, by far, compared to the version originally released even now, I think there's maybe the last cloud Sonya FF-V and GPT-IV-Zero that are performing it.”
Assertion Supported
Tay: Zero-shot benchmark scores at 1B model scale are random chance
“Every time some people propose like this, they run like some zero-shot score on like some LM event harness or something like that, and you know like at one B scale, all the numbers are random, basically. Like all your bull kill, they're all like random chance …”
Assertion Supported
Albrecht: Benchmark performance differences vanish once ambiguous questions are cleaned
“The main takeaway from any of the, like, actual performance is like, once you fix these ambiguous examples, a lot of these benchmarks are really saturated. Like, I think it's important to look at like, you know, like when you're talking about performance on NL…”
Assertion Supported
Albrecht: AgentBench paper's appendix examples are actually incorrect solutions
“Like we were looking at the agent bench paper, I think just last week for our paper club. And one of the things that we noticed is that actually like both of the examples in the appendix that are given as like traces where it got it right. This is actually not…”
Assertion Not checkable as stated
Frankle: No Databricks enterprise customer asks for abstract reasoning AI
“I don't think I have a single customer that's asking to, you know, have AI solve abstract reasoning problems.”
Assertion Not checkable as stated
Conover: AI model developers are absolutely overfitting to public evaluation benchmarks
“And I think the work around over, you know, overfitting on the test, I think is like that. 100% is happening.”
Assertion Partly supported
Bach: LLMs Demonstrate Theory of Mind from Conversational Context
“When you ask the LLM to make inferences about your mental state based on the conversation that you have, it's able to demonstrate that it has a theory of mind.”
Assertion Contradicted
Bach: AI has casually passed the Turing test in recent years
“At some point in the last few years, we casually skipped the Turing test, right? We broke through it.”
Assertion Supported
Murphy: Five-minute voice calls cost 6.5 cents on Deepgram versus ElevenLabs.
“And then on the text-to-speech side, and doing something like this with an 11 labs would be about maybe a dollar 20. And just to give you an idea of comparison. So you can do a five minute call here for about six and a half cents.”
Assertion Not checkable as stated
Chintala: Aggregate open source AI model usage rivals GPT
“Maybe open source models are being as used as GPT is at this point in, like, all kinds of, in a very fragmented way. Like, in aggregate, all the open source models together are probably being used as much as GPT is. Maybe, you know, close to that.”
Assertion Partly supported
Retool survey found GPT-4V NPS was around 14 vs 45 for GPT-4
“GPT-IV.V. MPS, I want to say it was, like, 14 or something, like, it was, like, not high, actually. But the GPT-IV MPS thing was, like, 45 or something like that”
Assertion Supported
Lambert: RLHF has not been shown to improve underlying model benchmark capabilities
“RLHF is not that shown to improve capabilities yet. I think one of the fun ones is from the GPT-IV technical report. They essentially listed their kind of bogus evaluations, because it's a hilarious table, because it's like LSAT AP exams, and then like AMC-X a…”
Assertion Not checkable as stated
Liu: Cody matches GitHub Copilot completion acceptance rates using open-source StarCoder
“Like today, Cody uses StarCoder for inline completions, and with the benefit of the context that we provide, we actually show, like, comparable completion acceptance rate metrics. It's kind of like the standard metric that folks use to evaluate inline completi…”
Assertion Supported
Patel: Hugging Face libraries achieve only 15% MBU for inference
“Hugging Face's libraries are actually very inefficient, like incredibly inefficient for inference. You get like, 15% MBU on, on, on, on some configurations, like eight, eight, eight, eight, eight, eight, 100, and LLAMA-seventy-beat, you get like, 15%, which is…”
Assertion Not checkable as stated
Patel: Google targets 2027 to exit Broadcom partnership after 2025 attempt failed
“Their latest target to get away from Broadcom is twenty-twenty-seven, right? But like, you know, that's four years from now. Chip design cycle is four years. So they already tried to get away in twenty-twenty-five, and that failed.”
Assertion Not checkable as stated
Patel: AMD MI300 costs more than twice as much to manufacture as Nvidia H100
“In the case of amd their manufacturing costs for mi-thruhundred or more than twice that of h-one hundred and it only beats h-one hundred by a little bit from you know performance stuff i've seen”
Assertion Not checkable as stated
Patel: Microsoft's custom AI chip performs worse than Nvidia H100
“Microsoft's going to announce their chip soon. It's worse performance than the H-H-E-N-H-E-D but the cost effectiveness of it is, is better for Microsoft internally, just because they don't have to pay the Nvidia tax.”
Assertion Supported
Patel: TSMC's Arizona fab still depends on Taiwan for masks and shipping
“TSMC is building a fab in Arizona. It's quite a bit smaller than the fabs in, in, in Taiwan, but even ignoring that, those fabs still have to ship everything to Taiwan back anyways. And also they have to get what's called a mask from Taiwan and get sent to get…”
Assertion Supported
Royzen: Phind built the first internet-scale LLM RAG search in 2022
“And to the best of my knowledge, I think that's the first example that I'm aware of a LLM search engine model that's effectively connected to, like, a large enough index that I would consider, like, an internet scale. So, so I think we were the first to releas…”
Assertion Not checkable as stated
Royzen: Users switch to Phind when ChatGPT-4 fails on code
“What really shocks us is that a lot of the people who do that they're coming from ChatGPT. So they tried it in ChatGPT with ChatGPT-IV. It didn't work. Maybe it required like some multi-step reasoning. Maybe it required to like, Some internet context or someth…”
Assertion Not checkable as stated
Royzen: Training on code unlocked general spatial and temporal reasoning
“We've seen emerging capabilities in the find model, whereby training it on high quality code, it can actually, like, reason better. It went from not being able to solve like, World problems where like riddles where like with like temporal and like low, like pl…”
Assertion Supported
Howard: Experiments show LLMs can memorize full datasets in one epoch
“And so we ran a bunch of experiments, and all of them supported the hypothesis that it was memorizing the data set in a single thing at once.”
Assertion Supported
RWKV Matches GPT-NeoX Performance at Equal Parameter and Data Scales
“RWKV is a modern recursive neural network with transformer-like level of LMM performance, which can be trained in a transformer mode. And this part has already been benchmarked against GPT-NeoX in the paper, And it has similar training performance compared to …”
Assertion Supported
RWKV Architecture Is Proven to Scale to Any Parameter Size
“What we have already proven is that it can be scaled and trained by a transformer.
How I do so, we'll cover later.
And this can be scaled to as many parameters as we want.”
Assertion Supported
Hotz: Tinygrad runs all ML models with only 25 primitive operations
“Tiny grad is, we are going to make a risk offset for all ML models. And yeah, it can run all ML models with basically 25 instead of the two 50 of XLA or PrimTorch. So about 10 X less complex.”
Assertion Supported
Swyx claims Jeff Dean just left Google
“Jeff Dean just left Google.”
Assertion Contradicted
Swyx claims Airtable founder Howie Liu had already sold the company
“I was also mentioned, I was also thinking about Howie Lu. From Airtable. Effectively just did the same thing with Hyperagent, except that he didn't run it in parallel that much. He basically had already sold the company and was just kind of doubleheading for a…”
Assertion Not checkable as stated
Jenik: Accelerated Understanding has trained physics models up to 1-trillion parameters
“We've trained, done hundreds of training runs. We've trained up to a trillion parameter models. So we've really shown the stuff to take off.”
Assertion Supported
Lie: Cerebras demoed GPT running at over 4,400 TPS at Hot Chips
“We here in this demo that we gave at hot chips we're showing GPT OSS running at over 4000 400 TPS, which is just mind blowing.”
Assertion Supported
Lie: Trillion-parameter models require thousands of Groq LPUs for weights
“To run a frontier level model, like, let's say, a few trillion parameters, you need thousands and thousands of Grok LPUs just to hold the weights, right?”
Assertion Supported
FourCastNet matches supercomputer weather accuracy 10,000 times faster on consumer GPUs
“To our surprise, we found that it's not only, you know, accurate, it's almost as close to the what the traditional weather models can do accurately, but also tens of thousands of times faster. So what would take a big supercomputer to run can now be run. And w…”
Assertion Supported
FourCastNet predicted Hurricane Lee's landfall days earlier than standard weather models
“For instance, our forecast net was able to correctly predict that the hurricane making the landfall several days earlier compared to the standard weather forecasting models.”
Assertion Supported
Spherical AI weather models achieve longer autoregressive rollouts than flat models
“These models that we have are able to do the longest rollouts compared to any of the other weather models that completely ignore spherical assumption and a range of other things.”
Assertion Supported
FourCastNet models trained on six-hour steps produce stable months-long forecasts
“To predict for the next six hours. And a little bit of multi-step fine tuning... Now we are showing for several months that it's able to do that.”
Assertion Supported
Park: Training on randomized controlled trials improves AI human-behavior prediction
“That by collecting a lot of these randomized control trials that are really well designed, we can make significant improvement in models capability to predict human behaviors.”
Assertion Not checkable as stated
Park: Simile observes empirical scaling laws when modeling human behavior
“What we are seeing is at Simile, so we do post-train our own model. The thing that we're actually seeing is the early glimpse of scaling law in simulations. The more data about humans and more compute you ingest, you actually start to get predictive and predic…”
Assertion Supported
Krentsel: OpenClaw, Pi, and Claude Code hardcode static agent policies
“These are all policy decisions that are static, that are defined for OpenClaw, or for Pi, or for Clawed code, if you look at their Source code. And so that is the kind of, that is the policy of what an agent is, the tools it can use, the skills it has, how it …”
Assertion Open · timeframe Aug 2026
Krentsel: Exo autonomously modified its code to inspect Pokémon game RAM
“We've had XO running, playing, playing Pokemon. And while it's running, the system itself decided to try inspecting the like RAM of the game and then went and mapped the RAM to, and people have reversed in the past, people have reverse engineered this manually…”
Assertion Open · timeframe Aug 2026
Krentsel: Exo autonomously re-architected its Discord adapter, cutting costs by 96%
“We asked it, Hey, I noticed, I asked, Hey, but how much did the last message cost in the discord adapter?
And it was like, it was.
It's like, are you serious?
16 cents.
That's actually crazy.
Like.
Go work on driving that down.
And so it went and re-architecte…”
Assertion Supported
McPartlon: Chai-2 generated binders for 25 targets with 20% hit rate
“So we designed antibodies to 50 targets for that paper, got binders to about half of them with being on average around a 20% hit rate for binding.”
Assertion Not checkable as stated
Patil: Chai-2 crossed the performance threshold for antibody design last year
“The ultimate goal here is to design medicines, right, and design new molecules, and I think CHI-2 really crossed the threshold of performance for doing that with antibodies a year ago.”
Assertion Supported
McPartlon: Chai-2 achieved 0.33 angstrom error in cryo-EM testing
“And in this case, it was a 0.33 angstrom error, which is one third the width of an atom.”
Assertion Supported
McPartlon: AlphaFold-Multimer gets antibody-antigen prediction right only 11% of the time
“Not really like outfold to got like, I think, 11%, the multiple version of this got like 11% of antibody antigen prediction cases. Correct. That means 90% of the time it's wrong.”