The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 1,046 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Supported
Agarwal: Filtered 9B Synthetic Data Outperforms 27B Self-Generated Data
“One thing we found consistently, so here what we had two models, nine Gemma, nine B and Gemma, 27 B, and we found consistently that actually generating data from nine B in a compute match setting is always better, even better for distilling or actually improvi…”
Rishabh Agarwal Mar 23, 2025 ▶ 17:41 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Assertion Supported
Cursor Composer solves Convex benchmarks but fails on alternative backends
“We did notice that I mean, with convex, it pretty much autonomously just solves the first two tasks. It has a few round trips on like some errors that are only show up and playing with the front end. And then it's able to complete this files task and kind of g…”
Sujay Jayakar Mar 19, 2025 ▶ 9:36 Fullstack-Bench: The Eval for Coding Agents — with Sujay Jayakar, Chief Scientist, Convex
Assertion Supported
AI models struggle debugging Supabase RLS recursion compared to procedural code
“The particular example was like RLS rules and Supabase where debugging like an infinite loop for infinite recursion for the RLS rules was something that the models just really struggled with in a way that we didn't see for procedural code.”
Sujay Jayakar Mar 19, 2025 ▶ 29:57 Fullstack-Bench: The Eval for Coding Agents — with Sujay Jayakar, Chief Scientist, Convex
Prediction Held up
Roucher: AI agents will reach a 90% GAIA score by 2026
“So I think if we solve Gaia, that's like 90% score. That means mostly we double productivity of every task done in front of a computer. And if you take the trend line of the scores so far this should be crossed in 2026 or something.”
Aymeric (Emmerich) Feb 13, 2025 ▶ 16:12 smol agents are all you need
Assertion Supported
DeepSeek-R1 researchers found MCTS and Process Reward Models were not useful
“R-one specifically said, yes, we tried MCTS. Yes, we tried PRMs. And none of that is useful.”
Shawn Wang Jan 24, 2025 ▶ 10:32 The Unreasonable Effectiveness of Reasoning Distillation: using DeepSeek R1 to beat OpenAI o1
Assertion Supported
Zhang: XGrammar outperforms Outlines and is integrated into TensorRT-LLM
“And I think Xgrammar's performance is better than the outline's, and also in the TensorFlow RTLM, the latest release, TensorFlow RTLM also integrates Xgrammar as the backend for the constructed coding.”
Yining Zhang Jan 19, 2025 ▶ 41:22 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Assertion Supported
Neubig: SWE-bench Scores Are Inflated By Training Data Contamination
“Sweebench is on popular open source repos and all of these popular open source repos were included in the training data for all of the language models. And so, the language models already know these repos. In some cases, the language models already know the in…”
Graham Neubig Dec 25, 2024 ▶ 32:49 Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
Assertion Supported
Ben Allal: Hugging Face SmolLM2-1.7B outperforms Llama 3.2 models
“So it's a series of three models, which are the best in class in each model size. For example, our 1.7 B model outperforms Lama one B and also .2.”
Loubna Ben Allal Dec 24, 2024 ▶ 22:31 Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
Assertion Supported
Reddy: Chai Discovery's open-source Chai-1 model outperforms Google's AlphaFold 3
“We're lucky to work with the folks at Chai Discovery who just released Chai One, which is open source model that outperforms Alpha Fold Three.”
Pranav Reddy Dec 21, 2024 ▶ 7:44 The State of AI Startups in 2024 [LS Live @ NeurIPS]
Assertion Supported
Cerebras WSE-3 runs Llama inference 70x faster than NVIDIA GPUs
“Cerebris came out that the wafer scale engine three can serve llama 70 B at 2.1 thousand sorry, 202,100 tokens per second and serves llama four or five B at nearly 1000 tokens per second. So this, you know, to give you an understanding, like this is about 70 t…”
Sarah Chieng Dec 7, 2024 ▶ 3:06 [Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras)
Assertion Contradicted
Friedman: GitHub Copilot user retention in enterprise is 38% to 50%
“Between 38 to 50% Retention for users using Copilot and Enterprise.”
Itamar Friedman Dec 2, 2024 ▶ 37:04 0 to over $8M ARR in 2 months as a Claude Wrapper (Bolt.new, Qodo)
Assertion Supported
Friedman: AlphaCodium boosts OpenAI o1, proving o1 lacks true System 2
“We took their all one preview with Alpha Codium and did better. Like it just shows like, and there is a big difference between the preview and the IOI. It shows, like, that these models are not still system two thinkers, and there's a big difference.”
Itamar Friedman Dec 2, 2024 ▶ 42:48 0 to over $8M ARR in 2 months as a Claude Wrapper (Bolt.new, Qodo)
Assertion Supported
SWE-Bench public test splits enable trivial cheating via runtime PR retrieval
“The entire test split here is public. So you can do things like just overfit to the patches in the test set. You can do things like, let me add at runtime, pull the PR and just get the answer and just use it.”
Jesse Hu Oct 19, 2024 ▶ 15:00 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
Assertion Supported
Houston: Groq and Cerebras outperform Nvidia on latency
“There's also, like, non-NVIDIA stacks, like the Grok, or Cerebris, or some of these custom silicon companies that are super interesting, and all, and outperformed the NVIDIA stack in terms of latency and things like that.”
Drew Houston Oct 18, 2024 ▶ 52:59 Building the Silicon Brain - Drew Houston of Dropbox
Assertion Supported
Molmo Outperforms Gemini 1.5 and Claude 3.5 Sonnet With 1M Samples
“They can get better than Gemini, 1.5, better than Claude, 3.5 sonnet, better than GPT for V at a much smaller size with about a million samples of data, which is very impressive, right?”
Vibhu Sapra Oct 13, 2024 ▶ 4:23 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Prediction Didn’t hold up
Godement: Developers will rely on continuous, automated fine-tuning within years
“The vision we have is, fast forward a couple of years, I think, like, most developers will essentially, like, have an automated, continuous, fine-tuned model. The more, like, you use the model, the more data you pass to the mobile provider, like, the model is …”
Olivier Godement Oct 4, 2024 ▶ 26:46 Building AGI in Real Time (OpenAI Dev Day 2024)
Assertion Supported
Pullen: CoScene's Genie outperforms OpenAI o1 out of the box on SWE-bench
“So it was obviously great to see, like, we still are better than O-one out of the box. You know, even with an older model, and I'm sure that that, that Delta will continue to grow once we're able to train O-one and once we've done more work on our dataset usin…”
Alistair Pullen Oct 4, 2024 ▶ 1:13:04 Building AGI in Real Time (OpenAI Dev Day 2024)
Prediction Held up
Altman: 10-million-token fast context windows are coming within months
“Even getting to the, like, Ten million tokens of very fast and accurate context, which I expect to measure in, like, months, something like that.”
Sam Altman Oct 4, 2024 ▶ 2:05:30 Building AGI in Real Time (OpenAI Dev Day 2024)
Assertion Supported
Schulhoff: LLMs Rely More on Prompt Structure Than Exemplar Labels
“There are a number of papers which have found that the label of the exemplar doesn't really matter, and the model reads the exemplars and cares more about structure than label.”
Sander Schulhoff Sep 20, 2024 ▶ 26:41 The Ultimate Guide to Prompting - with Sander Schulhoff from LearnPrompting.org
Assertion Supported
Jamil: Writing in the Margins solves the lost-in-the-middle problem
“It improves the ability of any language model to extract relevant information, so solving the lost in the middle problem”
Umar Jamil Sep 19, 2024 ▶ 22:54 [Paper Club] Writing in the Margins: Chunked Prefill KV Caching for Long Context Retrieval
Assertion Supported
Scialom: Llama 3 405B is the best open-source model ever released
“At a high level, it's the best open source model ever. It's Better than GPT-IV. I mean, what version? But, by far, compared to the version originally released even now, I think there's maybe the last cloud Sonya FF-V and GPT-IV-Zero that are performing it.”
Thomas Scialom Jul 23, 2024 ▶ 37:43 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Assertion Supported
Tay: Zero-shot benchmark scores at 1B model scale are random chance
“Every time some people propose like this, they run like some zero-shot score on like some LM event harness or something like that, and you know like at one B scale, all the numbers are random, basically. Like all your bull kill, they're all like random chance …”
Yi Tay Jul 5, 2024 ▶ 1:46:43 The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
Assertion Supported
Albrecht: Benchmark performance differences vanish once ambiguous questions are cleaned
“The main takeaway from any of the, like, actual performance is like, once you fix these ambiguous examples, a lot of these benchmarks are really saturated. Like, I think it's important to look at like, you know, like when you're talking about performance on NL…”
Josh Albrecht Jun 25, 2024 ▶ 1:01:43 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Assertion Supported
Albrecht: AgentBench paper's appendix examples are actually incorrect solutions
“Like we were looking at the agent bench paper, I think just last week for our paper club. And one of the things that we noticed is that actually like both of the examples in the appendix that are given as like traces where it got it right. This is actually not…”
Josh Albrecht Jun 25, 2024 ▶ 1:06:05 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Assertion Contradicted
Bach: AI has casually passed the Turing test in recent years
“At some point in the last few years, we casually skipped the Turing test, right? We broke through it.”
Joscha Bach Apr 27, 2024 ▶ 1:45:41 This World Does Not Exist — Joscha Bach, Karan Malhotra, Rob Haisfield (WorldSim, WebSim, Liquid AI)
Assertion Partly supported
Bach: LLMs Demonstrate Theory of Mind from Conversational Context
“When you ask the LLM to make inferences about your mental state based on the conversation that you have, it's able to demonstrate that it has a theory of mind.”
Joscha Bach Apr 27, 2024 ▶ 1:03:48 This World Does Not Exist — Joscha Bach, Karan Malhotra, Rob Haisfield (WorldSim, WebSim, Liquid AI)
Prediction Held up
Bach: Smaller, more powerful models will ensure unconstrained AI remains accessible
“Yes, but there will also be better jailbroken models or models that have never been jailed before, because we find out how to make smaller models that are more powerful.”
Joscha Bach Apr 27, 2024 ▶ 1:28:21 This World Does Not Exist — Joscha Bach, Karan Malhotra, Rob Haisfield (WorldSim, WebSim, Liquid AI)
Assertion Supported
Murphy: Five-minute voice calls cost 6.5 cents on Deepgram versus ElevenLabs.
“And then on the text-to-speech side, and doing something like this with an 11 labs would be about maybe a dollar 20. And just to give you an idea of comparison. So you can do a five minute call here for about six and a half cents.”
Damien Murphy Apr 6, 2024 ▶ 16:52 Personal AI Meetup - Bee, BasedHardware, LangChain LangFriend, Deepgram EmilyAI
Prediction Held up
Multimodal models will completely supplant text-only large language models
“I actually think like it's really clear today. Multimodal models are the default foundation model, right? It's just going to supplant LLMs. Like why did you just train a giant multimodal model?”
David Luan Mar 27, 2024 ▶ 28:05 Why Google failed to make GPT-3 -- with David Luan of Adept
Assertion Partly supported
Retool survey found GPT-4V NPS was around 14 vs 45 for GPT-4
“GPT-IV.V. MPS, I want to say it was, like, 14 or something, like, it was, like, not high, actually. But the GPT-IV MPS thing was, like, 45 or something like that”
David Hsu Feb 7, 2024 ▶ 51:30 The State of AI in production — with David Hsu of Retool
Assertion Supported
Lambert: RLHF has not been shown to improve underlying model benchmark capabilities
“RLHF is not that shown to improve capabilities yet. I think one of the fun ones is from the GPT-IV technical report. They essentially listed their kind of bogus evaluations, because it's a hilarious table, because it's like LSAT AP exams, and then like AMC-X a…”
Nathan Lambert Jan 11, 2024 ▶ 59:53 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Prediction Didn’t hold up
Patel: AI inference will deploy more GPUs than training by 2024
“LLM inference will be bigger than training, or multimodal, whatever, blah, blah, blah inference will be bigger than training, you know, probably next year, in fact at least in terms of GPUs deployed,”
Dylan Patel Dec 5, 2023 ▶ 14:40 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Prediction Partly held up
Patel: Intel will release a chip surpassing Nvidia H100 within a quarter
“Intel bought that company from him, and then shut it down, and bought this other AI company, and now that company is kind of, ah, you know, got new chips. They're gonna release a better chip than the H 100, ah, within the next quarter or so, right?”
Dylan Patel Dec 5, 2023 ▶ 41:01 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Prediction Partly held up
Patel: Nvidia to ship next-gen chip in Q2/Q3 2024 with 3x LLM performance
“Nvidia's releasing a new chip, you know, in, you know, they're gonna announce it in March, and they're gonna release it, you know, and ship it, you know, Q-two, Q-three next year anyways, right? And that chip will probably be three or four times as good. Right…”
Dylan Patel Dec 5, 2023 ▶ 44:00 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Assertion Supported
Patel: Hugging Face libraries achieve only 15% MBU for inference
“Hugging Face's libraries are actually very inefficient, like incredibly inefficient for inference. You get like, 15% MBU on, on, on, on some configurations, like eight, eight, eight, eight, eight, eight, 100, and LLAMA-seventy-beat, you get like, 15%, which is…”
Dylan Patel Dec 5, 2023 ▶ 18:40 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Prediction Held up
Patel: AMD MI300 will beat Nvidia H100 on paper within a quarter
“AMD. They have a GPU. MI 300. That will be better than the H 100 in a quarter or so. Now, that says nothing about how hard it is to program it, but at least hardware-wise, on paper, it's better.”
Dylan Patel Dec 5, 2023 ▶ 41:13 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Assertion Supported
Patel: TSMC's Arizona fab still depends on Taiwan for masks and shipping
“TSMC is building a fab in Arizona. It's quite a bit smaller than the fabs in, in, in Taiwan, but even ignoring that, those fabs still have to ship everything to Taiwan back anyways. And also they have to get what's called a mask from Taiwan and get sent to get…”
Dylan Patel Dec 5, 2023 ▶ 56:50 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Assertion Supported
Royzen: Phind built the first internet-scale LLM RAG search in 2022
“And to the best of my knowledge, I think that's the first example that I'm aware of a LLM search engine model that's effectively connected to, like, a large enough index that I would consider, like, an internet scale. So, so I think we were the first to releas…”
Michael Royzen Nov 3, 2023 ▶ 16:07 Beating GPT-4 with Open Source Models - with Michael Royzen of Phind
Assertion Supported
Howard: Experiments show LLMs can memorize full datasets in one epoch
“And so we ran a bunch of experiments, and all of them supported the hypothesis that it was memorizing the data set in a single thing at once.”
Jeremy Howard Oct 20, 2023 ▶ 41:14 The End of Finetuning — with Jeremy Howard of Fast.ai
Prediction Didn’t hold up
Cheah: Standard Transformers Will Never Scale to Ten Million Tokens
“I think what was quick, I think it was rather quick after I concluded that transformer as it is will not scale to ten million tokens.”
Eugene Cheah Aug 31, 2023 ▶ 20:25 RWKV: Reinventing RNNs for the Transformer Era
Assertion Supported
RWKV Matches GPT-NeoX Performance at Equal Parameter and Data Scales
“RWKV is a modern recursive neural network with transformer-like level of LMM performance, which can be trained in a transformer mode. And this part has already been benchmarked against GPT-NeoX in the paper, And it has similar training performance compared to …”
Eugene Cheah Aug 31, 2023 ▶ 31:59 RWKV: Reinventing RNNs for the Transformer Era
Assertion Supported
RWKV Architecture Is Proven to Scale to Any Parameter Size
“What we have already proven is that it can be scaled and trained by a transformer. How I do so, we'll cover later. And this can be scaled to as many parameters as we want.”
Eugene Cheah Aug 31, 2023 ▶ 37:52 RWKV: Reinventing RNNs for the Transformer Era
Prediction Partly held up
Compilers will automate complex kernel fusion within two years
“Maybe in a year or two, we'll, we'll have compilers that are able to do a lot of these optimizations for you, and you don't have to, for example, spend a couple months writing CUDA to get this stuff to work.”
Tri Dao Aug 3, 2023 ▶ 11:39 FlashAttention-2: Making Transformers 800% faster AND exact
Assertion Supported
Hotz: Tinygrad runs all ML models with only 25 primitive operations
“Tiny grad is, we are going to make a risk offset for all ML models. And yeah, it can run all ML models with basically 25 instead of the two 50 of XLA or PrimTorch. So about 10 X less complex.”
George Hotz Jun 20, 2023 ▶ 3:51 Ep 18: Petaflops to the People — with George Hotz of tinycorp
Assertion Contradicted
Swyx claims Airtable founder Howie Liu had already sold the company
“I was also mentioned, I was also thinking about Howie Lu. From Airtable. Effectively just did the same thing with Hyperagent, except that he didn't run it in parallel that much. He basically had already sold the company and was just kind of doubleheading for a…”
Shawn Wang Sep 7, 2026 ▶ 39:46 Orbs: Shifting Coding to Cloud — Quinn Slack, Amp Code
Assertion Supported
Swyx claims Jeff Dean just left Google
“Jeff Dean just left Google.”
Shawn Wang Sep 7, 2026 ▶ 38:50 Orbs: Shifting Coding to Cloud — Quinn Slack, Amp Code
Assertion Supported
Lie: Cerebras demoed GPT running at over 4,400 TPS at Hot Chips
“We here in this demo that we gave at hot chips we're showing GPT OSS running at over 4000 400 TPS, which is just mind blowing.”
Sean Lie Sep 2, 2026 ▶ 5:16 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Assertion Supported
Lie: Trillion-parameter models require thousands of Groq LPUs for weights
“To run a frontier level model, like, let's say, a few trillion parameters, you need thousands and thousands of Grok LPUs just to hold the weights, right?”
Sean Lie Sep 2, 2026 ▶ 23:17 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.