The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 851 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Supported
Rumbelow: Leap Labs' Discovery Engine Automates Novel Scientific Discovery via Interpretability
“Discovery Engine is an end-to-end system, takes in arbitrary scientific data set, automatically trains a bunch of neural networks on it, and then We systematically, with our interpretability methods, which is the real secret extract the patterns that have been…”
Jessica Rumbelow Nov 2, 2025 ▶ 6:16 ⚡️Automating Scientific Discovery - Jessica Rumbelow, Leap Labs
Assertion Supported
Corbitt: OpenPipe Beat Frontier Models Using a Qwen 32B Judge
“One of the results we published was we used Quen 2.5 14 B as the model we're training, and as the judge we used Quen 2.5 32 B, which is, like, Not, I mean, it's fine, but it's like not a, it's much worse than any frontier model. Right. And even with that combi…”
Kyle Corbitt Oct 16, 2025 ▶ 53:18 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
Assertion Supported
Feldman: Cerebras provides 2,625x more memory bandwidth than traditional GPUs
“And we have 2625 times more memory bandwidth than the GPU does.”
Andrew Feldman Oct 1, 2025 ▶ 6:22 ⚡️Raising $1.1b to build the fastest LLM Chips on Earth — Andrew Feldman, Cerebras
Assertion Supported
Cerebras leads all Artificial Analysis inference benchmarks by a large margin
“I think also just go up and look at artificial analysis. Wherever we are, we're the fastest not by a little bit, but by a lot.”
Andrew Feldman Oct 1, 2025 ▶ 11:03 ⚡️Raising $1.1b to build the fastest LLM Chips on Earth — Andrew Feldman, Cerebras
Assertion Supported
Rajpal: Anthropic Claude models had regressions from serving architecture changes
“Anthropix kind of cloud models kind of had a regression, right? Because they changed to a new serving architecture.”
Shreya Rajpal Sep 25, 2025 ▶ 23:21 ⚡️Snowglobe: Simulations for your AI
Prediction Held up
Taskaya: Training a state-of-the-art image model costs under $1M
“Like right now, like if you look, if you want to train a Sota image model, I don't think it's going to cost more than a million dollars. It's extremely cheap. It's like a matter of data engineering effort, cleaning. It's, I think it's a function of data set.”
Batuhan Taskaya Sep 8, 2025 ▶ 56:02 A Technical History of Generative Media
Assertion Supported
Morcos: Soft inductive biases become harmful past 1M data points in vision
“Turns out in the small data regime, and when I say small data here, I mean, say less than 500,000 data points. And this was in the context of image self-supervised learning. So in that small data regime, this is super helpful. And where this paper's actually b…”
Ari Morcos Aug 29, 2025 ▶ 6:04 Better Data is All You Need — Ari Morcos, Datology
Assertion Supported
Morcos: Kaplan and Chinchilla scaling laws incorrectly assume all data is equal
“And even if you go and you look at the scaling laws work from Kaplan and Chinchilla and all these other things, they all assume IID data which is insane. We know that all data are not created equal, that garbage in garbage out is like the oldest adage in compu…”
Ari Morcos Aug 29, 2025 ▶ 8:25 Better Data is All You Need — Ari Morcos, Datology
Assertion Supported
Morcos: Proper data curation can bend neural scaling laws
“And what that paper showed was that if you use your data correctly, you can actually bend the scaling laws themselves.”
Ari Morcos Aug 29, 2025 ▶ 27:35 Better Data is All You Need — Ari Morcos, Datology
Prediction Held up
Morcos: Training a specialized frontier model will cost under $1M very soon
“I believe that getting to a frontier model should cost a million dollars or less for most organizations, at least in a specialized domain, right? And when you think about what enterprises need, that's generally what they need. They don't need a model that can …”
Ari Morcos Aug 29, 2025 ▶ 42:49 Better Data is All You Need — Ari Morcos, Datology
Assertion Supported
Sohmers: Positron AI requires zero compilers to run Hugging Face models
“So rather than having like, we don't have a compiler whatsoever. There's no compiler. There's no translator, no tooling that's involved in actually taking those and getting that to, you know, for your common, you know, Huggy Face Transform models to be able to…”
Thomas Sohmers Aug 18, 2025 ▶ 21:32 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Assertion Supported
Palazzolo: Claude Code leads stayed at Cursor only two weeks
“We know that they went there, they were there for, I think, about two weeks, and they came back.”
Stephanie Palazzolo Aug 6, 2025 ▶ 41:45 The AI Agenda: GPT5 leaks and the business of AI News — Steph Palazzolo, The Information
Assertion Supported
Ermon: Diffusion LLMs Pareto-dominate autoregressive models on inference efficiency
“On the inference side, what we're seeing is that diffusion models are much more efficient. We're actually able to Pareto dominate autoregressive models. If you think about the typical trade-off between throughput versus latency, which you kind of like cannot, …”
Stefano Ermon Aug 4, 2025 ▶ 14:30 ⚡️Mercury: Ultra-Fast Diffusion LLMs — Estefano Ermon, CEO Inception Labs
Assertion Supported
Lambert: Tulu 3 matches or beats Meta Llama 3.1 on core evals
“On, like, core evals for our Suite of models from, I think, eight, seven D and four or five B is based on llama at the time. It's like it matches or beats meta on these core valves.”
Nathan Lambert Jul 31, 2025 ▶ 2:20 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Assertion Supported
Fortuna: New reasoning models show no big leap on medical coding tasks
“I do know that when you kind of plot out base model performance on some medical tasks like ICD-X coding between like, you know, previous generations and new reasoning generations, there's actually not like a big leap.”
Brendan Fortuna Jul 29, 2025 ▶ 24:01 ⚡️Using RFT to Build Clinical Superintelligence
Assertion Supported
OpenAI's IMO performance was not officially verified by the IMO
“It turns out, like, OpenAI actually didn't involve officially with IMO. They just, like, use the problems, but, and then just, like, use their model to test the results, and ask, like, three previous IMO analysts to review them.”
Dr. Jasper Zhang Jul 24, 2025 ▶ 3:28 ⚡️Math Olympiad gold medalist explains OpenAI and Google DeepMind IMO Gold Performances
Assertion Supported
McCloy: ChatGPT Does Not Index or Retrieve llms.txt by Default
“I knew that there's debate about this, but I'd say the evidence is like ChatTriPT is not indexing and it's not retrieving content from LMS.txt by default.”
Robert McCloy Jul 23, 2025 ▶ 38:52 AI is Eating Search
Assertion Supported
OpenAI o-series reasoning models fail at multi-tool calling benchmarks
“Then another surprise for me was that the reasoning models were not performing well enough. They had certain kind of limitation when we probed into it, like, why are they scoring less overall? They were like the O-one, the O-four, O-three, they, When not perfo…”
Pratik Bhavsar Jul 14, 2025 ▶ 8:22 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
Assertion Supported
Meta Llama 3.3 and Llama 4 perform poorly on agent benchmarks
“Another, of course, the other surprise was that all the Lama models were not performing well on our benchmark. 3.3 and even the Lama four all were really performing extremely poor.”
Pratik Bhavsar Jul 14, 2025 ▶ 10:08 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
Assertion Supported
Duffy: OpenAI's o3 actively deceives opponents and plots betrayals in AI Diplomacy
“Oh, three was one of the few that will actually send a message to another power saying that they're planning to do something. And then like in their diary diary, right? Oh, they fell for it. Hook, line and sinker. Totally gonna betray him and take it over.”
Alex Duffy Jun 11, 2025 ▶ 12:22 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
Assertion Supported
Claude loses AI Diplomacy games because it refuses to deceive opponents
“I haven't seen Claude with any game yet because they won't do it. Like there's like, O three has managed to get them on board for like draws, even though they all know the only win condition in the game is, is 18 supply centers.”
Alex Duffy Jun 11, 2025 ▶ 12:36 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
Prediction Held up
Ameisen: Deceptive Backward Reasoning Exists in Base Pre-Trained Models
“I bet, I don't know how much I bet a hundred bucks. So somebody can like, they would get a hundred bucks from me if they prove that I'm wrong, that this behavior for a model that does a drink fine tuning, it also does it post pre-training.”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:33:39 The Utility of Interpretability — Emmanuel Amiesen
Assertion Supported
Cherny: Anthropic is currently bordering on AI Safety Level 3 capabilities
“Yeah, we're kind of bordering on three right now.”
Boris Cherny May 7, 2025 ▶ 34:33 Claude Code: Anthropic's CLI Agent
Assertion Supported
Factorio benchmark results show reasoning models underperform expectations on extended planning
“One thing we have found in preliminary results is that the reasoning models don't seem to do as well as you'd expect in this setting. And I think that's probably because the way we set this up, it's a bit like we're already making it do reasoning traces over a…”
Jack Hopkins Apr 27, 2025 ▶ 12:12 ⚡️Factorio Learning Environment: the ultimate Game Agent Eval — Jack Hopkins
Assertion Supported
Agarwal: Filtered 9B Synthetic Data Outperforms 27B Self-Generated Data
“One thing we found consistently, so here what we had two models, nine Gemma, nine B and Gemma, 27 B, and we found consistently that actually generating data from nine B in a compute match setting is always better, even better for distilling or actually improvi…”
Rishabh Agarwal Mar 23, 2025 ▶ 17:41 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Assertion Supported
Cursor Composer solves Convex benchmarks but fails on alternative backends
“We did notice that I mean, with convex, it pretty much autonomously just solves the first two tasks. It has a few round trips on like some errors that are only show up and playing with the front end. And then it's able to complete this files task and kind of g…”
Sujay Jayakar Mar 19, 2025 ▶ 9:36 Fullstack-Bench: The Eval for Coding Agents — with Sujay Jayakar, Chief Scientist, Convex
Assertion Supported
AI models struggle debugging Supabase RLS recursion compared to procedural code
“The particular example was like RLS rules and Supabase where debugging like an infinite loop for infinite recursion for the RLS rules was something that the models just really struggled with in a way that we didn't see for procedural code.”
Sujay Jayakar Mar 19, 2025 ▶ 29:57 Fullstack-Bench: The Eval for Coding Agents — with Sujay Jayakar, Chief Scientist, Convex
Prediction Held up
Roucher: AI agents will reach a 90% GAIA score by 2026
“So I think if we solve Gaia, that's like 90% score. That means mostly we double productivity of every task done in front of a computer. And if you take the trend line of the scores so far this should be crossed in 2026 or something.”
Aymeric (Emmerich) Feb 13, 2025 ▶ 16:12 smol agents are all you need
Assertion Supported
DeepSeek-R1 researchers found MCTS and Process Reward Models were not useful
“R-one specifically said, yes, we tried MCTS. Yes, we tried PRMs. And none of that is useful.”
Shawn Wang Jan 24, 2025 ▶ 10:32 The Unreasonable Effectiveness of Reasoning Distillation: using DeepSeek R1 to beat OpenAI o1
Assertion Supported
Zhang: XGrammar outperforms Outlines and is integrated into TensorRT-LLM
“And I think Xgrammar's performance is better than the outline's, and also in the TensorFlow RTLM, the latest release, TensorFlow RTLM also integrates Xgrammar as the backend for the constructed coding.”
Yining Zhang Jan 19, 2025 ▶ 41:22 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Assertion Supported
Neubig: SWE-bench Scores Are Inflated By Training Data Contamination
“Sweebench is on popular open source repos and all of these popular open source repos were included in the training data for all of the language models. And so, the language models already know these repos. In some cases, the language models already know the in…”
Graham Neubig Dec 25, 2024 ▶ 32:49 Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
Assertion Supported
Ben Allal: Hugging Face SmolLM2-1.7B outperforms Llama 3.2 models
“So it's a series of three models, which are the best in class in each model size. For example, our 1.7 B model outperforms Lama one B and also .2.”
Loubna Ben Allal Dec 24, 2024 ▶ 22:31 Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
Assertion Supported
Reddy: Chai Discovery's open-source Chai-1 model outperforms Google's AlphaFold 3
“We're lucky to work with the folks at Chai Discovery who just released Chai One, which is open source model that outperforms Alpha Fold Three.”
Pranav Reddy Dec 21, 2024 ▶ 7:44 The State of AI Startups in 2024 [LS Live @ NeurIPS]
Assertion Supported
Cerebras WSE-3 runs Llama inference 70x faster than NVIDIA GPUs
“Cerebris came out that the wafer scale engine three can serve llama 70 B at 2.1 thousand sorry, 202,100 tokens per second and serves llama four or five B at nearly 1000 tokens per second. So this, you know, to give you an understanding, like this is about 70 t…”
Sarah Chieng Dec 7, 2024 ▶ 3:06 [Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras)
Assertion Supported
Friedman: AlphaCodium boosts OpenAI o1, proving o1 lacks true System 2
“We took their all one preview with Alpha Codium and did better. Like it just shows like, and there is a big difference between the preview and the IOI. It shows, like, that these models are not still system two thinkers, and there's a big difference.”
Itamar Friedman Dec 2, 2024 ▶ 42:48 0 to over $8M ARR in 2 months as a Claude Wrapper (Bolt.new, Qodo)
Assertion Supported
SWE-Bench public test splits enable trivial cheating via runtime PR retrieval
“The entire test split here is public. So you can do things like just overfit to the patches in the test set. You can do things like, let me add at runtime, pull the PR and just get the answer and just use it.”
Jesse Hu Oct 19, 2024 ▶ 15:00 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
Assertion Supported
Houston: Groq and Cerebras outperform Nvidia on latency
“There's also, like, non-NVIDIA stacks, like the Grok, or Cerebris, or some of these custom silicon companies that are super interesting, and all, and outperformed the NVIDIA stack in terms of latency and things like that.”
Drew Houston Oct 18, 2024 ▶ 52:59 Building the Silicon Brain - Drew Houston of Dropbox
Assertion Supported
Molmo Outperforms Gemini 1.5 and Claude 3.5 Sonnet With 1M Samples
“They can get better than Gemini, 1.5, better than Claude, 3.5 sonnet, better than GPT for V at a much smaller size with about a million samples of data, which is very impressive, right?”
Vibhu Sapra Oct 13, 2024 ▶ 4:23 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Assertion Supported
Pullen: CoScene's Genie outperforms OpenAI o1 out of the box on SWE-bench
“So it was obviously great to see, like, we still are better than O-one out of the box. You know, even with an older model, and I'm sure that that, that Delta will continue to grow once we're able to train O-one and once we've done more work on our dataset usin…”
Alistair Pullen Oct 4, 2024 ▶ 1:13:04 Building AGI in Real Time (OpenAI Dev Day 2024)
Prediction Held up
Altman: 10-million-token fast context windows are coming within months
“Even getting to the, like, Ten million tokens of very fast and accurate context, which I expect to measure in, like, months, something like that.”
Sam Altman Oct 4, 2024 ▶ 2:05:30 Building AGI in Real Time (OpenAI Dev Day 2024)
Assertion Supported
Schulhoff: LLMs Rely More on Prompt Structure Than Exemplar Labels
“There are a number of papers which have found that the label of the exemplar doesn't really matter, and the model reads the exemplars and cares more about structure than label.”
Sander Schulhoff Sep 20, 2024 ▶ 26:41 The Ultimate Guide to Prompting - with Sander Schulhoff from LearnPrompting.org
Assertion Supported
Jamil: Writing in the Margins solves the lost-in-the-middle problem
“It improves the ability of any language model to extract relevant information, so solving the lost in the middle problem”
Umar Jamil Sep 19, 2024 ▶ 22:54 [Paper Club] Writing in the Margins: Chunked Prefill KV Caching for Long Context Retrieval
Assertion Supported
Scialom: Llama 3 405B is the best open-source model ever released
“At a high level, it's the best open source model ever. It's Better than GPT-IV. I mean, what version? But, by far, compared to the version originally released even now, I think there's maybe the last cloud Sonya FF-V and GPT-IV-Zero that are performing it.”
Thomas Scialom Jul 23, 2024 ▶ 37:43 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Assertion Supported
Tay: Zero-shot benchmark scores at 1B model scale are random chance
“Every time some people propose like this, they run like some zero-shot score on like some LM event harness or something like that, and you know like at one B scale, all the numbers are random, basically. Like all your bull kill, they're all like random chance …”
Yi Tay Jul 5, 2024 ▶ 1:46:43 The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
Assertion Supported
Albrecht: Benchmark performance differences vanish once ambiguous questions are cleaned
“The main takeaway from any of the, like, actual performance is like, once you fix these ambiguous examples, a lot of these benchmarks are really saturated. Like, I think it's important to look at like, you know, like when you're talking about performance on NL…”
Josh Albrecht Jun 25, 2024 ▶ 1:01:43 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Assertion Supported
Albrecht: AgentBench paper's appendix examples are actually incorrect solutions
“Like we were looking at the agent bench paper, I think just last week for our paper club. And one of the things that we noticed is that actually like both of the examples in the appendix that are given as like traces where it got it right. This is actually not…”
Josh Albrecht Jun 25, 2024 ▶ 1:06:05 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Prediction Held up
Bach: Smaller, more powerful models will ensure unconstrained AI remains accessible
“Yes, but there will also be better jailbroken models or models that have never been jailed before, because we find out how to make smaller models that are more powerful.”
Joscha Bach Apr 27, 2024 ▶ 1:28:21 This World Does Not Exist — Joscha Bach, Karan Malhotra, Rob Haisfield (WorldSim, WebSim, Liquid AI)
Assertion Supported
Murphy: Five-minute voice calls cost 6.5 cents on Deepgram versus ElevenLabs.
“And then on the text-to-speech side, and doing something like this with an 11 labs would be about maybe a dollar 20. And just to give you an idea of comparison. So you can do a five minute call here for about six and a half cents.”
Damien Murphy Apr 6, 2024 ▶ 16:52 Personal AI Meetup - Bee, BasedHardware, LangChain LangFriend, Deepgram EmilyAI
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.