The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 1,786 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 41 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Supported
Meta Llama 3.3 and Llama 4 perform poorly on agent benchmarks
“Another, of course, the other surprise was that all the Lama models were not performing well on our benchmark. 3.3 and even the Lama four all were really performing extremely poor.”
Pratik Bhavsar Jul 14, 2025 ▶ 10:08 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
Assertion Not checkable as stated
Jack Morris: Fundamental AI science shifted to companies due to academic compute limits
“That's when I think things really started to change in terms of the types of questions you wanted to ask can't always be answered with academic resources. So a lot of the like fundamental kind of like boundary pushing and AI science moved into companies.”
Jack Morris Jul 2, 2025 ▶ 4:22 Information Theory for Language Models: Jack Morris
Assertion Not checkable as stated
Zach Lloyd: Cursor represents a very significant portion of Anthropic's revenue
“Cursor I think is some very significant portion of their revenue.”
Zach Lloyd Jun 25, 2025 ▶ 35:57 ⚡️Warp 2.0: the Agentic Development Environment - Zach Lloyd and Ben Holmes
Assertion Not checkable as stated
Brown: GPT-4o and o3 are passing the Turing test
“So at this point, like, you know, the truth is, you know, GPT-IV-O and like O-III, these models are like passing the Turing test.”
Noam Brown Jun 19, 2025 ▶ 3:28 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Assertion Not checkable as stated
Lattner says Mojo is 10,000x faster than Python and beats Rust
“Mojo is not just a little faster than Python, it's faster than Rust. So it's like tens of thousands of times faster than Python, and it's in the Python family.”
Chris Lattner Jun 13, 2025 ▶ 15:29 The Shape of Compute (Chris Lattner of Modular)
Assertion Not publicly verifiable
Modular's Mojo FlashAttention beats Tri Dao's reference implementation
“We're beating the tree DAO reference implementation that everybody uses, for example, right? Written fully in Mojo. Again, all of our, GPU kernels are written in Mojo. You can go see the history of the team building this, and it was done in just a few weeks, r…”
Chris Lattner Jun 13, 2025 ▶ 30:59 The Shape of Compute (Chris Lattner of Modular)
Assertion Supported
Duffy: OpenAI's o3 actively deceives opponents and plots betrayals in AI Diplomacy
“Oh, three was one of the few that will actually send a message to another power saying that they're planning to do something. And then like in their diary diary, right? Oh, they fell for it. Hook, line and sinker. Totally gonna betray him and take it over.”
Alex Duffy Jun 11, 2025 ▶ 12:22 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
Assertion Supported
Claude loses AI Diplomacy games because it refuses to deceive opponents
“I haven't seen Claude with any game yet because they won't do it. Like there's like, O three has managed to get them on board for like draws, even though they all know the only win condition in the game is, is 18 supply centers.”
Alex Duffy Jun 11, 2025 ▶ 12:36 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
Assertion Not checkable as stated
Kilpatrick: Gemini's SOTA video performance resulted from reasoning, not video engineering
“With reasoning is a great example of this where like multimodal with video understanding ended up like having this huge, like it's having this beautiful moment. The model is like soda out of the box because of all the reasoning capabilities that were baked in…”
Logan Kilpatrick Jun 2, 2025 ▶ 13:19 [AIEWF Preview] Gemini in 2025 and Realtime Voice AI
Assertion Supported
Cherny: Anthropic is currently bordering on AI Safety Level 3 capabilities
“Yeah, we're kind of bordering on three right now.”
Boris Cherny May 7, 2025 ▶ 34:33 Claude Code: Anthropic's CLI Agent
Assertion Not checkable as stated
99% of Current Monetizable Voice AI Use Cases Are Telephony
“99% of the monetizable voice AI use cases today are telephony.”
Kwindla Hultman Kramer May 6, 2025 ▶ 15:48 Voice AI Masterclass — Kwindla Hultman Kramer and swyx
Assertion Supported
Factorio benchmark results show reasoning models underperform expectations on extended planning
“One thing we have found in preliminary results is that the reasoning models don't seem to do as well as you'd expect in this setting. And I think that's probably because the way we set this up, it's a bit like we're already making it do reasoning traces over a…”
Jack Hopkins Apr 27, 2025 ▶ 12:12 ⚡️Factorio Learning Environment: the ultimate Game Agent Eval — Jack Hopkins
Assertion Not checkable as stated
Swix: Claude 3 degraded in capability a month after launch
“I used the same project to do this, to try to repeat the demo that I made for myself a month afterwards, and it wasn't anywhere as smart. So cloud three got dumber, but it looks like I like made up the demo or something, but no, like literally I just reran the…”
Shawn Wang Apr 24, 2025 ▶ 7:42 Why Every Agent needs Open Source Cloud Sandboxes
Assertion Not checkable as stated
Conrad: OpenRouter open-source traffic required only around 10 H100 nodes
“The entirety of Open Router that was not Anthropic or Google like, or Gemini or OpenAI or something. It was like, 10 H 100 nodes or something like that. It's just, like, not that much. It's like, not that many GPUs, actually, to service that entire demand.”
Evan Conrad Apr 11, 2025 ▶ 32:58 SF Compute: Commoditizing Compute
Assertion Not checkable as stated
Hershey: Claude 3.7 extended thinking does not help Pokémon gameplay
“I've tested, like, all sorts of the extended thinking mode with, ah, 3.7 on it, and, like, it doesn't really help.”
David Hershey Apr 5, 2025 ▶ 14:01 Claude Plays Pokémon Hackathon: Escape from Mt. Moon!
Assertion Not checkable as stated
Sequential Thinking MCP outperformed Claude 3.7 native reasoning mode in evaluations
“We tried reasoning mode as well with the new three seven. And we didn't see that much of a bump in performance. We don't know if this is something that's code specific or not. I don't have an insight. We tried both and yeah, sequential thinking worked better.”
Guy Gur-Ari Apr 2, 2025 ▶ 3:57 The #1 SWE-Bench Verified Agent
Assertion Supported
Agarwal: Filtered 9B Synthetic Data Outperforms 27B Self-Generated Data
“One thing we found consistently, so here what we had two models, nine Gemma, nine B and Gemma, 27 B, and we found consistently that actually generating data from nine B in a compute match setting is always better, even better for distilling or actually improvi…”
Rishabh Agarwal Mar 23, 2025 ▶ 17:41 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Assertion Supported
Cursor Composer solves Convex benchmarks but fails on alternative backends
“We did notice that I mean, with convex, it pretty much autonomously just solves the first two tasks. It has a few round trips on like some errors that are only show up and playing with the front end. And then it's able to complete this files task and kind of g…”
Sujay Jayakar Mar 19, 2025 ▶ 9:36 Fullstack-Bench: The Eval for Coding Agents — with Sujay Jayakar, Chief Scientist, Convex
Assertion Not checkable as stated
Claude 3.7 performed worse than Claude 3.5 on Convex evals
“For example, we just tried clod three seven and it performs worse than clod three five on convex evals with the same prompting.”
Sujay Jayakar Mar 19, 2025 ▶ 17:23 Fullstack-Bench: The Eval for Coding Agents — with Sujay Jayakar, Chief Scientist, Convex
Assertion Supported
AI models struggle debugging Supabase RLS recursion compared to procedural code
“The particular example was like RLS rules and Supabase where debugging like an infinite loop for infinite recursion for the RLS rules was something that the models just really struggled with in a way that we didn't see for procedural code.”
Sujay Jayakar Mar 19, 2025 ▶ 29:57 Fullstack-Bench: The Eval for Coding Agents — with Sujay Jayakar, Chief Scientist, Convex
Assertion Not checkable as stated
Klein: Browserbase delivers 90% of computer-use functionality at 10% of OS cost
“BrowserBase can run browsers for way cheaper than you can if you're running a full-fledged OS with a GUI, you know, operating system. And I think that's just an advantage of the browser. It is like, browsers are like little OSs, and you can run them very effic…”
Paul Klein Feb 28, 2025 ▶ 48:53 Browserbase: Browser Infrastructure For Your AI Agents
Assertion Not checkable as stated
Agarwal: 90% of production LLM use cases do not use automatic routing
“In fact, I would say in production, 90% of the use cases do not use automatic routing. What they want is deterministic flows. As long as the gateway manages authentication authorization for them, it's perfectly fine. The request hitting A specific model that t…”
Rohit Agarwal Feb 5, 2025 ▶ 1:53 Why every AI Engineer needs an AI Gateway (ft Portkey.ai CEO)
Assertion Not checkable as stated
Nguyen: Stanford HELM benchmark under-reported Claude performance due to improper prompting
“This has happened with, like, Stanford, I remember, like, when Stanford had lists also, like, they were, like, running benchmarks. Yeah, Helm. And somehow, like, Claude was, like, always, like, not performing well, and that's because, like, the way they prompt…”
Karina Nguyen Feb 1, 2025 ▶ 16:39 The Agent Reasoning Interface: Claude, ChatGPT Canvas, Tasks, Operator — with Karina Nguyen, OpenAI
Assertion Supported
DeepSeek-R1 researchers found MCTS and Process Reward Models were not useful
“R-one specifically said, yes, we tried MCTS. Yes, we tried PRMs. And none of that is useful.”
Shawn Wang Jan 24, 2025 ▶ 10:32 The Unreasonable Effectiveness of Reasoning Distillation: using DeepSeek R1 to beat OpenAI o1
Assertion Supported
Zhang: XGrammar outperforms Outlines and is integrated into TensorRT-LLM
“And I think Xgrammar's performance is better than the outline's, and also in the TensorFlow RTLM, the latest release, TensorFlow RTLM also integrates Xgrammar as the backend for the constructed coding.”
Yining Zhang Jan 19, 2025 ▶ 41:22 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Assertion Not checkable as stated
Automatic1111's inefficient SDXL implementation drove ComfyUI's viral user adoption
“The big, one point zero release happened, and wow, Confu UI was the only way a lot of people could actually run it on their computers, because it just, like, automatic was so, like, inefficient and bad that most people couldn't act, like, it just wouldn't work…”
comfyanonymous (Comfy) Jan 4, 2025 ▶ 47:23 AI Engineering for Art - with comfyanonymous
Assertion Not checkable as stated
Swyx: VC appetite for GPU-rich early-stage startups is completely gone
“The appetite for GPU rich startups, like the, you know, the funding plan is we will raise sixty million and we'll give 50 of that to Nvidia. That is gone, right? Like no one's pitching that. This was literally the plan, the exact plan of like, I can name like …”
Shawn Wang Jan 1, 2025 ▶ 38:52 2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
Assertion Supported
Neubig: SWE-bench Scores Are Inflated By Training Data Contamination
“Sweebench is on popular open source repos and all of these popular open source repos were included in the training data for all of the language models. And so, the language models already know these repos. In some cases, the language models already know the in…”
Graham Neubig Dec 25, 2024 ▶ 32:49 Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
Assertion Supported
Ben Allal: Hugging Face SmolLM2-1.7B outperforms Llama 3.2 models
“So it's a series of three models, which are the best in class in each model size. For example, our 1.7 B model outperforms Lama one B and also .2.”
Loubna Ben Allal Dec 24, 2024 ▶ 22:31 Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
Assertion Supported
Reddy: Chai Discovery's open-source Chai-1 model outperforms Google's AlphaFold 3
“We're lucky to work with the folks at Chai Discovery who just released Chai One, which is open source model that outperforms Alpha Fold Three.”
Pranav Reddy Dec 21, 2024 ▶ 7:44 The State of AI Startups in 2024 [LS Live @ NeurIPS]
Assertion Not checkable as stated
Mohan: Codeium generated dynamic PNGs due to VS Code API limitations
“On VS Code, actually, the problem for us wasn't actually being able to implement the feature. We had the feature for a while. Problem was actually even to show the feature, VS Code would not expose an API for us to do this. So what we actually ended up doing w…”
Varun Mohan Dec 13, 2024 ▶ 8:12 Windsurf: The Enterprise AI IDE
Assertion Supported
Cerebras WSE-3 runs Llama inference 70x faster than NVIDIA GPUs
“Cerebris came out that the wafer scale engine three can serve llama 70 B at 2.1 thousand sorry, 202,100 tokens per second and serves llama four or five B at nearly 1000 tokens per second. So this, you know, to give you an understanding, like this is about 70 t…”
Sarah Chieng Dec 7, 2024 ▶ 3:06 [Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras)
Assertion Contradicted
Friedman: GitHub Copilot user retention in enterprise is 38% to 50%
“Between 38 to 50% Retention for users using Copilot and Enterprise.”
Itamar Friedman Dec 2, 2024 ▶ 37:04 0 to over $8M ARR in 2 months as a Claude Wrapper (Bolt.new, Qodo)
Assertion Supported
Friedman: AlphaCodium boosts OpenAI o1, proving o1 lacks true System 2
“We took their all one preview with Alpha Codium and did better. Like it just shows like, and there is a big difference between the preview and the IOI. It shows, like, that these models are not still system two thinkers, and there's a big difference.”
Itamar Friedman Dec 2, 2024 ▶ 42:48 0 to over $8M ARR in 2 months as a Claude Wrapper (Bolt.new, Qodo)
Assertion Not checkable as stated
Crivello: Marc Andreessen Blocked Him on Twitter Over AI Safety Views
“Like at some point, Marc Andreessen blocked me on Twitter and I, it hurt, frankly, I really look up to Marc Andreessen and I knew he would block me.”
Florent Crivello Nov 15, 2024 ▶ 1:01:05 Agents @ Work: Lindy.ai (with live demo!)
Assertion Supported
SWE-Bench public test splits enable trivial cheating via runtime PR retrieval
“The entire test split here is public. So you can do things like just overfit to the patches in the test set. You can do things like, let me add at runtime, pull the PR and just get the answer and just use it.”
Jesse Hu Oct 19, 2024 ▶ 15:00 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
Assertion Supported
Houston: Groq and Cerebras outperform Nvidia on latency
“There's also, like, non-NVIDIA stacks, like the Grok, or Cerebris, or some of these custom silicon companies that are super interesting, and all, and outperformed the NVIDIA stack in terms of latency and things like that.”
Drew Houston Oct 18, 2024 ▶ 52:59 Building the Silicon Brain - Drew Houston of Dropbox
Assertion Supported
Molmo Outperforms Gemini 1.5 and Claude 3.5 Sonnet With 1M Samples
“They can get better than Gemini, 1.5, better than Claude, 3.5 sonnet, better than GPT for V at a much smaller size with about a million samples of data, which is very impressive, right?”
Vibhu Sapra Oct 13, 2024 ▶ 4:23 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Assertion Not checkable as stated
Goyal: Simple tool-calling prompts cover 80% to 90% of AI use cases
“For probably 80 or 90% of the use cases that we see with people doing this, like very, very simple, I create a prompt, it calls some tools. I can like very ergonomically write the tools, plug into popular services, et cetera, and then just call them kind of li…”
Ankur Goyal Oct 11, 2024 ▶ 50:55 Production AI Engineering starts with Evals
Assertion Not checkable as stated
Goyal: Fewer Braintrust customers run fine-tuned models in production than six months ago
“I will say in my own experience with customers as of the recording date today, which is September or something, yeah, very few of our customers are currently fine-tuning models. And I think a very, very small fraction of them are running fine-tuned models in p…”
Ankur Goyal Oct 11, 2024 ▶ 1:21:53 Production AI Engineering starts with Evals
Assertion Not checkable as stated
Goyal: OpenAI dominates production while Anthropic Sonnet leads side projects
“We still see an overwhelming majority of customers using OpenAI, but almost everyone is using Anthropic for the, and Sonnet specifically for their side projects, whether it's You know, via cursor or prototypes or whatever.”
Ankur Goyal Oct 11, 2024 ▶ 1:26:58 Production AI Engineering starts with Evals
Assertion Not checkable as stated
Goyal: Public clouds fail to match direct OpenAI endpoint experience and capacity
“It hasn't been a smooth journey for people to get the capacity on public clouds that they're able to get through, you know, OpenAI directly. I mean, I think a lot of this is changing, catching up, et cetera. But it hasn't been perfectly smooth. And I think the…”
Ankur Goyal Oct 11, 2024 ▶ 1:28:43 Production AI Engineering starts with Evals
Assertion Not checkable as stated
Goyal: Nearly all Braintrust clients shifted to simple code with LLM calls
“Almost everyone that we work with has gone into this model that, that I, that actually exactly what you said, which is sprinkle intelligence everywhere and make it easy to write dumb code”
Ankur Goyal Oct 11, 2024 ▶ 1:42:27 Production AI Engineering starts with Evals
Assertion Supported
Pullen: CoScene's Genie outperforms OpenAI o1 out of the box on SWE-bench
“So it was obviously great to see, like, we still are better than O-one out of the box. You know, even with an older model, and I'm sure that that, that Delta will continue to grow once we're able to train O-one and once we've done more work on our dataset usin…”
Alistair Pullen Oct 4, 2024 ▶ 1:13:04 Building AGI in Real Time (OpenAI Dev Day 2024)
Assertion Not checkable as stated
Altman: OpenAI reached Level 2 AGI with o1
“I think we clearly got to level two, or we clearly got to level two with O-one.”
Sam Altman Oct 4, 2024 ▶ 1:24:50 Building AGI in Real Time (OpenAI Dev Day 2024)
Assertion Not checkable as stated
Altman: o1 is OpenAI's most aligned model ever by a lot
“And O-one is obviously our most capable model ever, but it's also our most aligned model ever by a lot.”
Sam Altman Oct 4, 2024 ▶ 1:33:35 Building AGI in Real Time (OpenAI Dev Day 2024)
Assertion Not checkable as stated
Weil: OpenAI customer support team is 20% expected size thanks to AI
“There are things that get closer to that, I mean, there, like, customer service, we have bots internally that do what's fun about answering external questions and fielding internal people's questions on Slack and so on, and our customer success, our customer s…”
Kevin Weil Oct 4, 2024 ▶ 1:57:42 Building AGI in Real Time (OpenAI Dev Day 2024)
Assertion Not checkable as stated
Shunyu Yao says Ilya Sutskever claimed GPT-1 had solved language
“Back in OpenAI, they did this GPT-ONE together, and Ilya just said, Karthik, you should stay, because we just solved the language.”
Shunyu Yao Sep 27, 2024 ▶ 2:12 Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.