The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 2,445 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 100 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Supported
Agarwal: Filtered 9B Synthetic Data Outperforms 27B Self-Generated Data
“One thing we found consistently, so here what we had two models, nine Gemma, nine B and Gemma, 27 B, and we found consistently that actually generating data from nine B in a compute match setting is always better, even better for distilling or actually improvi…”
Rishabh Agarwal Mar 23, 2025 ▶ 17:41 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Prediction Not checkable as stated
Agarwal: Logit Distillation Can Match Giant Teacher Models on Reasoning
“My hunch is that the logic-based distillation can go even further, and you might be able to even close the gap with the biggest of the teachers you have, because I don't think you need a huge number of parameters, because the reasoning process is very, very, l…”
Rishabh Agarwal Mar 23, 2025 ▶ 42:39 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Assertion Supported
Cursor Composer solves Convex benchmarks but fails on alternative backends
“We did notice that I mean, with convex, it pretty much autonomously just solves the first two tasks. It has a few round trips on like some errors that are only show up and playing with the front end. And then it's able to complete this files task and kind of g…”
Sujay Jayakar Mar 19, 2025 ▶ 9:36 Fullstack-Bench: The Eval for Coding Agents — with Sujay Jayakar, Chief Scientist, Convex
Assertion Not checkable as stated
Claude 3.7 performed worse than Claude 3.5 on Convex evals
“For example, we just tried clod three seven and it performs worse than clod three five on convex evals with the same prompting.”
Sujay Jayakar Mar 19, 2025 ▶ 17:23 Fullstack-Bench: The Eval for Coding Agents — with Sujay Jayakar, Chief Scientist, Convex
Assertion Supported
AI models struggle debugging Supabase RLS recursion compared to procedural code
“The particular example was like RLS rules and Supabase where debugging like an infinite loop for infinite recursion for the RLS rules was something that the models just really struggled with in a way that we didn't see for procedural code.”
Sujay Jayakar Mar 19, 2025 ▶ 29:57 Fullstack-Bench: The Eval for Coding Agents — with Sujay Jayakar, Chief Scientist, Convex
Prediction Not checkable as stated
Kozlov: Wave of 'agent-first' businesses will emerge without traditional UIs
“I think that we're going to see a wave of, you know, going back to, again, Sunil, your point about mom and pop shops, these businesses turn up that are agent first, and that's kind of the only interface to using them as opposed to, you know, UIs and APIs and o…”
Rita Kozlov Mar 19, 2025 ▶ 22:52 npm install Agents — with Sunil Pai and Rita Kozlov (VP AI) of Cloudflare
Prediction Not checkable as stated
Solving Autonomous Coding Is the Direct Path to AGI
“Our core belief is that if you solve this problem, you solve the autonomous coding problem and build a super intelligent coding agent, that that thing will lead to super intelligence more broadly.”
Misha Laskin Mar 7, 2025 ▶ 3:50 Solve coding, solve AGI [Reflection.ai launch w/ CEO Misha Laskin]
Prediction Not checkable as stated
Superintelligence Cannot Be Trained Entirely From Scratch
“In the era of language models, I don't think you'll be able to train superintelligence from scratch.”
Misha Laskin Mar 7, 2025 ▶ 5:54 Solve coding, solve AGI [Reflection.ai launch w/ CEO Misha Laskin]
Prediction Not checkable as stated
Future UIs Will Be Built as Programmatic Interfaces for AI Models
“Over the coming years, there'll be more kind of AI friendly or language model friendly UIs. And what's friendly to a language model is, is code. So the way a model will be doing work, not just for coding and software engineering, Is by basically making functio…”
Misha Laskin Mar 7, 2025 ▶ 9:37 Solve coding, solve AGI [Reflection.ai launch w/ CEO Misha Laskin]
Prediction Not checkable as stated
Frontier Labs May Hoard Superintelligent Models and Release Nerfed Versions
“You can imagine you know, the world converging on a few companies have really powerful coding models. They basically release a nerfed version of that to the public at large, and basically have a competitive advantage by having, you know, a super intelligent co…”
Misha Laskin Mar 7, 2025 ▶ 16:17 Solve coding, solve AGI [Reflection.ai launch w/ CEO Misha Laskin]
Prediction Not checkable as stated
Klein: AI agents will use existing web interfaces instead of rebuilding APIs
“The thing that's hard is like reinventing the internet for agents. We don't want to rebuild the internet. That's an impossible task. And I think people often say like, well, we'll have this second layer of APIs built for agents. I'm like, we will for the top u…”
Paul Klein Feb 28, 2025 ▶ 27:31 Browserbase: Browser Infrastructure For Your AI Agents
Assertion Not checkable as stated
Klein: Browserbase delivers 90% of computer-use functionality at 10% of OS cost
“BrowserBase can run browsers for way cheaper than you can if you're running a full-fledged OS with a GUI, you know, operating system. And I think that's just an advantage of the browser. It is like, browsers are like little OSs, and you can run them very effic…”
Paul Klein Feb 28, 2025 ▶ 48:53 Browserbase: Browser Infrastructure For Your AI Agents
Prediction Open · timeframe Feb 2030
Klein: Browserbase will be a billion-dollar company within five years
“I can predict that Browserbase will be a billion dollar company one day. So let's check back in five years.”
Paul Klein Feb 28, 2025 ▶ 52:29 Browserbase: Browser Infrastructure For Your AI Agents
Prediction Not checkable as stated
Next billions of AI users will onboard via SMS and phone
“I don't think those people are going to come in through, you know, some front end website somewhere. Like those people are going to come in through audio from a telephone to texting to email. Like that's just the most obvious outcome.”
Logan Kilpatrick Feb 28, 2025 ▶ 20:15 Gemini 2.0 Flash and Flash Thinking: the new SOTA models for the agentic era
Prediction Not checkable as stated
Sutin: Traditional speech-to-text ASR will soon be obsolete and uninvestable
“Cause it's very clear that like all ASR, all speech to text is going to be pretty obsolete pretty soon. So like investing into that is probably kind of a dead end cause it's just going to be obsolete.”
Ethan Sutin Feb 17, 2025 ▶ 56:29 Bee AI: The Wearable Ambient Agent
Prediction Held up
Roucher: AI agents will reach a 90% GAIA score by 2026
“So I think if we solve Gaia, that's like 90% score. That means mostly we double productivity of every task done in front of a computer. And if you take the trend line of the scores so far this should be crossed in 2026 or something.”
Aymeric (Emmerich) Feb 13, 2025 ▶ 16:12 smol agents are all you need
Prediction Not checkable as stated
Bret Taylor: Open source will broadly win in AI developer tooling
“Nowadays the tools are changing so rapidly that I'm like not totally skeptical of tool makers, but I just think that open source will broadly win.”
Bret Taylor Feb 11, 2025 ▶ 25:36 The AI Architect: Bret Taylor
Prediction Not checkable as stated
Colvin: Gen AI observability will merge into general-purpose observability platforms
“Web observability stopped being a thing, not because the web stopped being a thing, but because all observability had to do web. If you were talking to people in 2010 or 2012, they would have talked about cloud observability. Now that's not a term because all …”
Samuel Colvin Feb 6, 2025 ▶ 49:00 Agent Engineering with Pydantic + Graphs — with Samuel Colvin, CEO of Pydantic Logfire
Assertion Not checkable as stated
Agarwal: 90% of production LLM use cases do not use automatic routing
“In fact, I would say in production, 90% of the use cases do not use automatic routing. What they want is deterministic flows. As long as the gateway manages authentication authorization for them, it's perfectly fine. The request hitting A specific model that t…”
Rohit Agarwal Feb 5, 2025 ▶ 1:53 Why every AI Engineer needs an AI Gateway (ft Portkey.ai CEO)
Assertion Not checkable as stated
Nguyen: Stanford HELM benchmark under-reported Claude performance due to improper prompting
“This has happened with, like, Stanford, I remember, like, when Stanford had lists also, like, they were, like, running benchmarks. Yeah, Helm. And somehow, like, Claude was, like, always, like, not performing well, and that's because, like, the way they prompt…”
Karina Nguyen Feb 1, 2025 ▶ 16:39 The Agent Reasoning Interface: Claude, ChatGPT Canvas, Tasks, Operator — with Karina Nguyen, OpenAI
Prediction Not checkable as stated
Nguyen: Website clicks will drop as internet access shifts to AI models
“In my opinion, like, people in, like, few years will click On, like, websites way less. I want to see the plot of, like, website clicks over time, but then my prediction is, like, it will go down and, like, people's access to the internet will be through the m…”
Karina Nguyen Feb 1, 2025 ▶ 56:35 The Agent Reasoning Interface: Claude, ChatGPT Canvas, Tasks, Operator — with Karina Nguyen, OpenAI
Prediction Not checkable as stated
Autonomous AI programmers will work effectively within the next two years
“I think that we will have autonomous AI programmers working really well for us, you know, within the next year or two.”
Shawn Lewis Jan 28, 2025 ▶ 33:01 Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
Assertion Supported
DeepSeek-R1 researchers found MCTS and Process Reward Models were not useful
“R-one specifically said, yes, we tried MCTS. Yes, we tried PRMs. And none of that is useful.”
Shawn Wang Jan 24, 2025 ▶ 10:32 The Unreasonable Effectiveness of Reasoning Distillation: using DeepSeek R1 to beat OpenAI o1
Assertion Supported
Zhang: XGrammar outperforms Outlines and is integrated into TensorRT-LLM
“And I think Xgrammar's performance is better than the outline's, and also in the TensorFlow RTLM, the latest release, TensorFlow RTLM also integrates Xgrammar as the backend for the constructed coding.”
Yining Zhang Jan 19, 2025 ▶ 41:22 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Assertion Not checkable as stated
Automatic1111's inefficient SDXL implementation drove ComfyUI's viral user adoption
“The big, one point zero release happened, and wow, Confu UI was the only way a lot of people could actually run it on their computers, because it just, like, automatic was so, like, inefficient and bad that most people couldn't act, like, it just wouldn't work…”
comfyanonymous (Comfy) Jan 4, 2025 ▶ 47:23 AI Engineering for Art - with comfyanonymous
Prediction Open · timeframe Jan 2028
Swyx predicts OpenAI will launch a $2,000 per month ChatGPT tier
“I think that 2000 dollars ChatGPT will come.”
Shawn Wang Jan 1, 2025 ▶ 18:02 2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
Assertion Not checkable as stated
Swyx: VC appetite for GPU-rich early-stage startups is completely gone
“The appetite for GPU rich startups, like the, you know, the funding plan is we will raise sixty million and we'll give 50 of that to Nvidia. That is gone, right? Like no one's pitching that. This was literally the plan, the exact plan of like, I can name like …”
Shawn Wang Jan 1, 2025 ▶ 38:52 2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
Prediction Not checkable as stated
Fanelli: 2025 will be the first year AI sets job skill floors
“And I think the skill floor more and more, I think, 20, 25 will be the first year where the AI sets the skill floor of a role.”
Alessio Fanelli Jan 1, 2025 ▶ 1:41:17 2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
Assertion Supported
Neubig: SWE-bench Scores Are Inflated By Training Data Contamination
“Sweebench is on popular open source repos and all of these popular open source repos were included in the training data for all of the language models. And so, the language models already know these repos. In some cases, the language models already know the in…”
Graham Neubig Dec 25, 2024 ▶ 32:49 Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
Prediction Not checkable as stated
Ben Allal: Properly curated synthetic data prevents model collapse
“And I think there's a lot of concerns about model collapse, and I'm going to talk about that later, but we'll see that like, if we use synthetic data properly and we curate it carefully that shouldn't happen.”
Loubna Ben Allal Dec 24, 2024 ▶ 2:09 Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
Assertion Supported
Ben Allal: Hugging Face SmolLM2-1.7B outperforms Llama 3.2 models
“So it's a series of three models, which are the best in class in each model size. For example, our 1.7 B model outperforms Lama one B and also .2.”
Loubna Ben Allal Dec 24, 2024 ▶ 22:31 Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
Prediction Not checkable as stated
Ben Allal: AI industry will shift to fine-tuning over prompt engineering
“And I think we're going back to fine tuning where we realize these models are really cosplay. It's better to use just a small model. We try to specialize it. So I think it's a little bit of a cycle and we're going to start to see like more of fine tuning and l…”
Loubna Ben Allal Dec 24, 2024 ▶ 27:47 Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
Prediction Open · timeframe Dec 2029
Fu: Real-time long-context video generation cannot use quadratic attention
“You're certainly not going to do a giant quadratic attention computation to try to run that.”
Dan Fu Dec 24, 2024 ▶ 31:33 2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
Assertion Supported
Reddy: Chai Discovery's open-source Chai-1 model outperforms Google's AlphaFold 3
“We're lucky to work with the folks at Chai Discovery who just released Chai One, which is open source model that outperforms Alpha Fold Three.”
Pranav Reddy Dec 21, 2024 ▶ 7:44 The State of AI Startups in 2024 [LS Live @ NeurIPS]
Prediction Not checkable as stated
Guo: Cheaper AI code generation will increase software volume, not replace developers
“If we take the cost of software and high quality software down two orders of magnitude, we're just gonna end up with more software in the world. We're not gonna end up with fewer people doing development.”
Sarah Guo Dec 21, 2024 ▶ 20:27 The State of AI Startups in 2024 [LS Live @ NeurIPS]
Assertion Not checkable as stated
Mohan: Codeium generated dynamic PNGs due to VS Code API limitations
“On VS Code, actually, the problem for us wasn't actually being able to implement the feature. We had the feature for a while. Problem was actually even to show the feature, VS Code would not expose an API for us to do this. So what we actually ended up doing w…”
Varun Mohan Dec 13, 2024 ▶ 8:12 Windsurf: The Enterprise AI IDE
Assertion Supported
Cerebras WSE-3 runs Llama inference 70x faster than NVIDIA GPUs
“Cerebris came out that the wafer scale engine three can serve llama 70 B at 2.1 thousand sorry, 202,100 tokens per second and serves llama four or five B at nearly 1000 tokens per second. So this, you know, to give you an understanding, like this is about 70 t…”
Sarah Chieng Dec 7, 2024 ▶ 3:06 [Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras)
Prediction Not checkable as stated
Friedman: Dedicated AI agents will prevail over general computer-use models in enterprise
“We're seeing it for a while, and I think it will stay like that despite the computer use, et cetera, that supposedly can just replace us, and it could, you can just, like, prompt it to be, hey, now be a QA, you know, or be a QA person or developer. I still thi…”
Itamar Friedman Dec 2, 2024 ▶ 6:43 0 to over $8M ARR in 2 months as a Claude Wrapper (Bolt.new, Qodo)
Assertion Contradicted
Friedman: GitHub Copilot user retention in enterprise is 38% to 50%
“Between 38 to 50% Retention for users using Copilot and Enterprise.”
Itamar Friedman Dec 2, 2024 ▶ 37:04 0 to over $8M ARR in 2 months as a Claude Wrapper (Bolt.new, Qodo)
Assertion Supported
Friedman: AlphaCodium boosts OpenAI o1, proving o1 lacks true System 2
“We took their all one preview with Alpha Codium and did better. Like it just shows like, and there is a big difference between the preview and the IOI. It shows, like, that these models are not still system two thinkers, and there's a big difference.”
Itamar Friedman Dec 2, 2024 ▶ 42:48 0 to over $8M ARR in 2 months as a Claude Wrapper (Bolt.new, Qodo)
Prediction Not checkable as stated
Crivello: A Few Very Big Horizontal AI Agent Platforms Will Dominate
“I think AI is going to completely penetrate every category of software, but then I also think there are going to be a few very, very, very big horizontal agents that serve a lot of functions for people.”
Florent Crivello Nov 15, 2024 ▶ 23:04 Agents @ Work: Lindy.ai (with live demo!)
Prediction Not checkable as stated
Crivello: AI Agent Software Will Create Trillions Replacing Human Labor
“It's going to sound like an exaggeration, but it is a fact that it's going to create trillions of dollars of value in a few years, right? It's going to, for the first time, we're actually having software directly replace human labor.”
Florent Crivello Nov 15, 2024 ▶ 33:52 Agents @ Work: Lindy.ai (with live demo!)
Prediction Not checkable as stated
Crivello: Infinite and cheap context windows will arrive within 18 months
“Now we just assume that infinite context windows are going to be here in a year or something, a year and a half and infinitely cheap as well. And dynamic compute is going to be here. Like we just assume all of these things are going to happen.”
Florent Crivello Nov 15, 2024 ▶ 36:16 Agents @ Work: Lindy.ai (with live demo!)
Assertion Not checkable as stated
Crivello: Marc Andreessen Blocked Him on Twitter Over AI Safety Views
“Like at some point, Marc Andreessen blocked me on Twitter and I, it hurt, frankly, I really look up to Marc Andreessen and I knew he would block me.”
Florent Crivello Nov 15, 2024 ▶ 1:01:05 Agents @ Work: Lindy.ai (with live demo!)
Prediction Not checkable as stated
Polu: Post-hyper-growth tech companies may increasingly eliminate traditional SaaS
“So it's interesting that we might see kind of a bad time for SaaS in post-hyper-growth tech companies. So it's still a big market, but it's not that big, because if you're not a tech company, You don't have the capabilities to reduce desk cost. If you're a hig…”
Stanislas Polu Nov 11, 2024 ▶ 50:35 Agents @ Work: Dust.tt — with Stanislas Polu
Prediction Not checkable as stated
Polu: The next generation will see billion-dollar companies with 20 engineers
“All generations of company might be the first billion dollar companies with engineering teams of 20 people. That would be so exciting as well. That would be so great. You know, you don't have the management hurdle. You're just 20 focused people with a lot of a…”
Stanislas Polu Nov 11, 2024 ▶ 53:53 Agents @ Work: Dust.tt — with Stanislas Polu
Assertion Supported
SWE-Bench public test splits enable trivial cheating via runtime PR retrieval
“The entire test split here is public. So you can do things like just overfit to the patches in the test set. You can do things like, let me add at runtime, pull the PR and just get the answer and just use it.”
Jesse Hu Oct 19, 2024 ▶ 15:00 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
Prediction Not checkable as stated
Houston: Fully autonomous AI knowledge workers will take a long time
“People sort of skip to like level five full autonomy, or we're going to have like an autonomous knowledge worker that's just going to take, that's going to, and then we won't need humans anymore kind of Projection that that's going to take a long time, but the…”
Drew Houston Oct 18, 2024 ▶ 7:26 Building the Silicon Brain - Drew Houston of Dropbox
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.