The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 790 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Opinion
Brown: Previous multi-agent research was heuristic and ignored Bitter Lesson
“I think that a lot of the approaches that have been taken have been very heuristic and haven't really been following like the bitter lesson approach to scaling and research.”
Noam Brown Jun 19, 2025 ▶ 44:58 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Opinion
Brown: Scaling self-play beyond zero-sum games will not be as easy as AlphaGo
“My point is that, like, this is where the AlphaGo analogy breaks down. And, not necessarily breaks down, but, like, it's not going to be as easy as self-play was in AlphaGo.”
Noam Brown Jun 19, 2025 ▶ 58:29 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Opinion
PyTorch democratized model training, but AI inference remains undemocratized
“And so things like PyTorch came on the scene and I think PyTorch gets all credit for democratizing model training, right? It's taught to pretty much every computer science student that graduates. That's a huge deal, but nobody democratized inference. Inference…”
Chris Lattner Jun 13, 2025 ▶ 35:44 The Shape of Compute (Chris Lattner of Modular)
Opinion
DeepSeek's open releases pulled global AI progress forward by six months
“But I think that what it did is it pulled forward progress in AI by like six months.”
Chris Lattner Jun 13, 2025 ▶ 57:54 The Shape of Compute (Chris Lattner of Modular)
Opinion
Ameisen: Current LLM Chain of Thought Is Unfaithful and Untrustworthy
“So I think there's like a sense in which right now the chain of thought is, is unfaithful, or at least you can't read the chain of thought and trust that that's how the model did it.”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:37:12 The Utility of Interpretability — Emmanuel Amiesen
Opinion
Docker stopped innovating after Solomon Hykes left the company
“At some point Docker basically stopped innovating and I left and the ecosystem continued around containers, but it didn't actually Pick up where we left off. Everything towards applying container tech to development kind of stopped.”
Solomon Hykes Jun 3, 2025 ▶ 7:43 [AIEWF Preview] Containing Agent Chaos — Solomon Hykes
Opinion
Hykes: AI agent tooling is trending toward proprietary vertical monoliths
“I'm seeing us heading in the direction of highly integrated, very vertical end to end monoliths, you know, and I don't want to name names, but it's just, it's a, it's definitely a market trend.”
Solomon Hykes Jun 3, 2025 ▶ 10:14 [AIEWF Preview] Containing Agent Chaos — Solomon Hykes
Opinion
Dockerfile was a 2013 stopgap and Compose was a cloned prototype
“Docker file was something we designed as a stopgap prototype. Thinking, oh, we'll clean this up later in, in 2013. You know, it's been more than 10 years. Compose was a clone of a clone that we acquired into the team and, like, stitched on top, and then as soo…”
Solomon Hykes Jun 3, 2025 ▶ 12:03 [AIEWF Preview] Containing Agent Chaos — Solomon Hykes
Opinion
Enterprises demand full AI delegation, not 15% to 20% speed gains
“I think that the product experience of delegation is really, really immature right now. And most enterprises though, see that as the holy grail, not like going 15% or 20% faster.”
Eno Reyes May 29, 2025 ▶ 9:55 The AI Coding Factory
Opinion
Cherny: Claude Code delivers up to 10x productivity gains for Anthropic engineers
“Anecdotally for me, it's probably two X my productivity. So I'm just like, I'm an engineer that codes all day, every day. For me, it's probably two X. Yeah. I think there's some engineers at Anthropic where It's probably 10 X their productivity”
Boris Cherny May 7, 2025 ▶ 1:05:57 Claude Code: Anthropic's CLI Agent
Opinion
Sobo: AI editor moats are shrinking rapidly due to better tool-calling LLMs
“So there's a, there was a lot of like integration work to bridge that gap, but now the LLMs, because they've taken on this tool calling and are getting better at tool calling, the story is the integration story has definitely gotten a lot easier. I think it's …”
Nathan Sobo May 7, 2025 ▶ 11:56 Zed Agents — with Zed Cofounders Nathan Sobo & Antonio Scandurra
Opinion
Sobo: Zed's competitive moat forces competitors to build an editor from scratch
“I mean, to me, the moat Z's moat is like implementing really nice product experiences with technical excellence and their performance, like just delivering a really great user experience that in order to compete with, you would have to Essentially build your o…”
Nathan Sobo May 7, 2025 ▶ 12:13 Zed Agents — with Zed Cofounders Nathan Sobo & Antonio Scandurra
Opinion
Bergum: pgvector is outpacing some standalone vector DBs in search capabilities
“So actually, what you can see, PJ Vector is doing more in the capabilities of vector search than some of the real vector database players, right?”
Jo Kristian Bergum Apr 19, 2025 ▶ 11:02 The Rise and Fall of the Vector DB category: Jo Kristian Bergum (ex-Chief Scientist, Vespa)
Opinion
Bergum: In-database ML like PostgresML is the wrong architectural direction
“No, I'm not. I'm sorry. I, I'm not. I think this is, yeah, we also seen other players that tries to, you know, move a lot of the logic into the database, agentic embedding inference and whatnot. I think if the right direction is to Keep infrastructure a little…”
Jo Kristian Bergum Apr 19, 2025 ▶ 17:24 The Rise and Fall of the Vector DB category: Jo Kristian Bergum (ex-Chief Scientist, Vespa)
Opinion
Conrad: SF Compute is the only true liquid bid-ask GPU market
“Turned into what is today SF compute, which is a compute market, which we think we are the functionally the most liquid GPU market of any capacity. Honestly, I think we're the only thing that actually is like a real market that there's like bids and asks and t…”
Evan Conrad Apr 11, 2025 ▶ 24:36 SF Compute: Commoditizing Compute
Opinion
Shah: The AI industry does not need a new language to replace Python
“He was talking about like, oh, we need a different language than Python or whatever that is like built for built for AI and built. It's like, No, Brett, I don't think we do actually. It's just fine. It deals with just fine, just expressive enough. And it's nic…”
Dharmesh Shah Mar 28, 2025 ▶ 50:27 The Agent Network — Dharmesh Shah, Agent.ai + CTO of HubSpot
Opinion
Swix: Descript still sucks despite major funding and OpenAI backing
“Descript is so much funding. They had OpenAI invested in them, and they still suck.”
Shawn Wang Mar 14, 2025 ▶ 39:59 Snipd: The AI Podcast App for Learning — with CEO Kevin Ben-Smith
Opinion
Swix: Most podcasts are bad because creators lack real-world experience
“The reason that most podcasts or YouTube videos are shit is they're made by people who don't have life experience, who are not that important in the world. They're not doing important jobs. And so what you want to actually enable is CEOs to each of them make t…”
Shawn Wang Mar 14, 2025 ▶ 1:14:55 Snipd: The AI Podcast App for Learning — with CEO Kevin Ben-Smith
Opinion
Husain: 80% of LLM-as-a-judge implementations are unhelpful
“I feel like a 75% LMS judge because it's low effort is kind of easy, but I would say out of the 75%, 80% is not helpful.”
Hamel Husain Mar 13, 2025 ▶ 14:58 [Lightning Pod] Evals: How to Improve AI Consistently — with Hamel Husain and Shreya Shankar
Opinion
Klein: OpenAI's brand constraints will prevent them from offering CAPTCHA solving
“I think it's going to be really hard for a company like OpenAI to do things like support CAPTCHA solving or like have proxies. Like, I think it's hard for them structurally. Imagine this New York Times headline, OpenAI CAPTCHA solving. Like, that would be a pr…”
Paul Klein Feb 28, 2025 ▶ 40:37 Browserbase: Browser Infrastructure For Your AI Agents
Opinion
Embeddings will not scale for cross-session memory in real-time AI systems
“I don't think that like embeddings are going to be able to scale to, I think they work well for some of this as like kind of the MVP version of the experience, but I think you're going to need a different experience and it's Probably something like really smar…”
Logan Kilpatrick Feb 28, 2025 ▶ 23:41 Gemini 2.0 Flash and Flash Thinking: the new SOTA models for the agentic era
Opinion
Swix: Deep research agents are the first agent category with true PMF
“What are the hard problems in this brand of agent that is like probably the first real product market fit agent. I will say more so than the computer use ones. This is the one where like, yeah, people are like, yeah, easily pays for 200 dollars worth a month w…”
Shawn Wang Feb 18, 2025 ▶ 58:12 Why is everyone cloning Deep Research?
Opinion
Sutin: The Humane AI Pin failed due to weight and overheating
“I think the Humane is like pretty incredible. Some of the engineering they did, but like, it wasn't kind of geared towards solving the problem. It was just it's too heavy. The swappable batteries is too much demand. The heat, the thermals is like too much.”
Ethan Sutin Feb 17, 2025 ▶ 43:56 Bee AI: The Wearable Ambient Agent
Opinion
Startups Cannot Profitably Pre-Train Foundation Models Due to Massive Compute Costs
“I think the idea that a company can make money sort of pre-training a foundation model is probably not true. It's hard to, you're competing with just you know, unreasonably large capex budgets.”
Bret Taylor Feb 11, 2025 ▶ 23:30 The AI Architect: Bret Taylor
Opinion
Bret Taylor Mocks Using AI to Write Unsafe, Slow Python Code
“I think that we're bringing the cost of writing code down to zero. So the fact that we're still writing Python with AI cracks me up just cause it's like literally was designed to be ergonomic to write, not, not safe to run or fast to run.”
Bret Taylor Feb 11, 2025 ▶ 38:49 The AI Architect: Bret Taylor
Opinion
Colvin: Jupyter Notebooks Are Worse Than or Similarly Bad to Excel
“I have very strong opinions about, you know, proper, like Jupyter notebooks, this idea that like you have to run the cells in the right order. I mean, a whole bunch of things. It's basically like worse than Excel or similarly bad to Excel.”
Samuel Colvin Feb 6, 2025 ▶ 56:25 Agent Engineering with Pydantic + Graphs — with Samuel Colvin, CEO of Pydantic Logfire
Opinion
Beauchamp: Crypto's only killer use cases are gambling and evading regulations
“As far as a gambling device is like the most fun form of gambling invented in like ever. Super fun. I thought as a way to evade monetary regulations and banking restrictions, I think it's also absolutely amazing. So it has two like Killer use cases.”
William Beauchamp Jan 26, 2025 ▶ 5:43 Outlasting Noam Shazeer, Crowdsourcing Chai AI w/ 1.4m DAU — with William Beauchamp, Chai Research
Opinion
Beauchamp: AI content should be created and trained by users, not Silicon Valley
“The number one problem for users in AI is this. All the AI is being generated by middle-aged men in Silicon Valley, right? That's all the content. You're interacting with this AI. You're speaking to it for 90 minutes on average. It's being trained by a middle-…”
William Beauchamp Jan 26, 2025 ▶ 43:14 Outlasting Noam Shazeer, Crowdsourcing Chai AI w/ 1.4m DAU — with William Beauchamp, Chai Research
Opinion
AMD GPUs perform horribly on Windows for generative AI tasks
“AMD works horribly on Windows. Like, on Linux, it works fine. It's lower than the price equivalent NVIDIA GPU. But it works, like, you can use it, generate images, everything works. On Linux, on Windows, You might have a hard time”
comfyanonymous (Comfy) Jan 4, 2025 ▶ 34:09 AI Engineering for Art - with comfyanonymous
Opinion
Swyx: DeepMind has roughly a four-year advantage over OpenAI in world modeling
“So like they have maybe four years advantage on world modeling that OpenAI does not have. Cause OpenAI basically only started Diffusion Transformers last year when they hired Build Peebles. So DeepMind has a bit of advantage here.”
Shawn Wang Jan 1, 2025 ▶ 51:53 2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
Opinion
Soldani: Web blocking disproportionately benefits incumbent closed AI labs
“And I think the problem is this blocking or ideas really, it impacts people in different ways. It disproportionately helps companies that have a head start, which are usually the closed labs, and it hurts incoming newcomer players where you either have now to …”
Luca Soldani Dec 23, 2024 ▶ 20:34 Best of 2024: Open Models [LS LIVE! at NeurIPS 2024]
Opinion
Robinson: YOLO architectures have hit a performance plateau
“So, for years, yellows have been the dominant way of doing real time object detection, and we can see here that they've essentially stagnated. The performance between 10 and 11 is not meaningfully different. At least, you know, in, in this type of high level c…”
Isaac Robinson Dec 22, 2024 ▶ 15:07 Best of 2024 in Vision [LS Live @ NeurIPS]
Opinion
Korupati: Vision-Language Models Are Lagging Behind LLMs in Reasoning
“LLMs are showing enormous progress in reasoning, especially with the latest set of models that we've seen, but we're not really seeing, I have a feeling that VLMs are lagging behind, as we can see with these tasks that should be very simple for a human to do t…”
Vik Korupati Dec 22, 2024 ▶ 53:01 Best of 2024 in Vision [LS Live @ NeurIPS]
Opinion
Schluntz: AI robotics today is where autonomous driving was 10 years ago
“I think where we are right now is where self-driving cars were 10 years ago. I think we have very cool demos that work. I mean, 10 years ago, you had videos of people driving a car on the highway, driving a car, you know, on a street with a safety driver, but …”
Erik Schluntz Nov 28, 2024 ▶ 1:02:19 The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic
Opinion
Schluntz: High vehicle costs make Waymo's per-car profitability doubtful
“Those cars are expensive. It's not about if you can hit profitability, it's about your cash conversion cycles. Like is building one Waymo, like how cheap can you make that compared to like how much you're earning sort of as the equivalent of what an Uber drive…”
Erik Schluntz Nov 28, 2024 ▶ 1:09:35 The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic
Opinion
Polu: DeepMind IMO breakthrough relied on scaling RL and autoformalization
“I think the DeepMind team just did a good job of scaling. I think there's nothing too magical in their approach, even if it hasn't been published as a Dan Silver talk from seven days ago, where it goes a little bit into more details. It feels like there's noth…”
Stanislas Polu Nov 11, 2024 ▶ 9:00 Agents @ Work: Dust.tt — with Stanislas Polu
Opinion
Polu: Anthropic split was driven by disagreement over OpenAI's API commercialization
“What I understood of it is that there was a disagreement of the commercialization of that technology. I think the focal point of the disagreement was the fact that we started working on the API and wanted to make those models available through an API. Is that …”
Stanislas Polu Nov 11, 2024 ▶ 17:03 Agents @ Work: Dust.tt — with Stanislas Polu
Opinion
Pullen: SWE-bench is a poor proxy for real-world AI coding competence
“I know Sweebench is, like, the most commonly talked about thing, and honestly, it's a very, it's an amazing project, but one of the things we've learned the most from actually shipping this product to users is, it's a pretty bad proxy at telling us how compete…”
Alistair Pullen Oct 4, 2024 ▶ 1:16:10 Building AGI in Real Time (OpenAI Dev Day 2024)
Opinion
Shawn Wang says Devin's breakthrough was Agent-Computer Interfaces, not advanced planning
“The planner is like actually pretty simple, but ACI. That they book through on.”
Shawn Wang Sep 27, 2024 ▶ 45:28 Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
Opinion
Schulhoff: Hiring Dedicated Prompt Engineers Makes No Sense for Most Companies
“I have always viewed prompt engineering as a skill that everybody should and will have, rather than a specialized role to hire for. That being said, there are definitely times where you do need just a prompt engineer. I think for AI companies, it's definitely …”
Sander Schulhoff Sep 20, 2024 ▶ 48:11 The Ultimate Guide to Prompting - with Sander Schulhoff from LearnPrompting.org
Opinion
Howard: Tech builds too many vanity foundation models over fine-tuning
“People are building too many vanity foundation models rather than taking better advantage of fine-tuning”
Jeremy Howard Aug 17, 2024 ▶ 25:06 Answer.ai & AI Magic with Jeremy Howard
Opinion
Howard: Building web apps is much worse now than 15 years ago
“Much to my, you know, horror, the story around creating web applications is much worse now than it was 10 or 15 years ago, in terms of, like, if I say to a data scientist, here's how to create and deploy a web application, You know, either you have to learn Ja…”
Jeremy Howard Aug 17, 2024 ▶ 48:50 Answer.ai & AI Magic with Jeremy Howard
Opinion
Howard: Cursor and VS Code shoehorn AI into legacy software paradigms
“It's like a convenience over the top of this incredibly complicated system that full-time, sophisticated software engineers have designed over the past few decades in a totally different environment as a way to build software, you know. And so we're trying to,…”
Jeremy Howard Aug 17, 2024 ▶ 59:58 Answer.ai & AI Magic with Jeremy Howard
Opinion
Scialom: AI has minted infrastructure unicorns but few successful application companies
“I see like now a lot of fundamental stacks that are like the unicorn of today. Foundational models, foundational like clusters, data notations, things like that. There's a lot, but less successful yet, for now at least, application company. And it's hard to bu…”
Thomas Scialom Jul 23, 2024 ▶ 1:02:32 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Opinion
Yi Tay: Llama 3 shows Meta may have caught up to Google
“So I think I don't really follow, like, fine much, but I think that, like, Lama Tree actually shows that, like, kind of, like, Meta got a pretty, like, a good stack around training these models you know, like, oh, and I've even started to feel like, oh, they a…”
Yi Tay Jul 5, 2024 ▶ 1:21:59 The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
Opinion
Tay: Mixture-of-Experts is fundamentally the right architecture for scaling
“Fundamentally, I just think that MOEs are just, like, the way to go in terms of, like, floppyram ratio, they bring the benefit from the scaling curve, if you do it right, if you, they bring the benefit from the scaling curve, right, and then, Like, that's, lik…”
Yi Tay Jul 5, 2024 ▶ 1:51:09 The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
Opinion
Tay: Hugging Face's Open LLM Leaderboard is a major problem
“The open LM leaderboard is, like, probably, like, the, a big, like, Problem, to be honest.”
Yi Tay Jul 5, 2024 ▶ 2:01:39 The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
Opinion
Albrecht: Training on AWS prevents diagnosing low-level hardware errors
“And if we're just using, you know, AWS or some other cloud provider, These errors are still going to be there, and you're gonna have no way to know and no way to debug this and no way to diagnose what's going wrong.”
Josh Albrecht Jun 25, 2024 ▶ 19:49 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.