The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 1,824 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Insight
Patel: Tech moats are shallower than ever due to massive capex
“I think moats are as shallow as they've ever been, right? Because how fast things are moving, the size of the numbers that are being thrown around now, right? It's hundreds of billions of dollars for each major hyperscaler. The size of the numbers are so large…”
Dylan Patel Feb 26, 2026 ▶ 44:56 Dylan Patel Explains the AI War While Cooking | In-Context Cooking
Insight
O'Laughlin: If AI automated coding, finance work will be automated too
“I think coding is a little harder, if I'm being honest with you, and you're telling me the hard one got automated. Why can't the easy one get automated? So I started to ask myself, how much can we do? And the answer is, it feels like a skill issue.”
Doug O'Laughlin Feb 24, 2026 ▶ 26:54 Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
Insight
O'Laughlin: Reviewing AI agents requires past hands-on manual domain experience
“If you didn't pay any like human cognition to get there, I don't think you're going to be a great reviewer. One of the reasons why, you know, what makes that, that human feet, that loop well is because once upon a time you did that and you could make the three…”
Doug O'Laughlin Feb 24, 2026 ▶ 56:10 Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
Insight
Casado: AI breaks the mythical man-month by turning money directly into capability
“Another thing that is very different this time than in the history of computer science is, is in the past, if you raised money, then you basically had to wait for engineering to catch up, which famously doesn't scale. Like, the mythical man must take a very lo…”
Martin Casado Feb 19, 2026 ▶ 8:08 Inside AI’s $10B+ Capital Flywheel — Martin Casado & Sarah Wang of a16z
Insight
Casado: Most hardware robotics startups inevitably become vertical sector companies
“When it comes to hardware most companies will end up verticalizing. Like, if you're If you're investing in a robot company for an app, for agriculture, you're investing in an ag company, because that's the competition and that's the pricing and that's the sup…”
Martin Casado Feb 19, 2026 ▶ 20:44 Inside AI’s $10B+ Capital Flywheel — Martin Casado & Sarah Wang of a16z
Insight
Casado: A $1B model training run economically justifies a custom ASIC
“No, no, no, a billion dollar training run, a one billion dollar training run, it makes sense to actually do a custom ASIC if you can do it in time. The question now is timeline. Not money. Cause just rough math. If it's a billion-dollar training run, then the …”
Martin Casado Feb 19, 2026 ▶ 23:17 Inside AI’s $10B+ Capital Flywheel — Martin Casado & Sarah Wang of a16z
Insight
Casado: Dedicated AI coding models fail since programming requires general intelligence
“And I think one of the conclusions is, is like there's no such thing as a coding model. You know, like, that's not a thing. Like, you're talking to another human being, and it's good at coding, but like, it's got to be good at everything.”
Martin Casado Feb 19, 2026 ▶ 36:04 Inside AI’s $10B+ Capital Flywheel — Martin Casado & Sarah Wang of a16z
Insight
Swyx: Agent labs will capture higher margins than model labs
“I think what I've been calling agent labs, which are people who build on top of all the other models. We'll probably have a better time with the margins because they price against the end user hours spent or like human labor. Whereas models get commodity price…”
Shawn Wang Feb 19, 2026 ▶ 53:56 Inside AI’s $10B+ Capital Flywheel — Martin Casado & Sarah Wang of a16z
Insight
Jeff Dean: Analog computing loses power advantages at digital boundaries
“I mean, I think there's still a, there's also sort of the more exotic things like analog based computing substrates as opposed to digital ones. I'm, you know, I think those are super interesting cause they can be potentially low power. but I think you often …”
Jeff Dean Feb 12, 2026 ▶ 41:23 The AI Frontier: from Gemini 3 Deep Think distilling to Flash — Jeff Dean
Insight
AI Excels at Structure Prediction but Fails to Model Physical Folding
“Folding is the more complex process of actually understanding, like, how it goes from, like, this disordered state into, like, a structured, like, state, and that I don't think we've made that much progress on, but the idea of, like, yeah, going straight to th…”
Jeremy Wohlwend Feb 12, 2026 ▶ 10:43 🔬Generating Molecules, Not Just Models
Insight
Core AI capabilities must be built during pre-training, not just fine-tuned
“If there is a core capability that you actually care about, that capability should be part of the foundation and not a fine-tuned artifact.”
Pratyush Maini Feb 10, 2026 ▶ 18:53 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
Insight
Small specialized pre-trained models can match capabilities of larger fine-tuned models
“When you think of the fact that by doing specialized pre-training, you can train a smaller model, which is as capable as a much larger model when fine-tuned.”
Pratyush Maini Feb 10, 2026 ▶ 19:38 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
Insight
Bissell: Activation Steering and In-Context Learning Are Quantitatively Equivalent
“He actually has a paper that, as well as some, you know, others from the team and elsewhere, that go into the essentially equivalence of activation steering and in-context learning, and how those are from a, he thinks of everything in a cognitive neuroscience …”
Mark Bissell Feb 5, 2026 ▶ 30:57 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Insight
Garg: Context graphs for unstructured data will consist of small models
“And so where the unstructured data is such a large part of it and conversation data is such a large part, most likely the context graph will actually consist of small models. The context graph itself is a model layer. Like you use that data to train a set of m…”
Ashu Garg Feb 4, 2026 ▶ 20:50 ⚡️Context Graphs: according to the authors — Jaya Gupta, Ashu Garg, Foundation Capital
Insight
White: AI Science Agents Do Not Require Bespoke Automated Labs
“We don't actually have to hold their hands so much anymore, or like they don't actually need to necessarily have an automated lab. They can like write an email to a CRO or something, or they can like tell you what experiment to do. And you can take a video of …”
Andrew White Jan 28, 2026 ▶ 16:56 🔬 From Red Teaming GPT-4 to Automating Drug Discovery: The Future of AI in Science — Andrew White
Insight
White: Simulating biological systems hits fundamental limits; empirical lab measurement is unavoidable
“Basically the best model whatever, Opus seven or GPT 10, like it really can only propose the first experiment, maybe slightly more clever, but at a certain point you just need information, right? Like some little calculations you can do that, like, there's mor…”
Andrew White Jan 28, 2026 ▶ 17:53 🔬 From Red Teaming GPT-4 to Automating Drug Discovery: The Future of AI in Science — Andrew White
Insight
White: RLHF fails on scientific hypotheses by ignoring impact and information gain
“We learned a lot about how bad our LHF is with people, just like people pay really attention to the tone, to the details, to like how many specific facts or figures on the hypothesis, right? Like actionability about like if the experiment is feasible, but what…”
Andrew White Jan 28, 2026 ▶ 20:35 🔬 From Red Teaming GPT-4 to Automating Drug Discovery: The Future of AI in Science — Andrew White
Insight
White: Verifiers in the loop provide higher signal than hypothesis rankings
“I have a lot more faith in these, like verifier-in-the-loop kind of scenarios where you have either data analysis, literature search, or you're running a unit test, or whatever, you're going and running the experiment. Anything like that, I think, is going to …”
Andrew White Jan 28, 2026 ▶ 24:58 🔬 From Red Teaming GPT-4 to Automating Drug Discovery: The Future of AI in Science — Andrew White
Insight
White: A world model serves as unifying glue for scientific agents
“In cosmos, we basically, we had all the pieces sitting around. We working on world models, we working on data analysis agent, working on literature agent. And then we're working on, you know, we built a platform for scientific agents. So we had things that can…”
Andrew White Jan 28, 2026 ▶ 38:21 🔬 From Red Teaming GPT-4 to Automating Drug Discovery: The Future of AI in Science — Andrew White
Insight
Yi Tay: Gradient descent learning paradigm is AI's bottleneck, not architecture
“It's not architecture itself. That's, that there's a problem that we, that is more of like the learning paradigm itself rather than the architecture itself. I think the architecture is just basically like the interface between the learning algorithm and the to…”
Yi Tay Jan 23, 2026 ▶ 49:07 Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay
Insight
Yi Tay: The 'Bitter Lesson' Is Overapplied; Architectural Ideas Fundamentally Matter
“The bitter lesson gets used too much in, like, too conveniently used around, but actually there's also a little bit of a, not a bit, there's also a sweet lesson where it's like, ideas matter.”
Yi Tay Jan 23, 2026 ▶ 52:19 Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay
Insight
Reggio: MCP belongs on conventional systems, not inter-agent communication
“I think of like the MCP and tool usage as being like the interface to all of our conventional imperative systems, not at the AI space.”
James Reggio Jan 17, 2026 ▶ 21:40 Brex’s AI Hail Mary — With CTO James Reggio (acquired for $5B by Capital One!)
Insight
Reggio: Agentic development amplifies bad architecture, yielding less net capacity increase
“I view agentic development as being something that amplifies all the good, just as much as it amplifies all the bad. And it amplifies sloppiness, poor architectural thinking misunderstanding of the requirements. Like there are, for all of the acceleration of g…”
James Reggio Jan 17, 2026 ▶ 1:04:27 Brex’s AI Hail Mary — With CTO James Reggio (acquired for $5B by Capital One!)
Insight
Reggio: Constraining LLMs to deterministic DAGs undersells their planning power
“My intuition has been that trying to craft LLMs into deterministic workflows and DAGs is, is kind of underselling like the power that they have to actually plan and execute more in a more sophisticated, like fluid way.”
James Reggio Jan 17, 2026 ▶ 1:08:04 Brex’s AI Hail Mary — With CTO James Reggio (acquired for $5B by Capital One!)
Insight
Hill-Smith: Widely tracked AI benchmarks improve without reflecting general intelligence gains
“Once an eval becomes the thing that everyone's looking at, schools can get better on it without there being a reflection of overall generalized intelligence of these models getting better. That has been true for the last couple of years. It'll be true for the …”
Micah Hill-Smith Jan 9, 2026 ▶ 15:22 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Insight
Cameron: Models perform better with minimal tools than rigid frameworks
“I think where we're getting to is that these models have gotten smart enough, they've gotten better, better tools that they can perform better when just given a minimalist set of tools and let them run, let the model Control the agentic workflow rather than us…”
George Cameron Jan 9, 2026 ▶ 51:27 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Insight
Bissell: Probing internal model features matches LLM-as-a-judge quality at 500x lower cost
“If you ask that model, try to like use it as an LLM as a judge, it's not very good. But if you probe its mind and you sort of detect when the features related to personally identifiable information are firing, that gets you the highest recall of anything. It's…”
Mark Bissell Dec 31, 2025 ▶ 9:53 [State of MechInterp] SAEs in Production, Circuit Tracing, AI4Science, "Pragmatic" Interp — Goodfire
Insight
Eysenbach: 1,000-layer RL requires reward-free objectives, not just architectural tricks
“I think the main conclusion is that using big networks not only requires these architectural tricks, but also, as Kevin mentioned before, it requires using a different objective. This objective doesn't actually use rewards in it, and so there's another word in…”
Benjamin Eysenbach Dec 31, 2025 ▶ 8:08 [NeurIPS Best Paper] 1000 Layer Networks for Self-Supervised RL — Kevin Wang et al, Princeton
Insight
Kevin Wang: Cross-entropy trajectory classification enables scalable deep reinforcement learning
“I think it's because we're fundamentally shifting the burden of learning from something like, Q-learning or, like, regressing to, like, TD errors, which we know is quite spurious and noisy and biased, to fundamentally, like, a classification problem. We're try…”
Kevin Wang Dec 31, 2025 ▶ 9:56 [NeurIPS Best Paper] 1000 Layer Networks for Self-Supervised RL — Kevin Wang et al, Princeton
Insight
McGrath: RLHF and RLVR differ by data quality, not optimization math
“Really, at the end of the day, like, RLHF, RLVR, They're both policy gradient methods, but the, what's different is just like the input data.”
Josh McGrath Dec 31, 2025 ▶ 9:02 [State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI
Insight
McGrath: DeepSeek Math's real breakthrough is verifiable reward trust, not GRPO
“As you said, it came out in the deep seek math paper, and like, it's an interesting optimization method, but it's like the more interesting thing that they have a new reward signal that they sort of like re that we can really, really trust. Like when, you know…”
Josh McGrath Dec 31, 2025 ▶ 12:42 [State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI
Insight
Nair: Academia rewards complex math over simple, generalizable solutions
“One of the pitfalls of academia is that it doesn't really reward, like, simple ideas that work, and instead kind of tends to reward, like, kind of mathier ideas. Those mathier ideas also give you these, like, kind of implicit knobs to tune that allow you to, l…”
Ashvin Nair Dec 30, 2025 ▶ 10:52 [State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
Insight
Nair: RL on LLMs is peaky and fails to generalize beyond training
“RL, the way it's applied to LLMs right now, is kind of a weird, funny tool where it doesn't really generalize beyond the training distribution that much. It generalizes to some extent, and generalizes in interesting ways, but It's like very peaky, right? Like …”
Ashvin Nair Dec 30, 2025 ▶ 12:26 [State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
Insight
Nair: Context integration, not model intelligence, bottlenecks useful automation
“A big thing that needs to happen is, like, it's not, it doesn't feel like intelligence of the models is the bottleneck. It's more like you just have products that bring the entire context of what someone wants to do into the product so that the LLM can, like, …”
Ashvin Nair Dec 30, 2025 ▶ 13:09 [State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
Insight
Nair: RLHF Is a Side Branch Because Compute Cannot Be Scaled
“I think human feedback is kind of like a bit of like a side branch, because you can't really pour that much compute Into it, right? It's like, you take the model, and you, like, elicit it to be a little bit better in terms of personality”
Ashvin Nair Dec 30, 2025 ▶ 23:16 [State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
Insight
Yegge: IDEs should only run in background for AI agents
“Yeah, so you, all you do is leave IntelliJ running, but you shouldn't look in it. It's a tool for the AI now, right?”
Steve Yegge Dec 26, 2025 ▶ 10:59 Steve Yegge's Vibe Coding Manifesto: Why Claude Code Isn't It & What Comes After the IDE
Insight
Yegge: 10x AI productivity turns merging into architectural rewrites
“As soon as you get to the point where like every developer is 10 times as productive, merging their code becomes this incredibly complicated problem because I, you and I work at the same time for two or three hours. We make, you know, 30,000 line change each. …”
Steve Yegge Dec 26, 2025 ▶ 19:18 Steve Yegge's Vibe Coding Manifesto: Why Claude Code Isn't It & What Comes After the IDE
Insight
Fioca: AI models develop operational habits during training analogous to muscle memory
“This is one of the coolest things about, like, model training is literally, like, they develop habits. It's just like a person does. Like, if you're, like, working on some podcasting tool, right, you're really good at editing, and then somebody makes you use a…”
Brian Fioca Dec 26, 2025 ▶ 8:37 ⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust — Brian Fioca + Bill Chen, OpenAI
Insight
Zhang: Superhuman computer vision requires RLHF rather than human SFT data
“But if you only do SFT and the SFT data is annotated by human, then your performance is funded by human. You cannot get, kind of, superhuman performance just by, kind of, this kind of data engine approach to use human annotated data and then learn from that. Y…”
Pengchuan Zhang Dec 18, 2025 ▶ 44:33 SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow)
Insight
AI will widen the performance gap between good and lazy engineers
“I do believe that AI will do only one thing. It will Separate faster the good engineers from the bad engineers. If you're a good engineer, and you're using AI well, you will be an amazing engineer. If you're a poor, lazy engineer, and you don't want to underst…”
Loïc Houssier Dec 11, 2025 ▶ 1:09:00 The Future of Email: Superhuman CTO on Your Inbox As the Real AI Agent (Not ChatGPT) — Loïc Houssier
Insight
Pure technical moats are gone; UX is now the only differentiator
“In the world where the technical moat is not that a moat anymore, because Like startups in two weeks, they can build something that is close to what you're building. The difference is like the, how you think about the user or the flow and all of that.”
Loïc Houssier Dec 11, 2025 ▶ 1:09:53 The Future of Email: Superhuman CTO on Your Inbox As the Real AI Agent (Not ChatGPT) — Loïc Houssier
Insight
Goyal: Publishing Public Benchmarks Is Marketing, Not Product Improvement
“It's just that the value proposition of publishing an eval is completely orthogonal to the value proposition of building evals in service of building a good product. I think the purpose of publishing benchmarks is marketing, and it's good marketing.”
Ankur Goyal Dec 7, 2025 ▶ 7:53 The Great Evals Debate — Ankur Goyal & Malte Ubl
Insight
Malte Ubl: When Vibes and Eval Data Disagree, Vibes Are Right
“I think that the common quip that if the vibes and the data disagree, the vibes are probably right. It's true, right? So you have to like, be honest with yourself, like, do they agree and kind of iterate On them over time.”
Malte Ubl Dec 7, 2025 ▶ 12:13 The Great Evals Debate — Ankur Goyal & Malte Ubl
Insight
Goyal: Providing eval criteria and examples is more effective than writing specs
“In many ways coming to the table of product building with representative examples and criteria that articulate what good versus bad is for a use case is just a more precise and usable form of product management than writing a spec.”
Ankur Goyal Dec 7, 2025 ▶ 18:36 The Great Evals Debate — Ankur Goyal & Malte Ubl
Insight
Proprietary data cannot be accurately valued without training models to evaluate capabilities
“I don't think you can value it unless you actually model it yourself and see what the capabilities are.”
Pim de Witte Dec 6, 2025 ▶ 35:44 World Models & General Intuition: Khosla's largest bet since LLMs & OpenAI
Insight
Founders negotiating large AI data deals should demand equity, says de Witte
“I would recommend, if you're gonna do large data deals, like, just try to get like a large chunk of equity in the company that you're doing it with if you can.”
Pim de Witte Dec 6, 2025 ▶ 36:22 World Models & General Intuition: Khosla's largest bet since LLMs & OpenAI
Insight
Making billions of gameplay clips playable bridges imitation learning to reinforcement learning
“Actually making every single clip on the platform playable at billions of clips scale is how we go from imitation learning to RL.”
Pim de Witte Dec 6, 2025 ▶ 1:00:30 World Models & General Intuition: Khosla's largest bet since LLMs & OpenAI
Insight
Johnson: Pixels offer a more lossless world representation than tokenized text
“And then like you actually lose something if you translate to this like purely tokenized representations that we use in LLMs, right? Like you lose the font, you lose the line breaks, you lose sort of the two D arrangement on the page. And for a lot of cases, f…”
Justin Johnson Nov 25, 2025 ▶ 22:11 After LLMs: Spatial Intelligence and World Models — Fei-Fei Li & Justin Johnson, World Labs
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.