The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 2,445 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 100 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Not checkable as stated
Kulik: No Current ML Potential Robustly Models All Materials Bonding
“The challenge is that you have a lot more than 20 building blocks when it comes to materials and so there's lots of different ways to think about chemical bonding, and right now no potentials are really robustly encoding all of that bonding, especially with re…”
Heather Kulik Mar 24, 2026 ▶ 24:42 🔬There Is No AlphaFold for Materials — AI for Materials Discovery with Heather Kulik
Assertion Not checkable as stated
Singleton: Stripe deployed some of the world's first production AI agent systems
“I was working at Stripe, as you mentioned, and we had the opportunity to put some of the very first AI agent systems in the world into production.”
David Singleton Mar 20, 2026 ▶ 3:59 Dreamer: the Agent OS for Everyone — David Singleton
Prediction Not checkable as stated
Rieseberg: Specialized AI wrapper apps won't survive as models generalize
“I think we're going to see a lot of like applications and companies that do very impressive things with AI that in the short term might seem very effective because they're very specialized to individual use cases. But I think once models get better at generali…”
Felix Rieseberg Mar 17, 2026 ▶ 25:29 Anthropic’s Felix Rieseberg on AI Coworkers, Local-First Agents, and the Future of Knowledge Work
Prediction Not checkable as stated
Rieseberg: AI takeoff will create an accelerating, self-reinforcing loop
“Big bang moment where things will accelerate so quickly that it becomes a self-reinforcing loop. And at that point it's sort of like off to the races and there will be no more like slowly catching up. You know, just have Claude being so good at everything.”
Felix Rieseberg Mar 17, 2026 ▶ 52:04 Anthropic’s Felix Rieseberg on AI Coworkers, Local-First Agents, and the Future of Knowledge Work
Assertion Supported
Eskildsen: Neon retrofitted Postgres for S3, while Turbopuffer built pure object-storage consensus
“I think neon neon was first to, and they're trying to retrofit it onto Postgres. And then they built this whole architecture where you have it in memory, and then you sort of like, you know, mmap back to S-III, and I think that was very novel at the time to do…”
Simon Eskildsen Mar 12, 2026 ▶ 16:38 Retrieval After RAG: Hybrid Search, Agents, and Database Design — Simon Eskildsen of Turbopuffer
Assertion Open · timeframe Mar 2027
Eskildsen: Turbopuffer outperforms Lucene on long LLM search queries
“Turbo Puffer today has a fairly start of the state of the art full text search engine. We beat Lucene on some queries, in particular, very long queries that we've optimized for, because those are the text search queries we see today.”
Simon Eskildsen Mar 12, 2026 ▶ 51:26 Retrieval After RAG: Hybrid Search, Agents, and Database Design — Simon Eskildsen of Turbopuffer
Assertion Supported
Shah: OpenClaw's 15-message replay fails prompt caching and costs 10x more
“The way OpenClaw does it is it essentially sends back the last 15 messages in the conversation and it essentially uses that back and forth. And I mean, the approach itself is not ideal because you will, like, you are not doing any, like, you're not utilizing a…”
Dhravya Shah Mar 9, 2026 ▶ 21:03 ⚡️ OpenClaw's Memory Sucks and the fix is simple — Dhravya Shah, Supermemory
Prediction Not checkable as stated
Algorithmic unhobblers could soon expand context windows to 100 million tokens
“I wouldn't be surprised if we do see the ability to, like, break through to, like, ten million, twenty million, a hundred million context through the, an unhobbler showing up.”
Kyle Kranen Mar 8, 2026 ▶ 52:37 Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
Prediction Open · timeframe Dec 2026
AI agents will achieve 24-hour self-consistent autonomous runtimes by late 2026
“We will see before the end of the year an agent that is capable of running for longer than 24 hours with like self consistency the entire time.”
Kyle Kranen Mar 8, 2026 ▶ 1:19:00 Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
Prediction Not checkable as stated
Nelle: Coding workflows will shift from inspecting diffs to previewing video demos
“That's going to happen again, where it goes from agents handing you back gifts and you're sort of like in the weeds and giving it, you know, 32nd to three minute tasks to you're giving it, you know, three minute to 30 minute to three hour tasks and you're gett…”
Jonas Nelle Mar 6, 2026 ▶ 20:09 Cursor's Third Era: Cloud Agents — ft. Sam Whitmore, Jonas Nelle, Cursor
Prediction Not checkable as stated
Nelle: 10-Person Startups Using Cloud Agents Will Need Enterprise-Scale Dev Pipelines
“As with cloud agents, you scale up this parallelism and how much code you generate, 10 person startups become Need the DevEx and pipelines that a 10,000 person company used to need.”
Jonas Nelle Mar 6, 2026 ▶ 26:49 Cursor's Third Era: Cloud Agents — ft. Sam Whitmore, Jonas Nelle, Cursor
Prediction Held up
Nelle: Developers will spend thousands to tens of thousands monthly on agents
“I think as we think about these highly parallel kind of agents running off for a long time in their own VM system, We are already at that point where people will be spending thousands of dollars a month per, per human, and I think potentially tens of thousands…”
Jonas Nelle Mar 6, 2026 ▶ 53:50 Cursor's Third Era: Cloud Agents — ft. Sam Whitmore, Jonas Nelle, Cursor
Prediction Not checkable as stated
Levie: A few wrong AI decisions could wipe software companies out
“You can just see how the AI tsunami could wipe you out. If you make just two, three, four or five wrong decisions in this space, like couple wrong architecture decisions, couple wrong AI feature decisions, couple wrong API platform decisions. And you might be …”
Aaron Levie Mar 5, 2026 ▶ 52:23 Why Every Agent Needs a Box — Aaron Levie, Box
Prediction Not checkable as stated
Levie: Distribution spending will match increased software code efficiency
“The thing that's going to happen on the ledger of software is we're going to produce far more output of code and thus features per dollar. But on the other end of this, we're going to actually end up spending probably just as much on how do you get all of that…”
Aaron Levie Mar 5, 2026 ▶ 1:13:15 Why Every Agent Needs a Box — Aaron Levie, Box
Prediction Not checkable as stated
Levie: AI agents will generate 10 to 100 times more code
“But no matter what, there's going to be 10 to a hundred times more code. So I think you can be very long engineering right now as just a, you know, purely on the dimension of software is going to become increasingly more important. Once agents are, you know, t…”
Aaron Levie Mar 5, 2026 ▶ 1:16:24 Why Every Agent Needs a Box — Aaron Levie, Box
Prediction Not checkable as stated
Becker: Operational Long Tail Will Delay Full AI R&D Automation
“There's this, Very long tail of things potentially involved in in R&D that would perhaps need to be fully automated in order to lead to capabilities explosion. I expect we're measuring, you know, in some ways only, only a small proportion of, only a small prop…”
Joel Becker Feb 27, 2026 ▶ 31:52 Measuring Exponential Trends Rising (in AI) — Joel Becker, METR
Prediction Not checkable as stated
Becker: Halving AI Compute Growth Halves Algorithmic Progress and Milestones
“And both of them both of those components half when compute halves sort of trivially, because compute is halving, and algorithmic progress halves because compute is this important input, and compute halves, then you might expect time horizon growth to half. An…”
Joel Becker Feb 27, 2026 ▶ 36:12 Measuring Exponential Trends Rising (in AI) — Joel Becker, METR
Prediction Open · timeframe Feb 2031
Dylan Patel: TSMC will obey US export controls even under KMT
“Even if the KMT wins, it's not like TSMC starts disobeying American export restrictions because The way American export restrictions are upheld is that Taiwan utilizes American banking systems, American equipment industries, and so they'll still have to uphold…”
Dylan Patel Feb 26, 2026 ▶ 9:48 Dylan Patel Explains the AI War While Cooking | In-Context Cooking
Assertion Supported
Patel: Claude Code's share of GitHub commits doubled to 4% in January
“Just in January, it went from four percent of or two percent of commits on GitHub to four percent of GitHub commits were done by Cloud Code, right?”
Dylan Patel Feb 26, 2026 ▶ 21:09 Dylan Patel Explains the AI War While Cooking | In-Context Cooking
Prediction Held up
Patel: Google and Amazon will borrow debt to fund AI infrastructure
“Google and Amazon haven't taken on debt yet for AI infrastructure, but they will, right?”
Dylan Patel Feb 26, 2026 ▶ 33:54 Dylan Patel Explains the AI War While Cooking | In-Context Cooking
Prediction Open · timeframe Dec 2027
Patel: Google will buy tons of GPUs through 2027 due to TPU limits
“When we look in 26, Google would buy a lot more TPUs, but they can't ramp production fast enough, right? And so they have to buy tons of GPUs. And we go to 27, it applies again, right? Google simply cannot buy enough TPUs, and they have to buy tons of GPUs.”
Dylan Patel Feb 26, 2026 ▶ 49:51 Dylan Patel Explains the AI War While Cooking | In-Context Cooking
Prediction Not checkable as stated
Patel: Semiconductor fab shortages will bottleneck AI through the end of the decade
“What's the bottleneck to building more fabs or to building more chips is more fabs, and people just have not built these fabs yet, right? And that's, I think, the big bottleneck now. And that's gonna persist through the end of the decade. Or until AI, you know…”
Dylan Patel Feb 26, 2026 ▶ 50:29 Dylan Patel Explains the AI War While Cooking | In-Context Cooking
Prediction Not checkable as stated
Welling: The Bitter Lesson of AI Scaling Will Overtake Materials Science
“The same bitter lessons or lessons that you can draw in LLM space are eventually going to be true in this space as well, I think.”
Max Welling Feb 25, 2026 ▶ 31:38 🔬Max Welling: Materials Underlie Everything
Assertion Supported
O'Laughlin: Anthropic does not train Claude agent teams with RL
“I have a controversial opinion that Claude does not do RL on the agent swarms or agent team.”
Doug O'Laughlin Feb 24, 2026 ▶ 33:33 Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
Assertion Not checkable as stated
O'Laughlin: Compiled a PhD-level chip cycle dataset in one day using AI
“I mean, this is like too much information to gather. It's like a lifetime of work. It's like a PhD project. I did it in a day.”
Doug O'Laughlin Feb 24, 2026 ▶ 46:46 Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
Prediction Not checkable as stated
O'Laughlin: Entry-level data analysis jobs are at risk from agentic AI
“I just can't imagine if I was an entry level worker doing data analysis that a 22 year old, an average 22 year old would murder the hell out of a relatively well thought out agentic system. And so you're like, yeah, that job actually does seem at risk.”
Doug O'Laughlin Feb 24, 2026 ▶ 50:25 Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
Prediction Not checkable as stated
O'Laughlin: Claude Code tools will become baseline for information work in 24 months
“I would argue what we'll see in the 24 month view, it will be a base level, I think. I think. Cloud code, co-work, whatever is going to be a base level of all information work very soon.”
Doug O'Laughlin Feb 24, 2026 ▶ 52:42 Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
Assertion Supported
O'Laughlin: AI build-out CapEx has massively passed the internet
“We've well massively passed the internet in terms of the absolute size of the build-out. It's not even close.”
Doug O'Laughlin Feb 24, 2026 ▶ 1:05:11 Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
Prediction Open · timeframe Dec 2026
O'Laughlin: AI Agents Will Write 25% to 50% of Public Code by Year-End
“I sandbagged the ever-living share of that. I just believe 25 is, Very, like, I, like, it's like a, the rate it's on is like, whatever, 50 or something like that, but I think, I feel, I wanted to give a 95 confidence interval. I think 25 is within the 95 confi…”
Doug O'Laughlin Feb 24, 2026 ▶ 1:15:51 Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
Prediction Not checkable as stated
O'Laughlin: TPU v8 will not compete as well against Nvidia Rubin
“V-A we just don't think will be as competitive to Ruben, and that's when your special window starts to close.”
Doug O'Laughlin Feb 24, 2026 ▶ 1:38:59 Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
Prediction Open · timeframe Feb 2028
O'Laughlin: Memory supply will not catch up with demand for two years
“And then boom, you're just looking at the supply demand and you're like, yeah, this is not gonna catch up for like two years.”
Doug O'Laughlin Feb 24, 2026 ▶ 1:44:05 Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
Prediction Not checkable as stated
O'Laughlin: DRAM prices could rise another 100%, causing demand destruction
“We, I, our posts, our conclusion is like, we could see DRAM prices like go up a hundred percent again. Like it's gonna be the point where, and this is like also example, like really interesting in the whole thing. Another hundred percent I think is demand dest…”
Doug O'Laughlin Feb 24, 2026 ▶ 1:44:32 Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
Prediction Not checkable as stated
O'Laughlin: Tech industry may face a CPU shortage from AI coding and RL
“You feel like we might actually be seeing a CPU shortage partially because of this refresh cycle, but partially also because like I legitimately believe the cloud code Cloud code is increasing software creation and then on top of that, there is real demand fro…”
Doug O'Laughlin Feb 24, 2026 ▶ 1:56:07 Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
Prediction Not checkable as stated
O'Laughlin: Memory shortages will price out low-end phones and gaming GPUs
“Memory prices are going to go up so much that we're going to have to choose which exactly. That's the crazy part to me. Historically, memory has never been a constraint like this where I said, actually, you're not going to get your low end. You're not going to…”
Doug O'Laughlin Feb 24, 2026 ▶ 1:57:11 Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
Assertion Supported
Watkins: Over half of SWE-bench problems investigated by OpenAI had test flaws
“In over half of the problems that were investigated in that deep dive, there was one problem or the other. I think the most common problem are, like, overly narrow tests where there's some particular implementation detail that the tests were looking for but wa…”
Olivia Watkins Feb 23, 2026 ▶ 7:26 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
Assertion Not checkable as stated
Watkins: SWE-bench Verified is contaminated across OpenAI, Claude, and Gemini models
“And in SweetBenchVerified, we found many instances of contamination across like, across OpenEye models, across, like, Quad Opus, 4.5, Gemini Flash, and all of these, we saw things like regurgitating the ground truth solutions, things like in some cases giving,…”
Olivia Watkins Feb 23, 2026 ▶ 11:54 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
Assertion Not checkable as stated
Glaese: OpenAI no longer trusts further score improvements on SWE-bench Verified
“Issues with the benchmark that means that now that we're at like 80%, we don't really trust like further improvements on it, but like it does measure something that is like a real like capability of models.”
Mia Glaese Feb 23, 2026 ▶ 14:34 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
Assertion Not checkable as stated
Casado: There is no compute supply overhang or 'dark GPUs'
“But we don't have a supply overhang. Like, there's no dark GPUs, right?”
Martin Casado Feb 19, 2026 ▶ 4:13 Inside AI’s $10B+ Capital Flywheel — Martin Casado & Sarah Wang of a16z
Assertion Not checkable as stated
Wang: L5 AI engineers can get offers in tens of millions
“You could be an L five and get an offer in the tens of millions.”
Sarah Wang Feb 19, 2026 ▶ 16:23 Inside AI’s $10B+ Capital Flywheel — Martin Casado & Sarah Wang of a16z
Assertion Not checkable as stated
Casado: Frontier Model Labs Are Gross Margin Negative Factoring Next-Gen Training
“If you look at the numbers of these companies, if you look at like the amount they're making and how much they spent training the last model, their gross margin positive, you're like, oh, that's really working. But if you look at like the current training that…”
Martin Casado Feb 19, 2026 ▶ 30:23 Inside AI’s $10B+ Capital Flywheel — Martin Casado & Sarah Wang of a16z
Assertion Not checkable as stated
Casado: Cursor built a near-SOTA coding model at 1/100th the cost
“So the interesting thing about cursors, they actually for, you know, a small fraction of the cost, a 100th the cost or less, developed an almost soda model, which for a period of time was the most popular coding model in the world, right?”
Martin Casado Feb 19, 2026 ▶ 52:29 Inside AI’s $10B+ Capital Flywheel — Martin Casado & Sarah Wang of a16z
Prediction Not checkable as stated
Dean: General AI models will win out over specialized ones
“I mean, I think general models will win out over specialized ones in most cases.”
Jeff Dean Feb 12, 2026 ▶ 49:39 The AI Frontier: from Gemini 3 Deep Think distilling to Flash — Jeff Dean
Assertion Open · timeframe Feb 2027
Modern LLMs verbatim regurgitate JEE exam questions from two-word prompts
“We consistently saw how many of these, like, models today are being, like, massively, like, kind of fine-tuned on problems from... Like, oversight? Massively worked with. Like, even, like, imagine if I ask you the light bulb, what comes next in your mind? It w…”
Pratyush Maini Feb 10, 2026 ▶ 3:36 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
Assertion Not checkable as stated
Labs train LLMs on benchmark questions for multiple epochs late in training
“It's very clear how the last stage of training for many of these models does have a massive amount of example or examine. Benchmaxing. Because the model will not, like, behaviorally complete exam questions with options if they have not really seen it at the en…”
Pratyush Maini Feb 10, 2026 ▶ 4:17 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
Assertion Not checkable as stated
Self-reflection training data is now core to all frontier foundation models
“What this suggests about the GPT training data is that the self-reflection data has now actually become pretty much core to the training of all frontier models, because we're seeing that happen in non-instruct models across the board.”
Pratyush Maini Feb 10, 2026 ▶ 15:26 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
Prediction Not checkable as stated
Enterprises will widely adopt specialized AI pre-training in 2026 and 2027
“So I think like, 26 and 27 are going to be the years where different enterprises start doing specialized pre-training, because the cost of pre-training really amortizes itself very fast.”
Pratyush Maini Feb 10, 2026 ▶ 19:27 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
Assertion Supported
Deng: Models internally represent uncertainty preceding hallucinatory behavior
“We've seen that models internally have some awareness of like uncertainty or some sort of like user pleasing behavior that leads to hallucinatory behavior.”
Myra Deng Feb 5, 2026 ▶ 27:50 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Prediction Not checkable as stated
Deng: Scaling alone will not achieve AI needed for mission-critical deployments
“Scale is not going to get us to the type of AI development that we want to be at in, in the future as these models get more powerful and get deployed and all these sorts of like mission critical contexts.”
Myra Deng Feb 5, 2026 ▶ 44:15 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.