The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 1,046 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Supported
FourCastNet matches supercomputer weather accuracy 10,000 times faster on consumer GPUs
“To our surprise, we found that it's not only, you know, accurate, it's almost as close to the what the traditional weather models can do accurately, but also tens of thousands of times faster. So what would take a big supercomputer to run can now be run. And w…”
Anima Anandkumar Aug 26, 2026 ▶ 33:56 🔬 Why Transformers Hit a Wall the Moment Physics Shows Up — Anima Anandkumar, Caltech
Assertion Supported
FourCastNet predicted Hurricane Lee's landfall days earlier than standard weather models
“For instance, our forecast net was able to correctly predict that the hurricane making the landfall several days earlier compared to the standard weather forecasting models.”
Anima Anandkumar Aug 26, 2026 ▶ 52:28 🔬 Why Transformers Hit a Wall the Moment Physics Shows Up — Anima Anandkumar, Caltech
Assertion Supported
Spherical AI weather models achieve longer autoregressive rollouts than flat models
“These models that we have are able to do the longest rollouts compared to any of the other weather models that completely ignore spherical assumption and a range of other things.”
Anima Anandkumar Aug 26, 2026 ▶ 57:12 🔬 Why Transformers Hit a Wall the Moment Physics Shows Up — Anima Anandkumar, Caltech
Assertion Supported
FourCastNet models trained on six-hour steps produce stable months-long forecasts
“To predict for the next six hours. And a little bit of multi-step fine tuning... Now we are showing for several months that it's able to do that.”
Anima Anandkumar Aug 26, 2026 ▶ 1:00:15 🔬 Why Transformers Hit a Wall the Moment Physics Shows Up — Anima Anandkumar, Caltech
Assertion Supported
Park: Training on randomized controlled trials improves AI human-behavior prediction
“That by collecting a lot of these randomized control trials that are really well designed, we can make significant improvement in models capability to predict human behaviors.”
Joon Sung Park Aug 21, 2026 ▶ 34:21 Simulating Humanity: from Generative Agents to 8 Billion Digital Twins — Joon Sung Park, Simile AI
Assertion Supported
Krentsel: OpenClaw, Pi, and Claude Code hardcode static agent policies
“These are all policy decisions that are static, that are defined for OpenClaw, or for Pi, or for Clawed code, if you look at their Source code. And so that is the kind of, that is the policy of what an agent is, the tools it can use, the skills it has, how it …”
Alex Krentsel Aug 15, 2026 ▶ 7:43 Exo: Harnesses should see their own code and logs — Alex Krentsel, UC Berekeley / Google Research
Assertion Supported
McPartlon: Chai-2 generated binders for 25 targets with 20% hit rate
“So we designed antibodies to 50 targets for that paper, got binders to about half of them with being on average around a 20% hit rate for binding.”
Matt McPartlon Aug 11, 2026 ▶ 22:51 🔬They Thought the Model Was Broken — Matt McPartlon & Neil Patil, Chai Discovery
Assertion Supported
McPartlon: Chai-2 achieved 0.33 angstrom error in cryo-EM testing
“And in this case, it was a 0.33 angstrom error, which is one third the width of an atom.”
Matt McPartlon Aug 11, 2026 ▶ 34:45 🔬They Thought the Model Was Broken — Matt McPartlon & Neil Patil, Chai Discovery
Assertion Supported
McPartlon: AlphaFold-Multimer gets antibody-antigen prediction right only 11% of the time
“Not really like outfold to got like, I think, 11%, the multiple version of this got like 11% of antibody antigen prediction cases. Correct. That means 90% of the time it's wrong.”
Matt McPartlon Aug 11, 2026 ▶ 53:57 🔬They Thought the Model Was Broken — Matt McPartlon & Neil Patil, Chai Discovery
Assertion Supported
Patil: Chai-2 demonstrated precise antibody-based GPCR agonist activity
“For example, in CHI-II, we showed like GPCR agonist activity, right? Where you can really hit the switch on a, you know, on a cell doorbell protein, so to speak, right? In a very precise way. Very, very, very hard to do that with antibodies if you can't be tha…”
Neil Patil Aug 11, 2026 ▶ 57:05 🔬They Thought the Model Was Broken — Matt McPartlon & Neil Patil, Chai Discovery
Assertion Supported
Kant: Poolside's 8B-active Laguna S solved Erdős 397 on DGX Spark
“A 118,000,000,008 B active model, which is not that large. It fits on a DGX spark and still runs at, you know, 3040 tokens a second on a spark is able to solve. Erdos three 97 independently. It's able to do complex programming tasks.”
Eiso Kant Jul 22, 2026 ▶ 39:09 The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
Assertion Supported
Chu: X-Cell Predicts Perturbation Effects In Unseen Activated T Cells
“Critically, Excel has not seen active cell T cells. And it's able to make accurate prediction, not only on the known biology, the TCR complex, predicting their effect accurately, that these are going to inactive the T cells, which It's exactly what we would ex…”
Ci Chu Jul 21, 2026 ▶ 57:35 🔬Causal Models Need Causal Data - Xaira’s X-Cell model (Bo Wang & Ci Chu)
Assertion Supported
Chu: Xaira is first to combine seven genome-wide Perturb-seq campaigns
“It is the first time that someone can put together not just one perturbseek, but seven genome-wide perturbseek campaigns together.”
Ci Chu Jul 21, 2026 ▶ 1:05:59 🔬Causal Models Need Causal Data - Xaira’s X-Cell model (Bo Wang & Ci Chu)
Assertion Supported
Bubna: Open-source DFlash matches proprietary speculative decoding performance
“Recently we shared our work on dflash, which is a block-based speculator, and we've open sourced all of it, so you can get, by using open source dflash, you can get the same performance as you would with one of the proprietary providers.”
Akshat Bubna Jul 8, 2026 ▶ 17:13 The Future of AI Infra: from Kubernetes to Agent Sandboxes — Akshat Bubna, Modal CTO
Assertion Supported
Cohen: OpenClaw logs all messages in plain text
“I started to see the size of the code base and the number of dependencies and some other things like logging all messages in plain text that just made me a bit apprehensive to use it for like production use cases to build a business on it.”
Gavriel Cohen Jun 29, 2026 ▶ 7:19 The Blueprint for Autonomous Work Agents | Gavriel Cohen, NanoClaw
Assertion Supported
Xin: Every major analytics database engine in traction is a decade old
“Actually, every single database engine out there, especially on the analytics side, are kind of a decade old. Pretty much everything that had reasonable traction are about a decade old.”
Reynold Xin Jun 24, 2026 ▶ 45:08 The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin
Assertion Supported
DeepMind, Microsoft, and Meta are building or using physical science labs
“You see people like Google DeepMind, Microsoft, other places like Meta, either building their own lab or running experiments at someone else's lab to get that data back.”
Joseph Krause Jun 17, 2026 ▶ 22:14 🔬 The Limits of AI in Science - Why We Need Self-Driving Labs — Joseph Krause, Radical AI
Assertion Supported
Andon Labs AI agent Luna published job listings and hired human employees
“So it has two, two people that it hired. It did job listings.”
Axel Backlund Jun 4, 2026 ▶ 1:07:04 When AI Agents Run Businesses — Lukas Petersson and Axel Backlund of Andon Labs
Assertion Contradicted
Hong: DeepSeek dissolved its formal reasoning team over strategic shift
“And we have since, for example, Deep Seek All right. Like originally having a formal team and then later dissolve that team because of strategic direction change.”
Carina Hong Jun 3, 2026 ▶ 1:14:16 Scaling Past Informal AI - Carina Hong, Axiom Math
Assertion Supported
Hong: Harmonic's Aristotle Verified an Erdős Problem Proof Found by GPT
“In fact, like, you know, GPT found a proof to an unsolved Erdos problem, and our competitor Harmonic, you know, Aristotle you know, verified it.”
Carina Hong Jun 3, 2026 ▶ 25:33 Scaling Past Informal AI - Carina Hong, Axiom Math
Assertion Supported
Hong: All OpenAI formal math researchers have left the company
“No, no, they all left.”
Carina Hong Jun 3, 2026 ▶ 1:05:45 Scaling Past Informal AI - Carina Hong, Axiom Math
Assertion Supported
Ethan He: Storing and moving video datasets costs millions per month
“So, so it's like just storing, storing the network, those costs, it's just I guess it would be a few millions per month to just storing everything, not to mention the GPU costs.”
Ethan He Jun 1, 2026 ▶ 35:49 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Assertion Supported
Sparse autoencoders found a single learned feature for nucleophilic elbows in ESMC
“You know, what we found basically is that the model has a kind of a single feature for this nucleophilic elbow and is activating across these like very evolutionarily diverse families, you know, really completely different structural topologies, proteins that …”
Alex Rives May 27, 2026 ▶ 12:59 🔬 The Bitter Lesson is Coming for Proteins - Alex Rives, BioHub
Assertion Supported
Incorporating metagenomic training data eliminated diminishing returns in ESMC protein foundation models
“And then, you know, what we saw basically is, is, is there are no longer diminishing returns to scale. So that's really saying that ESM two was kind of data limited rather than compute limited for ESMC.”
Alex Rives May 27, 2026 ▶ 24:23 🔬 The Bitter Lesson is Coming for Proteins - Alex Rives, BioHub
Assertion Supported
Rives: ESMC Has Been Used to Successfully Design scFv Antibodies
“We've been able to use this to actually now go and design many protein binders, but I think sort of most excitingly, we've been able to use this to actually design antibodies, SCFVs, and we're seeing really, I think, exciting success rates and a small number o…”
Alex Rives May 27, 2026 ▶ 28:26 🔬 The Bitter Lesson is Coming for Proteins - Alex Rives, BioHub
Assertion Supported
Burazin: Daytona built a Windows sandbox that spins up in one second
“And the only option right now is an EC two with Windows or on, on, on Azure. Both of them take anywhere from three to five minutes to spin up. We've created an actual sandbox. So it's a second instead of milliseconds, but you have like point in time snapshots,…”
Ivan Burazin May 21, 2026 ▶ 35:46 AI Agents Need Computers: 74% MoM Growth, 850K/Day Runs, & New Agent Cloud — Ivan Burazin, Daytona
Assertion Partly supported
Azhnyuk: Standard 10-Inch Ukrainian FPV Drones Strike at 30-40 Kilometers
“Many hits are happening between 30 and 40 kilometers, and that's what expected from a regular 10 inch, ah, FPV drone.”
Yaroslav Azhnyuk May 18, 2026 ▶ 22:10 FPV Drones -The Next War Is Already Here — Yaroslav Azhnyuk, The Fourth Law & Noah Smith, Noahpinion
Assertion Supported
Azhnyuk: Ukraine made 4M FPV drones last year with 7M targeted this year
“Last year alone, Ukraine manufactured about four million of these, and then Russia's maybe, like, 20% less than that. And for this year, the publicly voiced target was seven million on the Ukrainian side.”
Yaroslav Azhnyuk May 18, 2026 ▶ 39:19 FPV Drones -The Next War Is Already Here — Yaroslav Azhnyuk, The Fourth Law & Noah Smith, Noahpinion
Assertion Supported
Azhnyuk: Inexpensive local hardware now runs near-SOTA open-source AI models
“You can have basically an inexpensive computer running, you know, what was a state of the art model year and a half ago, running it locally on a device with an open source model, which also means that the Chinese can have it, the Russians can have it, the Nort…”
Yaroslav Azhnyuk May 18, 2026 ▶ 56:23 FPV Drones -The Next War Is Already Here — Yaroslav Azhnyuk, The Fourth Law & Noah Smith, Noahpinion
Assertion Supported
Smith: Success rate of large powers invading smaller states has dropped
“They fail a lot more now than they used to. The success rate of taking places over has gone way down.”
Noah Smith May 18, 2026 ▶ 1:55:41 FPV Drones -The Next War Is Already Here — Yaroslav Azhnyuk, The Fourth Law & Noah Smith, Noahpinion
Assertion Partly supported
Applied Intuition operates driverless L4 autonomous trucks in Japan
“We run like, as an example, we run driverless trucks in Japan right now, like as we speak, we can't have errors that are L four trucks.”
Qasar Younis Apr 27, 2026 ▶ 4:17 The $15B Physical AI Company: Simulation, Autonomy OS, Neural Sim, & 1K Engineers—Applied Intuition
Assertion Supported
Swix: YouTube's recommendation system is LLM-based using tokenized video codebooks
“The YouTube Rexxus is LLM-based, and they- Is it really? They obviously, yeah. That's cool. It, they re, they tokenize every video, and put it in a code book, and then they train a LLM on it, and then feed in your context, just like a regular LLM, and ask it t…”
Shawn Wang Apr 18, 2026 ▶ 22:14 ⚡️ How to turn Documents into Knowledge: Graphs in Modern AI — Emil Eifrem, CEO Neo4J
Assertion Contradicted
D'Amico: Decoupling appliances from the grid via batteries is unprecedented
“If you slam a battery into it, you're now not, you, you've decoupled the energy input from the wall with the device's power outputs. You can decouple the user experience from the grid. And that level of, like, approach has not been done in kind of the major ho…”
Sam D'Amico Mar 31, 2026 ▶ 8:56 The Stove Guy: Sam D'Amico Shows New AI Cooking Features on America's Most Powerful Stove at Impulse
Assertion Supported
D'Amico: Impulse cooktop boils a liter of water in 40s with 10,000W
“So we're putting 10,000 watts of power into a, like, medium-sized pan. Or pot. And, ah, we can boil about a liter of cold water in about 40 seconds.”
Sam D'Amico Mar 31, 2026 ▶ 0:40 The Stove Guy: Sam D'Amico Shows New AI Cooking Features on America's Most Powerful Stove at Impulse
Assertion Supported
D'Amico: Impulse stoves add a 3kWh battery for 15,000W boost power
“So basically we put a three kilowatt hour LFP battery, lithium iron phosphate inside home appliances, starting with stoves that gives us an additional 15,000 Watts of additional power. So that's like double what a typical induction cooktop has in terms of aggr…”
Sam D'Amico Mar 31, 2026 ▶ 3:45 The Stove Guy: Sam D'Amico Shows New AI Cooking Features on America's Most Powerful Stove at Impulse
Assertion Supported
Materials Project and Open Catalyst datasets rely on low-fidelity DFT calculations
“Materials project open catalyst project, these do provide good leaderboards, but some of the limitations that are the data comes from not very high fidelity density functional theory. So I'd say that's a second challenge is that we're all, all the smartest ML …”
Heather Kulik Mar 24, 2026 ▶ 18:19 🔬There Is No AlphaFold for Materials — AI for Materials Discovery with Heather Kulik
Assertion Partly supported
Colvin: Deno cannot control memory usage and prevents OOM protection
“Even if you don't allow that Dino does not have any way of controlling memory. So even if someone can't run arbitrary code, they can oom your machine as often as they like.”
Samuel Colvin Mar 14, 2026 ▶ 8:11 ⚡️Monty: the ultrafast Python interpreter by Agents for Agents — Samuel Colvin, Pydantic
Assertion Supported
Colvin: Monty executes Python in single-digit microseconds
“I mean, in a hot loop, we can run, we can go from code to execution result in under a microsecond in like 800 nanoseconds. In reality, it's like single digit microseconds to run code or single digit microseconds to run the next step of a REPL or single digit m…”
Samuel Colvin Mar 14, 2026 ▶ 5:13 ⚡️Monty: the ultrafast Python interpreter by Agents for Agents — Samuel Colvin, Pydantic
Assertion Supported
Google developers internally use a coding tool called Jet Ski
“They use an internal thing called jet ski. And I mean, the simple reason is they have internal versions of everything else. And so if they actually train the, they release the internal version to the external world, it would just make no sense. Like it would j…”
Shawn Wang Mar 14, 2026 ▶ 19:58 ⚡️Monty: the ultrafast Python interpreter by Agents for Agents — Samuel Colvin, Pydantic
Assertion Supported
Eskildsen: Turbopuffer ANN v3 searches 100B vectors with 40ms p50 latency
“ANN v. Three can search a hundred billion vectors with a p-fifty of around 40 milliseconds and a p-ninety-nine of 200 milliseconds. Maybe other people have done this. I'm sure Google and others have done this, but we haven't seen anyone at least not in like a …”
Simon Eskildsen Mar 12, 2026 ▶ 46:40 Retrieval After RAG: Hybrid Search, Agents, and Database Design — Simon Eskildsen of Turbopuffer
Assertion Supported
Eskildsen: Cursor moved 20 terabytes from Postgres to Turbopuffer
“Like at some point, cursor moved like 20 terabytes of Postgres data into Turbo Puffer, because it's like, it's there, it works, and these particular query plans we know work well, and so they just moved it all to defer sharding.”
Simon Eskildsen Mar 12, 2026 ▶ 55:44 Retrieval After RAG: Hybrid Search, Agents, and Database Design — Simon Eskildsen of Turbopuffer
Assertion Supported
Shah: Supermemory outperformed Claude Code and OpenClaw benchmarks by almost 50%
“So the Claude code one performed the worst, and OpenClaw slightly more than that, and SuperMemory is the highest, and you can see that, you know, it's like a pretty significant difference, like almost 50%.”
Dhravya Shah Mar 9, 2026 ▶ 14:33 ⚡️ OpenClaw's Memory Sucks and the fix is simple — Dhravya Shah, Supermemory
Assertion Supported
Shah: No other memory provider offers hybrid memory and raw RAG fallback
“Essentially, you will return the memories first, and then if there's any raw chunks that match up, we also return those to make sure the agent knows just enough information to answer the question. And no other provider does this right now”
Dhravya Shah Mar 9, 2026 ▶ 23:28 ⚡️ OpenClaw's Memory Sucks and the fix is simple — Dhravya Shah, Supermemory
Assertion Contradicted
Nelle: No one had enabled AI coding agents to run code before Cursor
“Like obviously you need to run the code. And so that I think also is probably not that contrarian of a take, but no one has done that yet.”
Jonas Nelle Mar 6, 2026 ▶ 1:38 Cursor's Third Era: Cloud Agents — ft. Sam Whitmore, Jonas Nelle, Cursor
Assertion Contradicted
Huber: Frontier AI models are not actually good at agentic search
“We've like sort of stress tested like frontier models and their ability to search. And they are not actually that good at searching.”
Jeff Huber Mar 5, 2026 ▶ 26:43 Why Every Agent Needs a Box — Aaron Levie, Box
Assertion Supported
Levie: Anthropic has forward-deployed engineers embedded at Goldman Sachs
“OpenAI probably is hiring FDEs to go into the enterprise and then Anthropic is embedded at Goldman Sachs.”
Aaron Levie Mar 5, 2026 ▶ 17:33 Why Every Agent Needs a Box — Aaron Levie, Box
Assertion Supported
Becker: Current Frontier Models Cannot Cause Catastrophic Harm
“We find, we think it's not capable enough, you know, on the basis of some of this capabilities evidence that you've alluded to commit these catastrophic harms.”
Joel Becker Feb 27, 2026 ▶ 2:34 Measuring Exponential Trends Rising (in AI) — Joel Becker, METR
Assertion Partly supported
Patel: Meta is taking on $40B in debt for its Louisiana AI cluster
“Meta, they've, they're already taking debt on for their largest AI cluster in Louisiana. You know, they're taking like forty billion dollars of debt on for that”
Dylan Patel Feb 26, 2026 ▶ 33:44 Dylan Patel Explains the AI War While Cooking | In-Context Cooking
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.