The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 851 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Supported
Azhnyuk: FPV Drones Cause 70% to 80% of Frontline Casualties
“Out of all the casualties on the frontline, between 70 and 80% are done by FPV drones.”
Yaroslav Azhnyuk May 18, 2026 ▶ 25:18 FPV Drones -The Next War Is Already Here — Yaroslav Azhnyuk, The Fourth Law & Noah Smith, Noahpinion
Assertion Supported
Sachs: AI Model Quality Varies Between First-Party APIs and Cloud Providers
“Companies that say they're selling the same model through different vendors, whether it be through first party or Bedrock, Azure, et cetera, we do see different qualities sometimes, and that's not necessarily what's advertised.”
Sarah Sachs Apr 15, 2026 ▶ 24:37 Notion’s Sarah Sachs & Simon Last on Custom Agents, Evals, and the Future of Work
Prediction Held up
Andreessen: Autonomous AI agents will inevitably hire humans for tasks
“The agent hiring the people, which of course is going to happen, right? It's obviously going to happen.”
Marc Andreessen Apr 3, 2026 ▶ 41:32 Marc Andreessen introspects on Death of the Browser, Pi + OpenClaw, and Why "This Time Is Different"
Assertion Supported
Bissell: CCP bias is identifiable in Qwen and DeepSeek-R1 representation spaces
“Well, there's, there are certainly internal, yeah, parts of the representation space where you can sort of see where that lives.”
Mark Bissell Feb 5, 2026 ▶ 10:08 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Assertion Supported
Bryk: Perplexity and ChatGPT Search rely on legacy Google and Bing APIs
“So these systems, there are a few of them now they basically rely on like traditional search engines like Google or Bing, and then they combine them with like LLMs at the end to, you know, output some power graphics answering your question. So they, Like, Sear…”
Will Bryk Jan 10, 2025 ▶ 21:16 Beating Google at Search with Neural PageRank and $5M of H200s — with Will Bryk of Exa.ai
Assertion Supported
Ben Allal: Recent web dumps improve model benchmarks despite synthetic data
“So what we did is we trained different models on these different dumps, and we then computed their performance on popular like NLP benchmarks, and then we computed the aggregated score. And surprisingly, you can see that the latest dumps are actually even bett…”
Loubna Ben Allal Dec 24, 2024 ▶ 4:12 Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
Assertion Supported
Joscha Bach: Only a Tiny Fraction of Wikimedia's Budget Goes to Servers
“The Wikimedia Foundation is publishing what they are paying the money for, and a very tiny fraction on this goes into running the servers, and the editors are working for free.”
Joscha Bach Apr 27, 2024 ▶ 1:53:10 This World Does Not Exist — Joscha Bach, Karan Malhotra, Rob Haisfield (WorldSim, WebSim, Liquid AI)
Assertion Supported
Lie: Cerebras runs OpenAI's flagship model 14x faster than GPUs
“We're running you know, frontier level, one of the most intelligent models, right? OpenAI's largest, most capable, most intelligent model right now at 14 times faster than their normal, you know, GPU speeds.”
Sean Lie Sep 2, 2026 ▶ 8:37 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Assertion Supported
Lie: Cerebras chips have 100x more memory than Groq LPUs
“One of our chips has, You know, order a hundred times more memory than one of their chips, right? So you got two orders of magnitude difference in scale kind of for free, right?”
Sean Lie Sep 2, 2026 ▶ 23:58 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Assertion Supported
AI models predict fusion reactor plasma disruption one million times faster
“You know, I talk about plasma and fusion reactor. You know, we barely have a few thousand samples, but we are able to accurately predict events like disruption very well. And we are able to do that a million times faster than what traditional simulations were …”
Anima Anandkumar Aug 26, 2026 ▶ 43:50 🔬 Why Transformers Hit a Wall the Moment Physics Shows Up — Anima Anandkumar, Caltech
Assertion Supported
Park: Generative agent digital twins replicate human behavior at 85% accuracy
“And this is where we basically could replicate people's behaviors and attitudes, 85% as accurately as people would replicate their own. So that actually was the first really paper that gave this validated results that we can actually model individuals in an ac…”
Joon Sung Park Aug 21, 2026 ▶ 29:20 Simulating Humanity: from Generative Agents to 8 Billion Digital Twins — Joon Sung Park, Simile AI
Assertion Supported
Kant: Major AI labs did not prioritize RL for LLMs three years ago
“And the second was that reinforcement learning was going to be the biggest driver for LLM capabilities. Today, very obvious three years ago was not an opinion held or direction held at either OpenAI or Google or Anthropic or others.”
Eiso Kant Jul 22, 2026 ▶ 6:03 The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
Assertion Supported
Kant: Laguna S outperforms models two to three times its size
“When you look at the benchmarks and start using it, you'll realize that we are outperforming models two or three times their size.”
Eiso Kant Jul 22, 2026 ▶ 57:50 The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
Assertion Supported
Chu: Descriptive Models Fail to Beat Linear Baselines on Causal Biology
“Models that are trained on descriptive data do not yet outperform linear models on causal tasks, perturbational tasks, what we call counterfactual tasks.”
Ci Chu Jul 21, 2026 ▶ 21:47 🔬Causal Models Need Causal Data - Xaira’s X-Cell model (Bo Wang & Ci Chu)
Assertion Supported
Wang: Diffusion Outperforms Autoregressive Models on Unseen Cellular Tasks
“We find that switching from autoregressive training to division language models give a significant improvement over some of the harder tasks, particularly generalized to unseen tasks.”
Bo Wang Jul 21, 2026 ▶ 52:00 🔬Causal Models Need Causal Data - Xaira’s X-Cell model (Bo Wang & Ci Chu)
Assertion Supported
Wang: X-Cell Can Predict Combinatorial Gene Perturbations In Silico
“This is also why we incorporate PPI networks as the prior knowledge into our model. And although the model right now are trained on single gene perturbations, but once the model is trained, you can actually predict combinatorial perturbations just on the model…”
Bo Wang Jul 21, 2026 ▶ 1:09:55 🔬Causal Models Need Causal Data - Xaira’s X-Cell model (Bo Wang & Ci Chu)
Assertion Supported
Perszyk: AI writing suggestions subconsciously shift users to opposing arguments
“There are studies that show that people will, even below their threshold of awareness, start with one argument and then be switched to a completely different, maybe opposing argument because of accepting all of these AI suggestions.”
Danielle Perszyk Jul 11, 2026 ▶ 35:35 Why AI Agents Don't Actually Understand You — Danielle Perszyk, Amazon AGI Lab
Assertion Supported
Perszyk: AI tools boost individual output but narrow overall scientific research
“Individual scientists who are using AI tools are benefiting because they are producing more papers. They are getting more grants accepted. But science as a whole is narrowing.”
Danielle Perszyk Jul 11, 2026 ▶ 36:16 Why AI Agents Don't Actually Understand You — Danielle Perszyk, Amazon AGI Lab
Assertion Supported
Feinberg: Studies show AlphaFold structures provided no value for drug docking
“There is this, ah, a few papers that came out, one was in Cell, I think last year, which showed that for all of the claims about AlphaFold-solving drug discovery, people try to take AlphaFold-produced protein structures, use them for traditional docking, and f…”
Evan Feinberg Jun 30, 2026 ▶ 24:24 🔬 "The Most Innovative Diffusion Research Is Happening in Drug Discovery, Not Image Generation"
Assertion Supported
Xin: Transcoding database rows to Parquet speeds object storage writes with zero compromise
“And as a matter of fact, once you transcode the data compresses better. So from those services writing to, for example, S three or other data lake, like object stores, you can actually write them faster because now they are now smaller. So there's no. Overhead…”
Reynold Xin Jun 24, 2026 ▶ 36:50 The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin
Assertion Supported
Krause: AI models cannot qualify new aerospace alloys without physical experiments
“A model can't figure out your way through the qualification pipeline for a new alloy for a jet turbine. You have to do experiments to do that”
Joseph Krause Jun 17, 2026 ▶ 3:59 🔬 The Limits of AI in Science - Why We Need Self-Driving Labs — Joseph Krause, Radical AI
Assertion Supported
Backlund: Opus 4.6 reasoning traces showed it deliberately lying about customer refunds
“And like for Opus 4.6, you could see that there was a customer, a simulated customer that wanted a refund because the product was faulty. And then the model lied that it would do the refund. And we could read in the traces that it actually was weighing like, o…”
Axel Backlund Jun 4, 2026 ▶ 47:42 When AI Agents Run Businesses — Lukas Petersson and Axel Backlund of Andon Labs
Assertion Supported
Hong: Axiom Math has solved open research problems across math subfields
“We have good performance, you know, having solved open research questions and number theory, commutative algebra, algebraic geometry, some discrete math that come into Rx and probability.”
Carina Hong Jun 3, 2026 ▶ 19:14 Scaling Past Informal AI - Carina Hong, Axiom Math
Assertion Supported
Hong: Axiom and Harmonic mistakenly claimed solved Erdős problems were new
“So actually what happened was our competitor, Harmonic, decided to publicize that they have solved unsolved problems, Erdos number one two four and four 81, and then we trusted their literature review, believing that these problems are really, truly unsolved. …”
Carina Hong Jun 3, 2026 ▶ 59:19 Scaling Past Informal AI - Carina Hong, Axiom Math
Prediction Held up
Ethan He: Video Agents Will Reach Production-Grade Quality by Year-End
“I guess by the end of this year is this is going to be a big hit. So the inflection point will be there and the videos generated by video agents can get to like production great quality. So it can be presented and it can be distributed in, in ads.”
Ethan He Jun 1, 2026 ▶ 1:30:54 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Assertion Supported
ESMC model search generates novel antibodies achieving therapeutic-grade binding affinity levels
“What we're able to see is that, you know, you can search ESMC and you can actually find antibodies that are reaching the level of affinity that are, I should say, are really at the level of affinity that is needed for therapeutic function and activity.”
Alex Rives May 27, 2026 ▶ 29:46 🔬 The Bitter Lesson is Coming for Proteins - Alex Rives, BioHub
Assertion Supported
Rives: ESMC is state of the art among open models for multimer prediction
“Yeah, I mean, I think we're state of the art for open models.”
Alex Rives May 27, 2026 ▶ 36:02 🔬 The Bitter Lesson is Coming for Proteins - Alex Rives, BioHub
Assertion Supported
Sanseviero: AI labs republished model merging techniques previously created on Reddit
“Yeah, like all of the FrankenMoe stuff, like all of the Axolotl library, like all of these tools, and there were papers published by different companies and research labs one or two years later that were rediscovering what was already done by The Reddit or Dis…”
Omar Sanseviero May 24, 2026 ▶ 23:45 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Assertion Supported
ChatGPT Pro Derived All Math in Recent Quantum Gravity Paper
“It's a real solid result in quantum gravity that was done pretty much completely by an AI. With humans steering it and asking kind of the right questions, but all the math was derived by ChatGPT Pro, the public model you can access.”
Alex Lupsasca May 5, 2026 ▶ 53:28 🔬How GPT‑5 derived new results in theoretical physics and quantum gravity — Alex Lupsasca, OpenAI
Assertion Supported
Ludwig: Tesla R&D vehicles still use LiDAR in the Bay Area
“If you see, for example, a Tesla R&D vehicle, it actually has LiDAR on it to this day, right? In, in the Bay Area, we see these you'll see like Model Ys or CyberCab that have LiDARs on them just driving around.”
Peter Ludwig Apr 27, 2026 ▶ 14:58 The $15B Physical AI Company: Simulation, Autonomy OS, Neural Sim, & 1K Engineers—Applied Intuition
Assertion Supported
Sun: Synthetic data matches real-world data for multimodal model pre-training
“We were actually generating a lot of synthetic data and showing that, hey, you can actually, these synthetic data are actually as useful as real-world data when it comes to multimodal pre-training.”
Fan-yun Sun Apr 2, 2026 ▶ 2:56 Moonlake: Interactive, Multimodal World Models — with Chris Manning and Fan-yun Sun
Assertion Supported
Reddy: Voxtral speech model is much stronger than Whisper
“And I think a big people, I think there's a big rich ecosystem of people finding whisper and people want the same thing with Voxer. It's much stronger than whisper.”
Pavan Kumar Reddy Mar 30, 2026 ▶ 26:34 Mistral: Voxtral TTS, Forge, Leanstral, & Mistral 4 — w/ Pavan Kumar Reddy & Guillaume Lample
Assertion Supported
Eskildsen: Neon retrofitted Postgres for S3, while Turbopuffer built pure object-storage consensus
“I think neon neon was first to, and they're trying to retrofit it onto Postgres. And then they built this whole architecture where you have it in memory, and then you sort of like, you know, mmap back to S-III, and I think that was very novel at the time to do…”
Simon Eskildsen Mar 12, 2026 ▶ 16:38 Retrieval After RAG: Hybrid Search, Agents, and Database Design — Simon Eskildsen of Turbopuffer
Assertion Supported
Shah: OpenClaw's 15-message replay fails prompt caching and costs 10x more
“The way OpenClaw does it is it essentially sends back the last 15 messages in the conversation and it essentially uses that back and forth. And I mean, the approach itself is not ideal because you will, like, you are not doing any, like, you're not utilizing a…”
Dhravya Shah Mar 9, 2026 ▶ 21:03 ⚡️ OpenClaw's Memory Sucks and the fix is simple — Dhravya Shah, Supermemory
Prediction Held up
Nelle: Developers will spend thousands to tens of thousands monthly on agents
“I think as we think about these highly parallel kind of agents running off for a long time in their own VM system, We are already at that point where people will be spending thousands of dollars a month per, per human, and I think potentially tens of thousands…”
Jonas Nelle Mar 6, 2026 ▶ 53:50 Cursor's Third Era: Cloud Agents — ft. Sam Whitmore, Jonas Nelle, Cursor
Assertion Supported
Patel: Claude Code's share of GitHub commits doubled to 4% in January
“Just in January, it went from four percent of or two percent of commits on GitHub to four percent of GitHub commits were done by Cloud Code, right?”
Dylan Patel Feb 26, 2026 ▶ 21:09 Dylan Patel Explains the AI War While Cooking | In-Context Cooking
Prediction Held up
Patel: Google and Amazon will borrow debt to fund AI infrastructure
“Google and Amazon haven't taken on debt yet for AI infrastructure, but they will, right?”
Dylan Patel Feb 26, 2026 ▶ 33:54 Dylan Patel Explains the AI War While Cooking | In-Context Cooking
Assertion Supported
O'Laughlin: Anthropic does not train Claude agent teams with RL
“I have a controversial opinion that Claude does not do RL on the agent swarms or agent team.”
Doug O'Laughlin Feb 24, 2026 ▶ 33:33 Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
Assertion Supported
O'Laughlin: AI build-out CapEx has massively passed the internet
“We've well massively passed the internet in terms of the absolute size of the build-out. It's not even close.”
Doug O'Laughlin Feb 24, 2026 ▶ 1:05:11 Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
Assertion Supported
Watkins: Over half of SWE-bench problems investigated by OpenAI had test flaws
“In over half of the problems that were investigated in that deep dive, there was one problem or the other. I think the most common problem are, like, overly narrow tests where there's some particular implementation detail that the tests were looking for but wa…”
Olivia Watkins Feb 23, 2026 ▶ 7:26 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
Assertion Supported
Deng: Models internally represent uncertainty preceding hallucinatory behavior
“We've seen that models internally have some awareness of like uncertainty or some sort of like user pleasing behavior that leads to hallucinatory behavior.”
Myra Deng Feb 5, 2026 ▶ 27:50 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Assertion Supported
White: ML trained on experimental data beat first-principles simulations by a large margin
“Two very well-resourced groups. They both tried different ideas, and the machine learning on experimental data beat out first principles simulation by You know, a very large margin.”
Andrew White Jan 28, 2026 ▶ 45:21 🔬 From Red Teaming GPT-4 to Automating Drug Discovery: The Future of AI in Science — Andrew White
Assertion Supported
Cameron: General model intelligence does not correlate with hallucination rates
“One interesting aspect is that we've found that there's not really a, not a strong correlation between intelligence and hallucination rate. That's to say that the smarter the models are in a generalist sense isn't correlated with their ability to, when they do…”
George Cameron Jan 9, 2026 ▶ 31:28 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Assertion Supported
Cameron: Model performance correlates with total parameters, not active parameters
“We, in our benchmark, see a lot of performance correlated more with total parameters than active, and not that correlated with how sparse like the models are. Our accuracy benchmark is part of a omniscience. It's very correlated with total. It's not correlated…”
George Cameron Jan 9, 2026 ▶ 1:05:08 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Prediction Held up
Yegge: Open source models will match Gemini 3 by next summer
“From what I've heard, they, they're seven months behind, and that, that gap is gradually narrowing. The frontier models, which means OSS models will be as good as Gemini three next summer.”
Steve Yegge Dec 26, 2025 ▶ 31:46 Steve Yegge's Vibe Coding Manifesto: Why Claude Code Isn't It & What Comes After the IDE
Assertion Supported
Pliny: Anthropic added a $20k–$30k bounty but withheld jailbreak data
“That whole thing ended with no open sourcing of data, but they did add a 30,000 or 20,000 dollar bounty, which I sort of sat myself out of, let the community go for it.”
Pliny the Liberator Dec 16, 2025 ▶ 19:08 ⚡️Jailbreaking AGI: Pliny the Liberator & John V on Red Teaming, BT6, and the Future of AI Security
Assertion Supported
Sam Altman Barred Investors Who Backed Glean From Investing in OpenAI
“Sam Altman once came out and said, if you're an investor in OpenAI and one of these five companies, including Glean, we don't want you as an investor.”
Deedy Das Nov 14, 2025 ▶ 9:02 Anthropic, Glean & OpenRouter: How AI Moats Are Built with Deedy Das of Menlo Ventures
Assertion Supported
AMD MI300X outperforms Nvidia H100 on FlashAttention-2 and memory-bound workloads
“We found that it's great for flash attention to specifically, we were able to be H-one hundred. We also found that like the less time you spend in like dense compute, like the less time you spend in tensor cores specifically, or less time you spend in lower bi…”
Quentin Anthony Nov 3, 2025 ▶ 3:19 How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.