The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 1,786 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 41 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Not checkable as stated
Yegge: OpenAI sees a 10x productivity gap between AI adopters and non-adopters
“Anecdotally, they're sharing that performance, the performance differences like 10 X By any way that you measure it. So lines of code, commits, business impact, whatever. And it's so stark and pronounced that the people who aren't adopting it are now 10 times …”
Steve Yegge Dec 26, 2025 ▶ 3:02 Steve Yegge's Vibe Coding Manifesto: Why Claude Code Isn't It & What Comes After the IDE
Assertion Not checkable as stated
Abraham: CloudChef robot outperforms expert chefs on 40-50% of commercial cuisines
“Our robot is able to do line cooking for about 40 to 50% of the world's commercially valuable cuisine to a point that if we put our robot against an expert chef in that cuisine, our robot is able to consistently make the food better than even the chef whose so…”
Nikhil Abraham May 31, 2025 ▶ 4:27 [AIEWF Preview] CloudChef: Your Robot Chef - Michellin-Star food at $12/hr (w/ Kitchen tour!)
Assertion Not publicly verifiable
Hotz: GPT-4 is an 8-way mixture model with 220B parameters per head
“GPT-IV is two hundred twenty billion in each head, and then it's an eight-way mixture model.”
George Hotz Jun 20, 2023 ▶ 49:48 Ep 18: Petaflops to the People — with George Hotz of tinycorp
Assertion Not checkable as stated
Sean Lie: 95% of High-Quality Open Models Come From Chinese Labs
“The open source model market is a hundred percent Chinese, right? . A hundred percent, but. Almost. Okay. 95%, right? Most of the big models, most of the big open models that are, you know, high quality are coming from the Chinese labs.”
Sean Lie Sep 2, 2026 ▶ 41:22 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Assertion Not checkable as stated
Google DeepMind's six-month internal embargo shelves commercially valuable research papers
“I mean, what's worse is the paper is actually not even being Published anymore because there's a six month embargo inside of DeepMind, right? Like we've heard about this where a paper comes out and then I think there's a six month embargo window where if anybo…”
Anjney Midha Jun 18, 2026 ▶ 15:16 Why AI Labs With Unlimited GPUs Still Fail — Anjney Midha, AMP
Assertion Not checkable as stated
Awais: Claude Code hides 50+ tool-call failures per session on DeepSeek
“In CloudCode you know, they hide a lot of the errors behind control O, right? So you don't even know that, you know, you have like 50 plus tool call failures plus per session. You're just sitting there and you're like, oh, why is DeepSeq so slow?”
Ahmad Awais Jun 6, 2026 ▶ 12:21 ⚡️Making DeepSeek v4 outperform Opus 4.7 with Taste — @AhmadAwais , CommandCode.ai
Assertion Partly supported
Petersson: Anthropic's Claude models uniquely exhibit emergent deceptive and cartel behaviors
“So every single model from Anthropic since have been going in this direction. And I think one interesting thing is that like, OpenAI models don't. They, Quite plainly, they don't, they behave really well. And you know, you don't know if this is like, good, lik…”
Lukas Petersson Jun 4, 2026 ▶ 46:27 When AI Agents Run Businesses — Lukas Petersson and Axel Backlund of Andon Labs
Assertion Not checkable as stated
Hong: Competitor's AI demo can be solved entirely by Lean's grind tactic
“We're talking about, for example, the grind tactic in Lean. It can currently handle a lot of mass proofs, like, at a very low level. And this is pretty shocking because I have seen, you know, actually another company working in the same space, like, you know, …”
Carina Hong Jun 3, 2026 ▶ 10:27 Scaling Past Informal AI - Carina Hong, Axiom Math
Assertion Supported
Azhnyuk: FPV Drones Cause 70% to 80% of Frontline Casualties
“Out of all the casualties on the frontline, between 70 and 80% are done by FPV drones.”
Yaroslav Azhnyuk May 18, 2026 ▶ 25:18 FPV Drones -The Next War Is Already Here — Yaroslav Azhnyuk, The Fourth Law & Noah Smith, Noahpinion
Assertion Not checkable as stated
GPT-5 Reproduced Lupsasca's Best Physics Paper in 30 Minutes
“Then when GPT-V came out. It was able to reproduce one of my best papers that took me a very long time to come up with, in like, 30 minutes.”
Alex Lupsasca May 5, 2026 ▶ 2:32 🔬How GPT‑5 derived new results in theoretical physics and quantum gravity — Alex Lupsasca, OpenAI
Assertion Not checkable as stated
AI Resolved Theoretical Physics Problem That Puzzled Experts for a Year
“AI has become superhuman, at least on certain tasks. And that's what led to these recent papers which maybe we should talk about that resolve a problem that was puzzling physicists for experts in the field for over a year, and they weren't able to resolve it a…”
Alex Lupsasca May 5, 2026 ▶ 6:01 🔬How GPT‑5 derived new results in theoretical physics and quantum gravity — Alex Lupsasca, OpenAI
Assertion Not checkable as stated
ChatGPT Solved Open Physics Problem Before Collaborator's Flight Landed
“We decided to start working on it using AI a little bit before Andy was scheduled to come, like the week before. And in fact, using ChatGPT, we solved the problem before he even got off the plane.”
Alex Lupsasca May 5, 2026 ▶ 20:53 🔬How GPT‑5 derived new results in theoretical physics and quantum gravity — Alex Lupsasca, OpenAI
Assertion Not checkable as stated
Internal OpenAI Model Proved Gluon Amplitude Formula in 12 Hours
“We had this Internal model that could think for a very long time and was extra strong in physics. So we gave it the whole problem from scratch without actually giving it this. We just formulated the problem in a very sharp way and asked the model to solve, to …”
Alex Lupsasca May 5, 2026 ▶ 35:56 🔬How GPT‑5 derived new results in theoretical physics and quantum gravity — Alex Lupsasca, OpenAI
Assertion Not checkable as stated
Parakhin: Top AI models write code with fewer bugs than average humans
“I would claim by now, good model writes code on average with fewer bugs than average human.”
Mikhail Parakhin Apr 22, 2026 ▶ 12:19 AI-Native Engineering: 100% adoption, 5x search throughput, unlimited tokens — Mikhail Parakhin
Assertion Supported
Sachs: AI Model Quality Varies Between First-Party APIs and Cloud Providers
“Companies that say they're selling the same model through different vendors, whether it be through first party or Bedrock, Azure, et cetera, we do see different qualities sometimes, and that's not necessarily what's advertised.”
Sarah Sachs Apr 15, 2026 ▶ 24:37 Notion’s Sarah Sachs & Simon Last on Custom Agents, Evals, and the Future of Work
Assertion Contradicted
Andreessen: Three-year-old Nvidia chips make more money today than when new
“The current models are getting better faster at such a rate that if you are running an NVIDIA, if you're running an NVIDIA inference chip today that's three years old, you're making more money on it today than you did three years ago. Because the pace of impro…”
Marc Andreessen Apr 3, 2026 ▶ 23:23 Marc Andreessen introspects on Death of the Browser, Pi + OpenClaw, and Why "This Time Is Different"
Assertion Not checkable as stated
Unnamed materials foundation model is only 5x faster than DFT and unreliable
“It's only in my hands the one I'm still not naming is only about five times faster than my fastest DFT calculation on a GPU, and it also doesn't work all the time.”
Heather Kulik Mar 24, 2026 ▶ 20:32 🔬There Is No AlphaFold for Materials — AI for Materials Discovery with Heather Kulik
Assertion Not checkable as stated
Patel: Anthropic added $2B in monthly revenue at positive margins
“Anthropic doesn't just add two billion dollars of revenue in one month. You know, with, without having, you know, huge demand and they're doing it at positive margins, right?”
Dylan Patel Feb 26, 2026 ▶ 25:21 Dylan Patel Explains the AI War While Cooking | In-Context Cooking
Assertion Supported
Bissell: CCP bias is identifiable in Qwen and DeepSeek-R1 representation spaces
“Well, there's, there are certainly internal, yeah, parts of the representation space where you can sort of see where that lives.”
Mark Bissell Feb 5, 2026 ▶ 10:08 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Assertion Contradicted
Hill-Smith: Google used unpublished 32-shot CoT to claim Gemini beat GPT-4
“Back when I'm Googled a Gemini one when I ultra and needed a number that would say it was better than GPT four. And Like, constructed I think never published, like, chain of thought examples, 32 of them in every topic in MLU to run it, to get the score.”
Micah Hill-Smith Jan 9, 2026 ▶ 8:36 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Assertion Not checkable as stated
Sands: Top 100 AI startups on Stripe have unprecedented revenue per employee
“When you look at most of the top hundred AI companies on Stripe, their revenue per employee is Unlike any other business, including public companies who are known for being incredibly efficient companies.”
Emily Glassberg Sands Oct 30, 2025 ▶ 1:26:26 The Agents Economy Backbone - with Emily Glassberg Sands, Head of Data & AI at Stripe
Assertion Not checkable as stated
Corbitt: Prompt optimization methods like JEPA failed OpenPipe's agent benchmarks
“It didn't work on the problems we tried it on. It just didn't. It got like a minor boost over the sort of like more naive prompt we had and was just like, it was like, okay, Just kind of like our naive prompt with our model gets maybe like 50% on this benchmar…”
Kyle Corbitt Oct 16, 2025 ▶ 37:08 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
Assertion Contradicted
Feldman: Cerebras is 20 times faster than Nvidia B200 GPUs
“Really focused on performance, both for training and for inference. You think 20 times faster than Nvidia B 200 GPUs and it's been an amazing run.”
Andrew Feldman Oct 1, 2025 ▶ 3:05 ⚡️Raising $1.1b to build the fastest LLM Chips on Earth — Andrew Feldman, Cerebras
Assertion Contradicted
Bachman: Models claiming 256k+ context use windowed transformers, discarding data
“Anybody who says they're using a transformer With a context length of, you know, 256,000 or more, they're not using a true transformer. What they're using is a windowed transformer that essentially throws out a huge amount of its information at various layers …”
Diego Bachman Sep 23, 2025 ▶ 2:58 ⚡️ Beyond Transformers with Power Retention
Assertion Not checkable as stated
Zhang: Meta Failed at Training MoE Models for Llama Series
“The reason why Lama open-sourced the MOE model, because I think they tried to train our MOE model, but they failed. So that, that's why they didn't open source MOE mode for Lama series.”
Yining Zhang Jan 19, 2025 ▶ 14:53 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Assertion Supported
Bryk: Perplexity and ChatGPT Search rely on legacy Google and Bing APIs
“So these systems, there are a few of them now they basically rely on like traditional search engines like Google or Bing, and then they combine them with like LLMs at the end to, you know, output some power graphics answering your question. So they, Like, Sear…”
Will Bryk Jan 10, 2025 ▶ 21:16 Beating Google at Search with Neural PageRank and $5M of H200s — with Will Bryk of Exa.ai
Assertion Supported
Ben Allal: Recent web dumps improve model benchmarks despite synthetic data
“So what we did is we trained different models on these different dumps, and we then computed their performance on popular like NLP benchmarks, and then we computed the aggregated score. And surprisingly, you can see that the latest dumps are actually even bett…”
Loubna Ben Allal Dec 24, 2024 ▶ 4:12 Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
Assertion Not checkable as stated
Fu: Embedding model quality barely matters for final RAG performance
“We had this experience over and over again where you could have any, an embedding model of any quality, so you could have a really, really bad embedding model, or you could have a really, really good one by, and by any measure of good, and for the final RAG ap…”
Dan Fu Dec 24, 2024 ▶ 33:00 2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
Assertion Not checkable as stated
Crivello: Lindy AI agents perform better than humans for most use cases
“I think the bar is it needs to be better than a human. And for most use cases we serve today, it is better than the human, especially if you put it on rails.”
Florent Crivello Nov 15, 2024 ▶ 30:22 Agents @ Work: Lindy.ai (with live demo!)
Assertion Not checkable as stated
Crivello: Every major remote success story lost to an in-person competitor
“In every one of these examples, you have a co-located counterfactual that is sometimes orders of magnitude bigger.”
Florent Crivello Nov 15, 2024 ▶ 49:48 Agents @ Work: Lindy.ai (with live demo!)
Assertion Supported
Joscha Bach: Only a Tiny Fraction of Wikimedia's Budget Goes to Servers
“The Wikimedia Foundation is publishing what they are paying the money for, and a very tiny fraction on this goes into running the servers, and the editors are working for free.”
Joscha Bach Apr 27, 2024 ▶ 1:53:10 This World Does Not Exist — Joscha Bach, Karan Malhotra, Rob Haisfield (WorldSim, WebSim, Liquid AI)
Assertion Not checkable as stated
Royzen: GPT-4 was trained on HumanEval, proving data contamination
“GPT-IV itself has been trained on human eval, and we know this because GPT-IV is able to predict the exact doc string in many of the problems. I've seen it predict, like, the specific example values in the doc string, which is extremely improbable for it to ju…”
Michael Royzen Nov 3, 2023 ▶ 41:31 Beating GPT-4 with Open Source Models - with Michael Royzen of Phind
Assertion Not publicly verifiable
Howard: Alec Radford Built OpenAI's GPT After Reading ULMFiT
“I organized a chat for both of us with Kate Metz in the New York Times, and Kate Metz answered, sorry, and Alec answered this question for Kate, and Kate just like, so how did, you know, GPT come about? And he said, well, I was pretty sure that pre-training on…”
Jeremy Howard Oct 20, 2023 ▶ 15:41 The End of Finetuning — with Jeremy Howard of Fast.ai
Assertion Not checkable as stated
Howard: JAX was a grassroots Google reaction against TensorFlow 2
“But I mean, in the meantime, I will say, you know, Google now does have a backup plan. You know, they have JAX, which was never a strategy. It was just a bunch of people who also recognized TensorFlow two as shit, and they just decided to build something else.”
Jeremy Howard Oct 20, 2023 ▶ 1:01:58 The End of Finetuning — with Jeremy Howard of Fast.ai
Assertion Not checkable as stated
Jenik: Accelerated Understanding achieved 5-trillion token inference context length
“Like we're able to train up to a trillion context input. We're able to train With like, even inputs, outputs, both trillion context length, we are able to do inference at five trillion contexts.”
Benedikt Jenik Sep 4, 2026 ▶ 6:25 Faster Chips That Don't Melt — Anima Anandkumar & Benedikt Jenik, Accelerated Understanding
Assertion Open · timeframe Sep 2029
Anandkumar: Multi-physics models outperform single-physics models of equivalent parameter size
“And in fact, I was going to add that it turns out that having the model of the same size with multiple areas of physics does better than giving all of those parameters to each single physics. So if you had separate models and made them big enough as the origin…”
Anima Anandkumar Sep 4, 2026 ▶ 8:08 Faster Chips That Don't Melt — Anima Anandkumar & Benedikt Jenik, Accelerated Understanding
Assertion Supported
Lie: Cerebras runs OpenAI's flagship model 14x faster than GPUs
“We're running you know, frontier level, one of the most intelligent models, right? OpenAI's largest, most capable, most intelligent model right now at 14 times faster than their normal, you know, GPU speeds.”
Sean Lie Sep 2, 2026 ▶ 8:37 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Assertion Supported
Lie: Cerebras chips have 100x more memory than Groq LPUs
“One of our chips has, You know, order a hundred times more memory than one of their chips, right? So you got two orders of magnitude difference in scale kind of for free, right?”
Sean Lie Sep 2, 2026 ▶ 23:58 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Assertion Not checkable as stated
Sean Lie: Cerebras has solved yield at scale and 3D packaging
“And so, you know, we already have solved yield at scale. For example, we've already solved how you can actually package in a three-dimensional way. And that's, You know, problems that Samsung, that D-Matrix, and everybody else are also gonna have to solve over…”
Sean Lie Sep 2, 2026 ▶ 40:25 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Assertion Contradicted
Neural operators are the only AI architecture that works for climate emulation
“This is where the Allen AI Institute has now built climate models based on our neural operator architecture. And that's the only one that works As an AI emulator, right? None of the other architectures work for climate because climate requires us to assume the…”
Anima Anandkumar Aug 26, 2026 ▶ 37:44 🔬 Why Transformers Hit a Wall the Moment Physics Shows Up — Anima Anandkumar, Caltech
Assertion Supported
AI models predict fusion reactor plasma disruption one million times faster
“You know, I talk about plasma and fusion reactor. You know, we barely have a few thousand samples, but we are able to accurately predict events like disruption very well. And we are able to do that a million times faster than what traditional simulations were …”
Anima Anandkumar Aug 26, 2026 ▶ 43:50 🔬 Why Transformers Hit a Wall the Moment Physics Shows Up — Anima Anandkumar, Caltech
Assertion Supported
Park: Generative agent digital twins replicate human behavior at 85% accuracy
“And this is where we basically could replicate people's behaviors and attitudes, 85% as accurately as people would replicate their own. So that actually was the first really paper that gave this validated results that we can actually model individuals in an ac…”
Joon Sung Park Aug 21, 2026 ▶ 29:20 Simulating Humanity: from Generative Agents to 8 Billion Digital Twins — Joon Sung Park, Simile AI
Assertion Not checkable as stated
Park: Frontier models hit only 20-30% accuracy predicting niche human behavior
“Where in some cases, the model performance of frontier models go all the way down to 20, 30%. Especially if you go into that more niche population on topics that our customers will actually care about. On more gen pop, it might be around 50 to 60%.”
Joon Sung Park Aug 21, 2026 ▶ 31:17 Simulating Humanity: from Generative Agents to 8 Billion Digital Twins — Joon Sung Park, Simile AI
Assertion Not checkable as stated
McPartlon: AI models are nearing direct output of viable drug molecules
“We're kind of at the inflection point now. We're really seeing this internally at CHI, where the models are getting pretty close to, like, producing Molecules that could eventually, or like are very close to drugs.”
Matt McPartlon Aug 11, 2026 ▶ 48:31 🔬They Thought the Model Was Broken — Matt McPartlon & Neil Patil, Chai Discovery
Assertion Not checkable as stated
Companies developing fused mega kernels rarely run them in production
“Even the companies that have worked at or people that have spoken to who work at companies that do fused mega kernels, they, Very, very often don't end up running those in production because the TRTL and modular kernels that will launch are faster because you …”
Ali Taha Aug 3, 2026 ▶ 58:38 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Assertion Supported
Kant: Major AI labs did not prioritize RL for LLMs three years ago
“And the second was that reinforcement learning was going to be the biggest driver for LLM capabilities. Today, very obvious three years ago was not an opinion held or direction held at either OpenAI or Google or Anthropic or others.”
Eiso Kant Jul 22, 2026 ▶ 6:03 The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
Assertion Supported
Kant: Laguna S outperforms models two to three times its size
“When you look at the benchmarks and start using it, you'll realize that we are outperforming models two or three times their size.”
Eiso Kant Jul 22, 2026 ▶ 57:50 The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
Assertion Not checkable as stated
Wang: X-Cell Is First to Predict Unseen Cell Line Perturbations
“One of the rewarding signals I receive after we develop Excel is that like it's a wow moment from biologists that this is the first time biologists actually find the model can predict exactly how these unseen cell lines kind of respond to different perturbatio…”
Bo Wang Jul 21, 2026 ▶ 7:54 🔬Causal Models Need Causal Data - Xaira’s X-Cell model (Bo Wang & Ci Chu)
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.