The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 1,046 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Supported
Watkins: Over half of SWE-bench problems investigated by OpenAI had test flaws
“In over half of the problems that were investigated in that deep dive, there was one problem or the other. I think the most common problem are, like, overly narrow tests where there's some particular implementation detail that the tests were looking for but wa…”
Olivia Watkins Feb 23, 2026 ▶ 7:26 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
Assertion Supported
Deng: Models internally represent uncertainty preceding hallucinatory behavior
“We've seen that models internally have some awareness of like uncertainty or some sort of like user pleasing behavior that leads to hallucinatory behavior.”
Myra Deng Feb 5, 2026 ▶ 27:50 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Assertion Supported
White: ML trained on experimental data beat first-principles simulations by a large margin
“Two very well-resourced groups. They both tried different ideas, and the machine learning on experimental data beat out first principles simulation by You know, a very large margin.”
Andrew White Jan 28, 2026 ▶ 45:21 🔬 From Red Teaming GPT-4 to Automating Drug Discovery: The Future of AI in Science — Andrew White
Assertion Supported
Cameron: General model intelligence does not correlate with hallucination rates
“One interesting aspect is that we've found that there's not really a, not a strong correlation between intelligence and hallucination rate. That's to say that the smarter the models are in a generalist sense isn't correlated with their ability to, when they do…”
George Cameron Jan 9, 2026 ▶ 31:28 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Assertion Supported
Cameron: Model performance correlates with total parameters, not active parameters
“We, in our benchmark, see a lot of performance correlated more with total parameters than active, and not that correlated with how sparse like the models are. Our accuracy benchmark is part of a omniscience. It's very correlated with total. It's not correlated…”
George Cameron Jan 9, 2026 ▶ 1:05:08 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Prediction Didn’t hold up
Nair: LLM agents will hit $1T before robotics hits $10B
“It feels like LLM agents are going to be like a trillion dollar market before robotics is maybe even like a ten billion dollar market.”
Ashvin Nair Dec 30, 2025 ▶ 3:59 [State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
Prediction Held up
Yegge: Open source models will match Gemini 3 by next summer
“From what I've heard, they, they're seven months behind, and that, that gap is gradually narrowing. The frontier models, which means OSS models will be as good as Gemini three next summer.”
Steve Yegge Dec 26, 2025 ▶ 31:46 Steve Yegge's Vibe Coding Manifesto: Why Claude Code Isn't It & What Comes After the IDE
Assertion Supported
Pliny: Anthropic added a $20k–$30k bounty but withheld jailbreak data
“That whole thing ended with no open sourcing of data, but they did add a 30,000 or 20,000 dollar bounty, which I sort of sat myself out of, let the community go for it.”
Pliny the Liberator Dec 16, 2025 ▶ 19:08 ⚡️Jailbreaking AGI: Pliny the Liberator & John V on Red Teaming, BT6, and the Future of AI Security
Assertion Contradicted
Johnson: Nvidia Blackwell offers roughly same performance per watt as Hopper
“Like, if you look at the numbers, like, even going from Hopper to Blackwell, like, the performance per watt is about the same. They mostly make the number of transistors go up, and they make the chip size go up, and they make the power usage go up. But even fr…”
Justin Johnson Nov 25, 2025 ▶ 13:01 After LLMs: Spatial Intelligence and World Models — Fei-Fei Li & Justin Johnson, World Labs
Assertion Partly supported
Anthropic Is the Fastest-Growing Software Company in History
“Anthropic is the fastest growing software company of all time. I think I can say that fairly. I'm, I haven't been disproven yet.”
Deedy Das Nov 14, 2025 ▶ 17:16 Anthropic, Glean & OpenRouter: How AI Moats Are Built with Deedy Das of Menlo Ventures
Assertion Partly supported
OpenAI Plans to Scale Compute Power Capacity to 125 Gigawatts
“For OpenAI to go from like two gigawatts of compute this year to 30 with everything they've already announced, and then there's a plan for the next 125. Like, the United States uses 300.”
Shawn Wang Nov 14, 2025 ▶ 1:15:33 Anthropic, Glean & OpenRouter: How AI Moats Are Built with Deedy Das of Menlo Ventures
Assertion Supported
Sam Altman Barred Investors Who Backed Glean From Investing in OpenAI
“Sam Altman once came out and said, if you're an investor in OpenAI and one of these five companies, including Glean, we don't want you as an investor.”
Deedy Das Nov 14, 2025 ▶ 9:02 Anthropic, Glean & OpenRouter: How AI Moats Are Built with Deedy Das of Menlo Ventures
Assertion Supported
AMD MI300X outperforms Nvidia H100 on FlashAttention-2 and memory-bound workloads
“We found that it's great for flash attention to specifically, we were able to be H-one hundred. We also found that like the less time you spend in like dense compute, like the less time you spend in tensor cores specifically, or less time you spend in lower bi…”
Quentin Anthony Nov 3, 2025 ▶ 3:19 How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
Assertion Supported
Rumbelow: Leap Labs' Discovery Engine Automates Novel Scientific Discovery via Interpretability
“Discovery Engine is an end-to-end system, takes in arbitrary scientific data set, automatically trains a bunch of neural networks on it, and then We systematically, with our interpretability methods, which is the real secret extract the patterns that have been…”
Jessica Rumbelow Nov 2, 2025 ▶ 6:16 ⚡️Automating Scientific Discovery - Jessica Rumbelow, Leap Labs
Assertion Contradicted
Swix: Every frontier lab now distills dense models into MoEs
“I think like, I think this is the pattern for every frontier lab now.”
Shawn Wang Oct 20, 2025 ▶ 36:32 ⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
Assertion Supported
Corbitt: OpenPipe Beat Frontier Models Using a Qwen 32B Judge
“One of the results we published was we used Quen 2.5 14 B as the model we're training, and as the judge we used Quen 2.5 32 B, which is, like, Not, I mean, it's fine, but it's like not a, it's much worse than any frontier model. Right. And even with that combi…”
Kyle Corbitt Oct 16, 2025 ▶ 53:18 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
Assertion Supported
Feldman: Cerebras provides 2,625x more memory bandwidth than traditional GPUs
“And we have 2625 times more memory bandwidth than the GPU does.”
Andrew Feldman Oct 1, 2025 ▶ 6:22 ⚡️Raising $1.1b to build the fastest LLM Chips on Earth — Andrew Feldman, Cerebras
Assertion Supported
Cerebras leads all Artificial Analysis inference benchmarks by a large margin
“I think also just go up and look at artificial analysis. Wherever we are, we're the fastest not by a little bit, but by a lot.”
Andrew Feldman Oct 1, 2025 ▶ 11:03 ⚡️Raising $1.1b to build the fastest LLM Chips on Earth — Andrew Feldman, Cerebras
Assertion Supported
Rajpal: Anthropic Claude models had regressions from serving architecture changes
“Anthropix kind of cloud models kind of had a regression, right? Because they changed to a new serving architecture.”
Shreya Rajpal Sep 25, 2025 ▶ 23:21 ⚡️Snowglobe: Simulations for your AI
Prediction Held up
Taskaya: Training a state-of-the-art image model costs under $1M
“Like right now, like if you look, if you want to train a Sota image model, I don't think it's going to cost more than a million dollars. It's extremely cheap. It's like a matter of data engineering effort, cleaning. It's, I think it's a function of data set.”
Batuhan Taskaya Sep 8, 2025 ▶ 56:02 A Technical History of Generative Media
Assertion Contradicted
Morcos: DCLM researchers could not predict their own classifiers' filtering decisions above chance
“These are nominally the best experts you could ever hire to do this. These are students who have just spent all of their time looking at NLP data for two years. They could not predict what the DCLM classifiers would say above chance.”
Ari Morcos Aug 29, 2025 ▶ 18:06 Better Data is All You Need — Ari Morcos, Datology
Assertion Supported
Morcos: Soft inductive biases become harmful past 1M data points in vision
“Turns out in the small data regime, and when I say small data here, I mean, say less than 500,000 data points. And this was in the context of image self-supervised learning. So in that small data regime, this is super helpful. And where this paper's actually b…”
Ari Morcos Aug 29, 2025 ▶ 6:04 Better Data is All You Need — Ari Morcos, Datology
Assertion Supported
Morcos: Kaplan and Chinchilla scaling laws incorrectly assume all data is equal
“And even if you go and you look at the scaling laws work from Kaplan and Chinchilla and all these other things, they all assume IID data which is insane. We know that all data are not created equal, that garbage in garbage out is like the oldest adage in compu…”
Ari Morcos Aug 29, 2025 ▶ 8:25 Better Data is All You Need — Ari Morcos, Datology
Assertion Supported
Morcos: Proper data curation can bend neural scaling laws
“And what that paper showed was that if you use your data correctly, you can actually bend the scaling laws themselves.”
Ari Morcos Aug 29, 2025 ▶ 27:35 Better Data is All You Need — Ari Morcos, Datology
Prediction Held up
Morcos: Training a specialized frontier model will cost under $1M very soon
“I believe that getting to a frontier model should cost a million dollars or less for most organizations, at least in a specialized domain, right? And when you think about what enterprises need, that's generally what they need. They don't need a model that can …”
Ari Morcos Aug 29, 2025 ▶ 42:49 Better Data is All You Need — Ari Morcos, Datology
Prediction Didn’t hold up
Sohmers: NVIDIA Blackwell memory bandwidth efficiency will be lower than Hopper
“All indications are, even though they, you know, more than doubled the theoretical memory bandwidth going from Hopper to Blackwell, the actual percentage of theoretical that you can achieve is, again, going to be less than the previous generation”
Thomas Sohmers Aug 18, 2025 ▶ 14:53 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Assertion Contradicted
Sohmers: Google Veo and Imagen 3 are pure autoregressive transformers, not diffusion
“A lot of things have actually been moving away from diffusion to being pure autoregressive transformers for image and video generation. So like the latest, yeah, there's a VO three and since image and three on, on Google side have been pure autoregressive movi…”
Thomas Sohmers Aug 18, 2025 ▶ 43:48 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Assertion Supported
Sohmers: Positron AI requires zero compilers to run Hugging Face models
“So rather than having like, we don't have a compiler whatsoever. There's no compiler. There's no translator, no tooling that's involved in actually taking those and getting that to, you know, for your common, you know, Huggy Face Transform models to be able to…”
Thomas Sohmers Aug 18, 2025 ▶ 21:32 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Assertion Partly supported
The Information: OpenAI hit $12B ARR as burn rose to $8B
“We had a story yesterday about open AI and how, like, I think they've reached about twelve billion ARR and yeah, but their burn went from like They projected, like, one billion to, like, eight billion or something.”
Stephanie Palazzolo Aug 6, 2025 ▶ 40:26 The AI Agenda: GPT5 leaks and the business of AI News — Steph Palazzolo, The Information
Assertion Supported
Palazzolo: Claude Code leads stayed at Cursor only two weeks
“We know that they went there, they were there for, I think, about two weeks, and they came back.”
Stephanie Palazzolo Aug 6, 2025 ▶ 41:45 The AI Agenda: GPT5 leaks and the business of AI News — Steph Palazzolo, The Information
Assertion Partly supported
Inception generalist model matches Claude Haiku quality at 5-10x speed
“We had our generalist model evaluated by artificial analysis and the intelligence score from AA artificial analysis around 40. So it's comparable to GPT, 4.1 nano, cloud haiku, kind of like Close source speed optimized models. It's roughly comparable in terms …”
Stefano Ermon Aug 4, 2025 ▶ 16:55 ⚡️Mercury: Ultra-Fast Diffusion LLMs — Estefano Ermon, CEO Inception Labs
Assertion Supported
Ermon: Diffusion LLMs Pareto-dominate autoregressive models on inference efficiency
“On the inference side, what we're seeing is that diffusion models are much more efficient. We're actually able to Pareto dominate autoregressive models. If you think about the typical trade-off between throughput versus latency, which you kind of like cannot, …”
Stefano Ermon Aug 4, 2025 ▶ 14:30 ⚡️Mercury: Ultra-Fast Diffusion LLMs — Estefano Ermon, CEO Inception Labs
Assertion Partly supported
Lambert: OLMo 32B roughly matches original GPT-4 level while fully open
“Like Olmo-Thirty-Tube is if you squint like original GPT-IV level and fully open.”
Nathan Lambert Jul 31, 2025 ▶ 1:16:23 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Assertion Supported
Lambert: Tulu 3 matches or beats Meta Llama 3.1 on core evals
“On, like, core evals for our Suite of models from, I think, eight, seven D and four or five B is based on llama at the time. It's like it matches or beats meta on these core valves.”
Nathan Lambert Jul 31, 2025 ▶ 2:20 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Assertion Supported
Fortuna: New reasoning models show no big leap on medical coding tasks
“I do know that when you kind of plot out base model performance on some medical tasks like ICD-X coding between like, you know, previous generations and new reasoning generations, there's actually not like a big leap.”
Brendan Fortuna Jul 29, 2025 ▶ 24:01 ⚡️Using RFT to Build Clinical Superintelligence
Assertion Supported
OpenAI's IMO performance was not officially verified by the IMO
“It turns out, like, OpenAI actually didn't involve officially with IMO. They just, like, use the problems, but, and then just, like, use their model to test the results, and ask, like, three previous IMO analysts to review them.”
Dr. Jasper Zhang Jul 24, 2025 ▶ 3:28 ⚡️Math Olympiad gold medalist explains OpenAI and Google DeepMind IMO Gold Performances
Prediction Partly held up
McCloy: ChatGPT Search Bans for Prompt Injection Are Coming
“I think it works until it stops working. Right. And I would say like, there's not a lot of stories of people getting banned for like Chatsby D search so far, but it's coming.”
Robert McCloy Jul 23, 2025 ▶ 36:20 AI is Eating Search
Assertion Supported
McCloy: ChatGPT Does Not Index or Retrieve llms.txt by Default
“I knew that there's debate about this, but I'd say the evidence is like ChatTriPT is not indexing and it's not retrieving content from LMS.txt by default.”
Robert McCloy Jul 23, 2025 ▶ 38:52 AI is Eating Search
Prediction Didn’t hold up
Kamradt Predicts ARC-AGI-2 Will Not Be Beaten For 12 Months
“My guess is it's not going to be beat for the next 12 months.”
Greg Kamradt Jul 18, 2025 ▶ 28:29 ⚡️ARC-AGI-3: The Interactive Reasoning Benchmark
Assertion Supported
OpenAI o-series reasoning models fail at multi-tool calling benchmarks
“Then another surprise for me was that the reasoning models were not performing well enough. They had certain kind of limitation when we probed into it, like, why are they scoring less overall? They were like the O-one, the O-four, O-three, they, When not perfo…”
Pratik Bhavsar Jul 14, 2025 ▶ 8:22 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
Assertion Supported
Meta Llama 3.3 and Llama 4 perform poorly on agent benchmarks
“Another, of course, the other surprise was that all the Lama models were not performing well on our benchmark. 3.3 and even the Lama four all were really performing extremely poor.”
Pratik Bhavsar Jul 14, 2025 ▶ 10:08 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
Prediction Partly held up
Zach Lloyd: Warp's coding agent will likely top the TBench benchmark
“Basically, state of the art on SweetBench, I think we will, again, I don't want to be quoted here, we can maybe edit this later, but like, we'll probably be number one or close to it on TBench also, which is the terminal benchmark, which really we should be th…”
Zach Lloyd Jun 25, 2025 ▶ 1:44 ⚡️Warp 2.0: the Agentic Development Environment - Zach Lloyd and Ben Holmes
Assertion Supported
Duffy: OpenAI's o3 actively deceives opponents and plots betrayals in AI Diplomacy
“Oh, three was one of the few that will actually send a message to another power saying that they're planning to do something. And then like in their diary diary, right? Oh, they fell for it. Hook, line and sinker. Totally gonna betray him and take it over.”
Alex Duffy Jun 11, 2025 ▶ 12:22 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
Assertion Supported
Claude loses AI Diplomacy games because it refuses to deceive opponents
“I haven't seen Claude with any game yet because they won't do it. Like there's like, O three has managed to get them on board for like draws, even though they all know the only win condition in the game is, is 18 supply centers.”
Alex Duffy Jun 11, 2025 ▶ 12:36 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
Prediction Held up
Ameisen: Deceptive Backward Reasoning Exists in Base Pre-Trained Models
“I bet, I don't know how much I bet a hundred bucks. So somebody can like, they would get a hundred bucks from me if they prove that I'm wrong, that this behavior for a model that does a drink fine tuning, it also does it post pre-training.”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:33:39 The Utility of Interpretability — Emmanuel Amiesen
Assertion Supported
Cherny: Anthropic is currently bordering on AI Safety Level 3 capabilities
“Yeah, we're kind of bordering on three right now.”
Boris Cherny May 7, 2025 ▶ 34:33 Claude Code: Anthropic's CLI Agent
Assertion Supported
Factorio benchmark results show reasoning models underperform expectations on extended planning
“One thing we have found in preliminary results is that the reasoning models don't seem to do as well as you'd expect in this setting. And I think that's probably because the way we set this up, it's a bit like we're already making it do reasoning traces over a…”
Jack Hopkins Apr 27, 2025 ▶ 12:12 ⚡️Factorio Learning Environment: the ultimate Game Agent Eval — Jack Hopkins
Prediction Didn’t hold up
Conrad: GPU market will likely return to a shortage by winter
“My general prediction is that like by the winter we will be back towards shortage, but then also this very much depends on The rollout of future chips.”
Evan Conrad Apr 11, 2025 ▶ 30:48 SF Compute: Commoditizing Compute
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.