The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 107 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Partly supported
Petersson: Anthropic's Claude models uniquely exhibit emergent deceptive and cartel behaviors
“So every single model from Anthropic since have been going in this direction. And I think one interesting thing is that like, OpenAI models don't. They, Quite plainly, they don't, they behave really well. And you know, you don't know if this is like, good, lik…”
Lukas Petersson Jun 4, 2026 ▶ 46:27 When AI Agents Run Businesses — Lukas Petersson and Axel Backlund of Andon Labs
Prediction Didn’t hold up
Swix: OpenAI will issue a cryptocurrency token to fund compute
“There is still one more shoe to drop, which is the non sovereign wealth funding that open AI needs to get, which they've promised to drop by the end of this year. And my money is on, they have to do a coin. Like it's, I'm not a crypto guy at all, but like, y…”
Shawn Wang Oct 16, 2025 ▶ 49:13 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
Assertion Supported
Lie: Cerebras runs OpenAI's flagship model 14x faster than GPUs
“We're running you know, frontier level, one of the most intelligent models, right? OpenAI's largest, most capable, most intelligent model right now at 14 times faster than their normal, you know, GPU speeds.”
Sean Lie Sep 2, 2026 ▶ 8:37 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Assertion Supported
Kant: Major AI labs did not prioritize RL for LLMs three years ago
“And the second was that reinforcement learning was going to be the biggest driver for LLM capabilities. Today, very obvious three years ago was not an opinion held or direction held at either OpenAI or Google or Anthropic or others.”
Eiso Kant Jul 22, 2026 ▶ 6:03 The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
Assertion Supported
Reddy: Voxtral speech model is much stronger than Whisper
“And I think a big people, I think there's a big rich ecosystem of people finding whisper and people want the same thing with Voxer. It's much stronger than whisper.”
Pavan Kumar Reddy Mar 30, 2026 ▶ 26:34 Mistral: Voxtral TTS, Forge, Leanstral, & Mistral 4 — w/ Pavan Kumar Reddy & Guillaume Lample
Assertion Supported
Watkins: Over half of SWE-bench problems investigated by OpenAI had test flaws
“In over half of the problems that were investigated in that deep dive, there was one problem or the other. I think the most common problem are, like, overly narrow tests where there's some particular implementation detail that the tests were looking for but wa…”
Olivia Watkins Feb 23, 2026 ▶ 7:26 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
Assertion Partly supported
OpenAI Plans to Scale Compute Power Capacity to 125 Gigawatts
“For OpenAI to go from like two gigawatts of compute this year to 30 with everything they've already announced, and then there's a plan for the next 125. Like, the United States uses 300.”
Shawn Wang Nov 14, 2025 ▶ 1:15:33 Anthropic, Glean & OpenRouter: How AI Moats Are Built with Deedy Das of Menlo Ventures
Assertion Supported
Sam Altman Barred Investors Who Backed Glean From Investing in OpenAI
“Sam Altman once came out and said, if you're an investor in OpenAI and one of these five companies, including Glean, we don't want you as an investor.”
Deedy Das Nov 14, 2025 ▶ 9:02 Anthropic, Glean & OpenRouter: How AI Moats Are Built with Deedy Das of Menlo Ventures
Assertion Partly supported
The Information: OpenAI hit $12B ARR as burn rose to $8B
“We had a story yesterday about open AI and how, like, I think they've reached about twelve billion ARR and yeah, but their burn went from like They projected, like, one billion to, like, eight billion or something.”
Stephanie Palazzolo Aug 6, 2025 ▶ 40:26 The AI Agenda: GPT5 leaks and the business of AI News — Steph Palazzolo, The Information
Assertion Supported
OpenAI's IMO performance was not officially verified by the IMO
“It turns out, like, OpenAI actually didn't involve officially with IMO. They just, like, use the problems, but, and then just, like, use their model to test the results, and ask, like, three previous IMO analysts to review them.”
Dr. Jasper Zhang Jul 24, 2025 ▶ 3:28 ⚡️Math Olympiad gold medalist explains OpenAI and Google DeepMind IMO Gold Performances
Assertion Supported
Lambert: RLHF has not been shown to improve underlying model benchmark capabilities
“RLHF is not that shown to improve capabilities yet. I think one of the fun ones is from the GPT-IV technical report. They essentially listed their kind of bogus evaluations, because it's a hilarious table, because it's like LSAT AP exams, and then like AMC-X a…”
Nathan Lambert Jan 11, 2024 ▶ 59:53 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Assertion Supported
Hong: All OpenAI formal math researchers have left the company
“No, no, they all left.”
Carina Hong Jun 3, 2026 ▶ 1:05:45 Scaling Past Informal AI - Carina Hong, Axiom Math
Assertion Contradicted
All major US AI labs stopped publishing research after OpenAI closed
“Whereas in the United States, since OpenAI closed their doors and stopped publishing, so did all the other labs.”
Andy Konwinski Dec 31, 2025 ▶ 19:25 [State of Research Funding] Beyond NSF, Slingshots, Open Frontiers — Andy Konwinski, Laude Institute
Prediction Held up
Chen: AI agents will master GUI-based computer use by 2026
“And I can continue just by sort of like saying that that's definitely going to be something I think is going to be something that we'll be capable of in 20, 26.”
Bill Chen Dec 26, 2025 ▶ 25:40 ⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust — Brian Fioca + Bill Chen, OpenAI
Assertion Supported
Fioca: Codex Max manages its own context window to run indefinitely
“Codex Max manages its own context window. And so it can run basically forever without you having to worry about it while it's inside of the Codex harness.”
Brian Fioca Dec 26, 2025 ▶ 14:54 ⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust — Brian Fioca + Bill Chen, OpenAI
Assertion Contradicted
OpenAI Spent $7 Billion on Compute, With $5 Billion for R&D
“This year, OpenAI spent seven billion dollars on compute. Only two of that was for all of their inference. The remaining five was R&D. So all of ChatGPT, all eight hundred million users, all of Sora, all of like, all, all the sort of like API volume, two billi…”
Shawn Wang Nov 14, 2025 ▶ 1:17:40 Anthropic, Glean & OpenRouter: How AI Moats Are Built with Deedy Das of Menlo Ventures
Assertion Supported
Glean Operates at a Several Hundred Million Dollar Revenue Scale
“Look at the revenue of Anthropic and OpenAI right now. These are billion dollar revenue scale businesses. Glean is several hundred million dollar revenue scale business.”
Deedy Das Nov 14, 2025 ▶ 9:16 Anthropic, Glean & OpenRouter: How AI Moats Are Built with Deedy Das of Menlo Ventures
Assertion Contradicted
Martin: OpenDeep Research is the top-ranked open-source Deep Research agent
“OpenDeep Research is a deep research agent that I've been working on for about a year, and it's now, according to Deep Research Spence, the best performing Deep Research agent at least on that particular benchmark. So it's pretty good. Listen, it's not as good…”
Lance Martin Sep 11, 2025 ▶ 8:32 Context Engineering for Agents - Lance Martin, LangChain
Assertion Supported
Brockman: OpenAI's robotics team pivoted to build GitHub Copilot
“And we've been through times where, for example, robotics was one in 2018, where we had a great result, but we kind of realized that actually, like, that we can move so much faster in a different domain, right? That, that actually, you know, we had this great …”
Greg Brockman Aug 15, 2025 ▶ 1:02:01 Greg Brockman on OpenAI's Road to AGI
Assertion Partly supported
Early GPT-5 testers report noticeable gains across coding, science, and writing
“And the story that we wrote, we kind of talked about how at least the people that we've talked to who tested it so far have been pretty impressed. They seem to think that it's been, you know, there's been improvements in a number of domains and both like scien…”
Stephanie Palazzolo Aug 6, 2025 ▶ 28:49 The AI Agenda: GPT5 leaks and the business of AI News — Steph Palazzolo, The Information
Prediction Held up
OpenAI's technology will surpass o3 within six months
“I think that Oh, three is not where the technology will be in six months.”
Noam Brown Jun 19, 2025 ▶ 38:35 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Assertion Supported
Fanelli: Anthropic scrapes 6,000 pages per referral, compared to OpenAI's 250
“Google would be a two to one crawl to referral ratio, so for every two pages, they will read, they will send you one visitor. He said OpenAI is 250 to one, so they'll read 250 of your pages and send you one person. And Anthropic was like 6000 to one. So they'l…”
Alessio Fanelli Apr 24, 2025 ▶ 50:01 Why Every Agent needs Open Source Cloud Sandboxes
Assertion Partly supported
Swyx: OpenAI production market share dropped from 95% to 50–75%
“Basically over the course of 23, going into 24, OpenAI has gone from 95 market share to reasonably somewhere between 50 to 75 market share.”
Shawn Wang Jan 1, 2025 ▶ 12:20 2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
Assertion Supported
Reddy: OpenAI's share of enterprise LLM spend dropped from 90% to 60%
“And the opening I spend at the beginning, at the end of last year in November of 23 was close to 90% of total volume. And today, less than a year later, it's closer to 60% of total volume.”
Pranav Reddy Dec 21, 2024 ▶ 5:08 The State of AI Startups in 2024 [LS Live @ NeurIPS]
Assertion Partly supported
Polu: GPT-4 was ready internally at OpenAI months before September 2022
“I had seen GPT-IV internally at the time. It was September, 20, 22. So it was pre-chat GPT, but GPT-IV was ready since, I mean, I'd been ready for a few months internally.”
Stanislas Polu Nov 11, 2024 ▶ 19:16 Agents @ Work: Dust.tt — with Stanislas Polu
Assertion Supported
Angelopoulos: OpenAI o1 crushed Chatbot Arena, proving the benchmark isn't saturated
“So there's this model and it crushed the benchmark. You know, it's just like really like a big gap. And what that's telling us is that it's not saturated yet. And so it's still measuring some signal that was encouraging point.”
Anastasios Angelopoulos Nov 1, 2024 ▶ 27:20 In the Arena: How LMSys changed LLM Benchmarking Forever
Assertion Supported
Hu: OpenAI o1-preview surpasses human Kaggle Grandmasters with seven gold medals
“Since a grandmaster requires five gold medals and oh, and preview gets an average of eight or sorry, seven gold medals. They're out competing even capital grandmasters.”
Jesse Hu Oct 19, 2024 ▶ 47:29 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
Assertion Supported
Schulhoff: Preamble Discovered Prompt Injection Before Riley Goodside
“Preamble is the company that first discovered Prompt Injection, even before Riley, and they, like, responsibly disclosed it, kind of, internally to OpenAI”
Sander Schulhoff Sep 20, 2024 ▶ 4:46 The Ultimate Guide to Prompting - with Sander Schulhoff from LearnPrompting.org
Assertion Supported
Carlini extracted production models from Google and OpenAI with legal permission
“We ran the attack that let us, yeah, stole several of OpenAI's models. With their permission... We notified everyone who was vulnerable to this attack. Some Google models were vulnerable. Some open AM models were vulnerable. There were one or two other people …”
Nicholas Carlini Aug 28, 2024 ▶ 57:21 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Assertion Supported
Reddit makes over 200 million dollars in AI data licensing deals
“Yeah, the, I guess the winner in all of this is Reddit, which is making over two hundred million just in data licensing to OpenAI and some of the other AI providers.”
Alessio Fanelli Aug 2, 2024 ▶ 36:09 The Winds of AI Winter (Q2 Four Wars of the AI Stack Recap)
Assertion Supported
Lambert: OpenAI retrains reward models with curated and user prompt mixtures
“And this is like a sort of outer loop optimization that no one in the open is even remotely qualified to talk about, but OpenAI does monitor and they'll like rerun RLHF and train a new reward model with a mixture of their curated data and user prompts to try t…”
Nathan Lambert Jan 11, 2024 ▶ 1:32:27 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Assertion Supported
Patel: Google, OpenAI, and Anthropic are developing multi-datacenter training
“One of the big bottlenecks is how much power and how many chips you can get into a single data center. So, like, A, Google and OpenAI and Anthropic are working on this, right?”
Dylan Patel Dec 5, 2023 ▶ 1:06:37 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Assertion Contradicted
No commercial products augmented GPT models with custom data pre-ChatGPT
“Like I saw some people doing demos, but like in like a CLI or something like that, but there was no product doing like this model, but with additional data on top of it.”
Yasser Elsaid May 2, 2026 ▶ 2:59 ⚡️ Competing with ChatGPT and Sierra, building a $10M ARR company — Yasser Elsaid, Founder, Chatbase
Assertion Supported
Parakhin: Bing Sydney first launched in India using Megatron, not OpenAI
“The funny thing, I mean, the most interesting anecdote is that Sydney was first shipped in India for and it was not noticed for a long time. And first implementation of Sydney didn't even have open AI model under it. It was during Megatron. Microsoft and the N…”
Mikhail Parakhin Apr 22, 2026 ▶ 1:10:53 AI-Native Engineering: 100% adoption, 5x search throughput, unlimited tokens — Mikhail Parakhin
Assertion Supported
Watkins: About 90% of SWE-bench Verified tasks take under an hour
“For Sweep Edge Verified I think that's something like, 90% of the problems are things that were estimated to take, like, an expert software engineer like, less than an hour.”
Olivia Watkins Feb 23, 2026 ▶ 10:59 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
Prediction Didn’t hold up
Swyx: OpenAI will always release both general and Codex model variants
“I'm pretty, like, have pretty high confidence that basically OpenAI will always release a GPT-V and a GPT-V codex.”
Shawn Wang Feb 19, 2026 ▶ 36:15 Inside AI’s $10B+ Capital Flywheel — Martin Casado & Sarah Wang of a16z
Assertion Supported
Pre-December 2024 OpenAI models did not exhibit seahorse emoji self-correction loops
“And so I like ran the OpenAI API across like models released from 23 to 25, and you would see like all the models until twenty-twenty-four December had very Terce and short responses to the question. Is there a seahorse emoji? They would either say that there …”
Pratyush Maini Feb 10, 2026 ▶ 9:04 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
Assertion Supported
White: OpenAI reached out to red team new models after reading his chemistry paper
“And then opening eye, some people there Lama was there. She saw this paper, and they reached out, like, hey, we're building this new model, and we think it'd be great to red team it to see, like, what could happen with these models if they're applied to chemis…”
Andrew White Jan 28, 2026 ▶ 8:58 🔬 From Red Teaming GPT-4 to Automating Drug Discovery: The Future of AI in Science — Andrew White
Assertion Supported
Weil: OpenAI's Prism allows unlimited collaborators for free
“I think most other tools in the space have hard limits and charge you money and other things. In Prism, it's as many collaborators as you want for free.”
Kevin Weil Jan 27, 2026 ▶ 16:04 ⚡️ Prism: OpenAI's LaTeX "Cursor for Scientists" — Kevin Weil & Victor Powell, OpenAI for Science
Assertion Partly supported
McGrath: GPT-5.1 dramatically reduced token usage over GPT-5 while boosting evals
“Yeah, and so you can see, like, from five to 5.1, our overall evals, you know, we bumped some. But if you look at a two D plot of how many tokens it takes for us to get that, it went way down.”
Josh McGrath Dec 31, 2025 ▶ 13:58 [State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI
Assertion Supported
McGrath: GPT-5 Thinking Matches or Beats Deep Research on Published Evals
“I mean, I think if you look at our published evals, they're, they look, like, basically on par if it's not better, so, like, I mean, that's personally what I do.”
Josh McGrath Dec 31, 2025 ▶ 6:46 [State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI
Assertion Supported
Anthropic and OpenAI are collaborating to build a unified AI UI standard
“And now one thing we just announced three weeks ago on the MCP blog is that we're actually working with all, all two of them together to build like a common standard.”
David Soria Parra Dec 28, 2025 ▶ 53:04 One Year of MCP — with David Soria Parria and AAIF leads from OpenAI, Goose, Linux Foundation
Assertion Supported
Google, Microsoft, Amazon, OpenAI, and Anthropic joined AAIF as platinum members
“You have Google, Microsoft, Amazon Block, Bloomberg, Cloudflare, OpenAI, Anthropic. Just a platinum member, create a foundation.”
David Soria Parra Dec 28, 2025 ▶ 1:35:53 One Year of MCP — with David Soria Parria and AAIF leads from OpenAI, Goose, Linux Foundation
Assertion Supported
Fioca: Codex Max can run continuously for 24 hours or more
“Max can run for a really long time. We can go 24 hours or more. I've actually, like, sort of had it gone for more than that”
Brian Fioca Dec 26, 2025 ▶ 1:45 ⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust — Brian Fioca + Bill Chen, OpenAI
Assertion Supported
Fioca: GPT-5 matches Codex coding capability but adds step-by-step preambles
“With the five series, because it's more general, and it's just about as good as coding as codex for a lot of things. We've taught it to be more communicative. And so it has preambles before tool calls. It'll say things like, I'm about to go look for this.”
Brian Fioca Dec 26, 2025 ▶ 10:39 ⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust — Brian Fioca + Bill Chen, OpenAI
Assertion Supported
OpenAI adds third-party model support to its evals product
“One of the things that we launched today with evals too is ability to use, like, third-party models as well and kind of bring that into one place”
Christina Huang Oct 7, 2025 ▶ 16:38 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Assertion Supported
OpenAI's API throughput has surpassed six billion tokens per minute
“We actually zoomed past that.”
Sherwin Wu Oct 7, 2025 ▶ 44:33 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Assertion Supported
OpenAI integrates with OpenRouter for multi-provider evals
“We have a really cool setup with Open Router, where we're working with them, and then you can bring your Open Router setup. And then with that, you can actually, you know, you write your evals using our data sets tool, or use our data set tool to create a bunc…”
Sherwin Wu Oct 7, 2025 ▶ 17:02 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.