The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 90 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Supported
Lie: Cerebras runs OpenAI's flagship model 14x faster than GPUs
“We're running you know, frontier level, one of the most intelligent models, right? OpenAI's largest, most capable, most intelligent model right now at 14 times faster than their normal, you know, GPU speeds.”
Sean Lie Sep 2, 2026 ▶ 8:37 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Assertion Supported
Kant: Major AI labs did not prioritize RL for LLMs three years ago
“And the second was that reinforcement learning was going to be the biggest driver for LLM capabilities. Today, very obvious three years ago was not an opinion held or direction held at either OpenAI or Google or Anthropic or others.”
Eiso Kant Jul 22, 2026 ▶ 6:03 The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
Assertion Supported
Reddy: Voxtral speech model is much stronger than Whisper
“And I think a big people, I think there's a big rich ecosystem of people finding whisper and people want the same thing with Voxer. It's much stronger than whisper.”
Pavan Kumar Reddy Mar 30, 2026 ▶ 26:34 Mistral: Voxtral TTS, Forge, Leanstral, & Mistral 4 — w/ Pavan Kumar Reddy & Guillaume Lample
Assertion Supported
Watkins: Over half of SWE-bench problems investigated by OpenAI had test flaws
“In over half of the problems that were investigated in that deep dive, there was one problem or the other. I think the most common problem are, like, overly narrow tests where there's some particular implementation detail that the tests were looking for but wa…”
Olivia Watkins Feb 23, 2026 ▶ 7:26 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
Assertion Supported
Sam Altman Barred Investors Who Backed Glean From Investing in OpenAI
“Sam Altman once came out and said, if you're an investor in OpenAI and one of these five companies, including Glean, we don't want you as an investor.”
Deedy Das Nov 14, 2025 ▶ 9:02 Anthropic, Glean & OpenRouter: How AI Moats Are Built with Deedy Das of Menlo Ventures
Assertion Supported
OpenAI's IMO performance was not officially verified by the IMO
“It turns out, like, OpenAI actually didn't involve officially with IMO. They just, like, use the problems, but, and then just, like, use their model to test the results, and ask, like, three previous IMO analysts to review them.”
Dr. Jasper Zhang Jul 24, 2025 ▶ 3:28 ⚡️Math Olympiad gold medalist explains OpenAI and Google DeepMind IMO Gold Performances
Assertion Supported
Lambert: RLHF has not been shown to improve underlying model benchmark capabilities
“RLHF is not that shown to improve capabilities yet. I think one of the fun ones is from the GPT-IV technical report. They essentially listed their kind of bogus evaluations, because it's a hilarious table, because it's like LSAT AP exams, and then like AMC-X a…”
Nathan Lambert Jan 11, 2024 ▶ 59:53 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Assertion Supported
Hong: All OpenAI formal math researchers have left the company
“No, no, they all left.”
Carina Hong Jun 3, 2026 ▶ 1:05:45 Scaling Past Informal AI - Carina Hong, Axiom Math
Prediction Held up
Chen: AI agents will master GUI-based computer use by 2026
“And I can continue just by sort of like saying that that's definitely going to be something I think is going to be something that we'll be capable of in 20, 26.”
Bill Chen Dec 26, 2025 ▶ 25:40 ⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust — Brian Fioca + Bill Chen, OpenAI
Assertion Supported
Fioca: Codex Max manages its own context window to run indefinitely
“Codex Max manages its own context window. And so it can run basically forever without you having to worry about it while it's inside of the Codex harness.”
Brian Fioca Dec 26, 2025 ▶ 14:54 ⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust — Brian Fioca + Bill Chen, OpenAI
Assertion Supported
Glean Operates at a Several Hundred Million Dollar Revenue Scale
“Look at the revenue of Anthropic and OpenAI right now. These are billion dollar revenue scale businesses. Glean is several hundred million dollar revenue scale business.”
Deedy Das Nov 14, 2025 ▶ 9:16 Anthropic, Glean & OpenRouter: How AI Moats Are Built with Deedy Das of Menlo Ventures
Assertion Supported
Brockman: OpenAI's robotics team pivoted to build GitHub Copilot
“And we've been through times where, for example, robotics was one in 2018, where we had a great result, but we kind of realized that actually, like, that we can move so much faster in a different domain, right? That, that actually, you know, we had this great …”
Greg Brockman Aug 15, 2025 ▶ 1:02:01 Greg Brockman on OpenAI's Road to AGI
Prediction Held up
OpenAI's technology will surpass o3 within six months
“I think that Oh, three is not where the technology will be in six months.”
Noam Brown Jun 19, 2025 ▶ 38:35 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Assertion Supported
Fanelli: Anthropic scrapes 6,000 pages per referral, compared to OpenAI's 250
“Google would be a two to one crawl to referral ratio, so for every two pages, they will read, they will send you one visitor. He said OpenAI is 250 to one, so they'll read 250 of your pages and send you one person. And Anthropic was like 6000 to one. So they'l…”
Alessio Fanelli Apr 24, 2025 ▶ 50:01 Why Every Agent needs Open Source Cloud Sandboxes
Assertion Supported
Reddy: OpenAI's share of enterprise LLM spend dropped from 90% to 60%
“And the opening I spend at the beginning, at the end of last year in November of 23 was close to 90% of total volume. And today, less than a year later, it's closer to 60% of total volume.”
Pranav Reddy Dec 21, 2024 ▶ 5:08 The State of AI Startups in 2024 [LS Live @ NeurIPS]
Assertion Supported
Angelopoulos: OpenAI o1 crushed Chatbot Arena, proving the benchmark isn't saturated
“So there's this model and it crushed the benchmark. You know, it's just like really like a big gap. And what that's telling us is that it's not saturated yet. And so it's still measuring some signal that was encouraging point.”
Anastasios Angelopoulos Nov 1, 2024 ▶ 27:20 In the Arena: How LMSys changed LLM Benchmarking Forever
Assertion Supported
Hu: OpenAI o1-preview surpasses human Kaggle Grandmasters with seven gold medals
“Since a grandmaster requires five gold medals and oh, and preview gets an average of eight or sorry, seven gold medals. They're out competing even capital grandmasters.”
Jesse Hu Oct 19, 2024 ▶ 47:29 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
Assertion Supported
Schulhoff: Preamble Discovered Prompt Injection Before Riley Goodside
“Preamble is the company that first discovered Prompt Injection, even before Riley, and they, like, responsibly disclosed it, kind of, internally to OpenAI”
Sander Schulhoff Sep 20, 2024 ▶ 4:46 The Ultimate Guide to Prompting - with Sander Schulhoff from LearnPrompting.org
Assertion Supported
Carlini extracted production models from Google and OpenAI with legal permission
“We ran the attack that let us, yeah, stole several of OpenAI's models. With their permission... We notified everyone who was vulnerable to this attack. Some Google models were vulnerable. Some open AM models were vulnerable. There were one or two other people …”
Nicholas Carlini Aug 28, 2024 ▶ 57:21 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Assertion Supported
Reddit makes over 200 million dollars in AI data licensing deals
“Yeah, the, I guess the winner in all of this is Reddit, which is making over two hundred million just in data licensing to OpenAI and some of the other AI providers.”
Alessio Fanelli Aug 2, 2024 ▶ 36:09 The Winds of AI Winter (Q2 Four Wars of the AI Stack Recap)
Assertion Supported
Lambert: OpenAI retrains reward models with curated and user prompt mixtures
“And this is like a sort of outer loop optimization that no one in the open is even remotely qualified to talk about, but OpenAI does monitor and they'll like rerun RLHF and train a new reward model with a mixture of their curated data and user prompts to try t…”
Nathan Lambert Jan 11, 2024 ▶ 1:32:27 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Assertion Supported
Patel: Google, OpenAI, and Anthropic are developing multi-datacenter training
“One of the big bottlenecks is how much power and how many chips you can get into a single data center. So, like, A, Google and OpenAI and Anthropic are working on this, right?”
Dylan Patel Dec 5, 2023 ▶ 1:06:37 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Assertion Supported
Parakhin: Bing Sydney first launched in India using Megatron, not OpenAI
“The funny thing, I mean, the most interesting anecdote is that Sydney was first shipped in India for and it was not noticed for a long time. And first implementation of Sydney didn't even have open AI model under it. It was during Megatron. Microsoft and the N…”
Mikhail Parakhin Apr 22, 2026 ▶ 1:10:53 AI-Native Engineering: 100% adoption, 5x search throughput, unlimited tokens — Mikhail Parakhin
Assertion Supported
Watkins: About 90% of SWE-bench Verified tasks take under an hour
“For Sweep Edge Verified I think that's something like, 90% of the problems are things that were estimated to take, like, an expert software engineer like, less than an hour.”
Olivia Watkins Feb 23, 2026 ▶ 10:59 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
Assertion Supported
Pre-December 2024 OpenAI models did not exhibit seahorse emoji self-correction loops
“And so I like ran the OpenAI API across like models released from 23 to 25, and you would see like all the models until twenty-twenty-four December had very Terce and short responses to the question. Is there a seahorse emoji? They would either say that there …”
Pratyush Maini Feb 10, 2026 ▶ 9:04 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
Assertion Supported
White: OpenAI reached out to red team new models after reading his chemistry paper
“And then opening eye, some people there Lama was there. She saw this paper, and they reached out, like, hey, we're building this new model, and we think it'd be great to red team it to see, like, what could happen with these models if they're applied to chemis…”
Andrew White Jan 28, 2026 ▶ 8:58 🔬 From Red Teaming GPT-4 to Automating Drug Discovery: The Future of AI in Science — Andrew White
Assertion Supported
Weil: OpenAI's Prism allows unlimited collaborators for free
“I think most other tools in the space have hard limits and charge you money and other things. In Prism, it's as many collaborators as you want for free.”
Kevin Weil Jan 27, 2026 ▶ 16:04 ⚡️ Prism: OpenAI's LaTeX "Cursor for Scientists" — Kevin Weil & Victor Powell, OpenAI for Science
Assertion Supported
McGrath: GPT-5 Thinking Matches or Beats Deep Research on Published Evals
“I mean, I think if you look at our published evals, they're, they look, like, basically on par if it's not better, so, like, I mean, that's personally what I do.”
Josh McGrath Dec 31, 2025 ▶ 6:46 [State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI
Assertion Supported
Anthropic and OpenAI are collaborating to build a unified AI UI standard
“And now one thing we just announced three weeks ago on the MCP blog is that we're actually working with all, all two of them together to build like a common standard.”
David Soria Parra Dec 28, 2025 ▶ 53:04 One Year of MCP — with David Soria Parria and AAIF leads from OpenAI, Goose, Linux Foundation
Assertion Supported
Google, Microsoft, Amazon, OpenAI, and Anthropic joined AAIF as platinum members
“You have Google, Microsoft, Amazon Block, Bloomberg, Cloudflare, OpenAI, Anthropic. Just a platinum member, create a foundation.”
David Soria Parra Dec 28, 2025 ▶ 1:35:53 One Year of MCP — with David Soria Parria and AAIF leads from OpenAI, Goose, Linux Foundation
Assertion Supported
Fioca: Codex Max can run continuously for 24 hours or more
“Max can run for a really long time. We can go 24 hours or more. I've actually, like, sort of had it gone for more than that”
Brian Fioca Dec 26, 2025 ▶ 1:45 ⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust — Brian Fioca + Bill Chen, OpenAI
Assertion Supported
Fioca: GPT-5 matches Codex coding capability but adds step-by-step preambles
“With the five series, because it's more general, and it's just about as good as coding as codex for a lot of things. We've taught it to be more communicative. And so it has preambles before tool calls. It'll say things like, I'm about to go look for this.”
Brian Fioca Dec 26, 2025 ▶ 10:39 ⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust — Brian Fioca + Bill Chen, OpenAI
Assertion Supported
OpenAI adds third-party model support to its evals product
“One of the things that we launched today with evals too is ability to use, like, third-party models as well and kind of bring that into one place”
Christina Huang Oct 7, 2025 ▶ 16:38 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Assertion Supported
OpenAI's API throughput has surpassed six billion tokens per minute
“We actually zoomed past that.”
Sherwin Wu Oct 7, 2025 ▶ 44:33 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Assertion Supported
OpenAI integrates with OpenRouter for multi-provider evals
“We have a really cool setup with Open Router, where we're working with them, and then you can bring your Open Router setup. And then with that, you can actually, you know, you write your evals using our data sets tool, or use our data set tool to create a bunc…”
Sherwin Wu Oct 7, 2025 ▶ 17:02 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Assertion Supported
OpenAI launches Agent Kit to build, deploy, and optimize agents
“We launched Agent Kit today. Full set of solutions to build, deploy, and optimize agents.”
Christina Huang Oct 7, 2025 ▶ 9:26 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Assertion Supported
Wu: OpenAI was first to launch stateful responses API
“Obviously we were the first one to launch responses API, but like a couple of other people have kind of adopted, I think Grok has it in their API. I think I saw LMSYS just did something”
Sherwin Wu Oct 7, 2025 ▶ 15:46 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Assertion Supported
Feldman: Sam Altman and Ilya Sutskever invested in Cerebras' early rounds
“In 2016, we met with Sam Altman and Ilya Suskovard at OpenAI and they were an idea and we were PowerPoint, right? That's amazing. And what AI was doing was identifying cats in pictures. And I think they ended up investing in us, both of them and many of their …”
Andrew Feldman Oct 1, 2025 ▶ 1:47 ⚡️Raising $1.1b to build the fastest LLM Chips on Earth — Andrew Feldman, Cerebras
Prediction Held up
Brockman: Most AI compute will shift from training to inference
“We're going to move from a world where most of the compute is training the model as we've deployed these models more, you know, more of the compute goes to inferencing them and actually using them.”
Greg Brockman Aug 15, 2025 ▶ 14:13 Greg Brockman on OpenAI's Road to AGI
Assertion Supported
DeepMind and OpenAI eliminated formal Lean translation for 2025 IMO solutions
“What surprised me is this time they don't use formal language, but instead they just use LM. And so last year when they tried to do the IMO, they need like a like a human to kind of translate the natural language. Problems to Lean, and then they use Lean to ki…”
Dr. Jasper Zhang Jul 24, 2025 ▶ 6:10 ⚡️Math Olympiad gold medalist explains OpenAI and Google DeepMind IMO Gold Performances
Assertion Supported
OpenAI increases prompt caching discount from 50% to 75% on GPT-4.1
“We've increased our prompt caching discount from 50% to 75% on these models.”
Michelle Pokrass Apr 15, 2025 ▶ 43:27 GPT 4.1: The New OpenAI Workhorse
Assertion Supported
GPT-4.1 reduces extraneous edit rate to 2%, down from GPT-4o's 9%
“And we found that from four O, which got nine percent, which is pretty crazy, nine percent of the time making an extraneous edit is a lot. 4.1 is at two percent, so it's a pretty big improvement.”
Michelle Pokrass Apr 15, 2025 ▶ 23:43 GPT 4.1: The New OpenAI Workhorse
Assertion Supported
OpenAI launches GPT-4.1 model lineup featuring 1M-token context window
“Yeah, I'll just say we released three new models today, GPT-Fort.one, GPT-Fort.one mini, and GPT-Fort.one data, and the real focus on these were just making the models that were great for developers so we improved instruction following, coding, and shipped our…”
Michelle Pokrass Apr 15, 2025 ▶ 1:27 GPT 4.1: The New OpenAI Workhorse
Assertion Supported
Swix: Google's entire SigLIP vision team left to join OpenAI
“I think the most recent notable move, I think the entire vision team from Google Lucas Beyer and all the other authors of Siglip left Google to join OpenAI”
Shawn Wang Mar 14, 2025 ▶ 3:24 Snipd: The AI Podcast App for Learning — with CEO Kevin Ben-Smith
Assertion Supported
Swyx: OpenAI Assistants API target sunset is H1 2026
“And assistance API we've has a target sunset date of first half of 26.”
Shawn Wang Mar 11, 2025 ▶ 3:27 The new OpenAI Agents Platform: CUA, Web Search, Responses API, Agents SDK!!
Assertion Supported
Nikunj Handa: OpenAI distilled o-series models into GPT-4o search
“They use, like, synthetic data techniques. They've done, like, O-series model distillation to, like, make these four or fine tunes really good.”
Nikunj Handa Mar 11, 2025 ▶ 9:26 The new OpenAI Agents Platform: CUA, Web Search, Responses API, Agents SDK!!
Assertion Supported
Colvin: AI providers are centralizing around OpenAI's API standard
“I think the truth is that everyone is centralizing around OpenAI's SD API as the one to do. So DeepSeek support that. Grok with a K support that. Olama also does it. Well, I mean, if there is that library right now, it's more or less the OpenAI SDK.”
Samuel Colvin Feb 6, 2025 ▶ 30:32 Agent Engineering with Pydantic + Graphs — with Samuel Colvin, CEO of Pydantic Logfire
Assertion Supported
Lambert: Reinforcement fine-tuning requires only dozens of labeled samples
“This reinforcement fine tuning does many passes over the data, which is why they can say you only need dozens of labeled samples to actually learn from it, which is very different than. Previous training regimes”
Nathan Lambert Jan 2, 2025 ▶ 9:57 The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024]
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.