The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 476 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 11 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Not checkable as stated
Yegge: OpenAI sees a 10x productivity gap between AI adopters and non-adopters
“Anecdotally, they're sharing that performance, the performance differences like 10 X By any way that you measure it. So lines of code, commits, business impact, whatever. And it's so stark and pronounced that the people who aren't adopting it are now 10 times …”
Steve Yegge Dec 26, 2025 ▶ 3:02 Steve Yegge's Vibe Coding Manifesto: Why Claude Code Isn't It & What Comes After the IDE
What-if
Google would have crushed OpenAI by giving Noam Shazeer half its TPUs
“That muscle did not exist during my time at Google. And I think had they had it, what they would have done would be say, hey, Noam Shazir, you're a brilliant guy. You know how to scale these things up? Like, here's half of all of our TPUs. And then I think the…”
David Luan Mar 27, 2024 ▶ 8:29 Why Google failed to make GPT-3 -- with David Luan of Adept
Assertion Not publicly verifiable
Hotz: GPT-4 is an 8-way mixture model with 220B parameters per head
“GPT-IV is two hundred twenty billion in each head, and then it's an eight-way mixture model.”
George Hotz Jun 20, 2023 ▶ 49:48 Ep 18: Petaflops to the People — with George Hotz of tinycorp
Opinion
Hotz: Meta attracts researchers who want to publish while OpenAI keeps ideologues
“OpenAI can keep ideologues who, you know, believe ideological stuff, and Facebook can keep every researcher who's like, dude, I just want to build AI and publish it.”
George Hotz Jun 20, 2023 ▶ 55:16 Ep 18: Petaflops to the People — with George Hotz of tinycorp
Opinion
OpenAI built a significantly better GPU than Nvidia with Jalapeno chip
“I think that, like, they, you know, they pushed a lot on the performance and the fact that they're significantly better than, you know, better performance than the GPU than NVIDIA. But what I see is that they've built a significantly better GPU. And that in it…”
Sean Lie Sep 2, 2026 ▶ 16:30 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Disclosure
Chen: OpenAI's three-year goal is models conducting end-to-end research
“When we look at our kind of three-year roadmap, right the end goal that we want to reach is one where You know, the models are just doing end-to-end research, and I think a part of that problem is just being able to have the model come up with good taste.”
Mark Chen Jun 25, 2026 ▶ 34:15 Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen
Assertion Partly supported
Petersson: Anthropic's Claude models uniquely exhibit emergent deceptive and cartel behaviors
“So every single model from Anthropic since have been going in this direction. And I think one interesting thing is that like, OpenAI models don't. They, Quite plainly, they don't, they behave really well. And you know, you don't know if this is like, good, lik…”
Lukas Petersson Jun 4, 2026 ▶ 46:27 When AI Agents Run Businesses — Lukas Petersson and Axel Backlund of Andon Labs
Assertion Not checkable as stated
GPT-5 Reproduced Lupsasca's Best Physics Paper in 30 Minutes
“Then when GPT-V came out. It was able to reproduce one of my best papers that took me a very long time to come up with, in like, 30 minutes.”
Alex Lupsasca May 5, 2026 ▶ 2:32 🔬How GPT‑5 derived new results in theoretical physics and quantum gravity — Alex Lupsasca, OpenAI
Assertion Not checkable as stated
Internal OpenAI Model Proved Gluon Amplitude Formula in 12 Hours
“We had this Internal model that could think for a very long time and was extra strong in physics. So we gave it the whole problem from scratch without actually giving it this. We just formulated the problem in a very sharp way and asked the model to solve, to …”
Alex Lupsasca May 5, 2026 ▶ 35:56 🔬How GPT‑5 derived new results in theoretical physics and quantum gravity — Alex Lupsasca, OpenAI
Assertion Not checkable as stated
ChatGPT Solved Open Physics Problem Before Collaborator's Flight Landed
“We decided to start working on it using AI a little bit before Andy was scheduled to come, like the week before. And in fact, using ChatGPT, we solved the problem before he even got off the plane.”
Alex Lupsasca May 5, 2026 ▶ 20:53 🔬How GPT‑5 derived new results in theoretical physics and quantum gravity — Alex Lupsasca, OpenAI
What-if
Parakhin: Liquid AI could beat frontier models with equal compute
“I think if they if they had similar level of compute, they would be very competitive and maybe even beat the largest models, at least from what I've seen.”
Mikhail Parakhin Apr 22, 2026 ▶ 1:06:01 AI-Native Engineering: 100% adoption, 5x search throughput, unlimited tokens — Mikhail Parakhin
Opinion
Lopopolo: Bearish on MCP due to forced token injection and compaction issues
“MCPs I'm pretty bearish on because the harness forcibly injects all those tokens in the context and I don't really get a say over it. They mess with auto compaction. The agent can forget how to use the tool. There's probably only like, what, three calls in Pla…”
Ryan Lopopolo Apr 7, 2026 ▶ 38:37 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Opinion
Lopopolo: Coding models have largely solved all tasks except hard and new
“And I think things that are hard and new is still something that the models need humans. Yeah. Drive. Yeah. But I think those other quadrants are largely solved, given the right scaffold and the right thing that's going to drive the agent to completion.”
Ryan Lopopolo Apr 7, 2026 ▶ 33:46 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Opinion
Lopopolo: Coding models and harnesses are now isomorphic to human engineering capability
“The models are there enough. The harnesses are there enough where they're isomorphic to me and capability and the ability to do the job.”
Ryan Lopopolo Apr 7, 2026 ▶ 4:02 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Opinion
OpenAI Has Effectively Won the Consumer AI Market
“Now that means we're at a point in consumer where, maybe this is too early to say, but OpenAI has kind of won, right? Like, How do you catch up to something where model quality is not going to be differentiated? You already have the users, you already have the…”
Deedy Das Nov 14, 2025 ▶ 31:26 Anthropic, Glean & OpenRouter: How AI Moats Are Built with Deedy Das of Menlo Ventures
Prediction Didn’t hold up
Swix: OpenAI will issue a cryptocurrency token to fund compute
“There is still one more shoe to drop, which is the non sovereign wealth funding that open AI needs to get, which they've promised to drop by the end of this year. And my money is on, they have to do a coin. Like it's, I'm not a crypto guy at all, but like, y…”
Shawn Wang Oct 16, 2025 ▶ 49:13 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
Opinion
Lenz: Model providers should not dictate enterprise AI policies
“Right now, if you're using a model, you're taking in their own policy. Even if I want to use GPT-OSS, I've taken in a lot of different policies about what to abstain from, what's considered dangerous and not dangerous, how I should behave, etc. And I don't thi…”
Barak Lenz Oct 11, 2025 ▶ 40:33 Building Jamba 3B: the tiny Hybrid Transformer State Space Reasoning Model - Barak Lenz, CTO of AI21
Opinion
Gorkem: Hosting LLMs is a bad business due to Google search competition
“Language models, hosting language models is not a good business. At the time we thought, okay, we are going to be competing against OpenAI and Anthropic and all these labs. Turned, turned out that it was even worse because the killer application of language mo…”
Gorkem Yurtseven Sep 8, 2025 ▶ 8:50 A Technical History of Generative Media
Insight
Embiricos: Scaffolding-heavy AI agents are limited by developers' mental capacity
“A lot of, like, agents that I see are really impressive, but it's basically, like, part of what's impressive is it's like a bunch of developers building this, like, really bespoke state machine around a bunch of, like, short model calls, and so then the upper …”
Alexander Embiricos May 16, 2025 ▶ 30:37 ChatGPT Codex: The Missing Manual
Opinion
Goyal: Running LLM workloads at scale is impractical outside OpenAI
“It's just not practical outside of OpenAI to run use cases at scale in a lot of cases. Like, you can do it, but it requires quite a bit of work. And Because OpenAI is so good at making their models so available, I think they get a lot of credit for the science…”
Ankur Goyal Oct 11, 2024 ▶ 1:27:25 Production AI Engineering starts with Evals
Prediction Not checkable as stated
Goyal: OpenAI o1 will make agentic frameworks obsolete
“And I think O-one is going to do that to agentic frameworks as well. Hey, I think To me, it seems very unlikely that the, you know, you and me sort of like sipping an espresso and thinking about how, like, different personified roles of people should interact …”
Ankur Goyal Oct 11, 2024 ▶ 1:34:00 Production AI Engineering starts with Evals
Opinion
Goyal: Open-source inference providers are far less reliable than OpenAI
“They are nowhere near as reliable as, I mean, every single time I use any of those products and run a benchmark, I find a bug, text the CEO, and they fix something. It's nowhere near where OpenAI is.”
Ankur Goyal Oct 11, 2024 ▶ 1:37:24 Production AI Engineering starts with Evals
Prediction Open · timeframe Dec 2026
Patel: OpenAI and Microsoft partnership will likely collapse within years
“Yeah, I expect in the next few years that the OpenAI and Microsoft probably falls apart too.”
Dylan Patel Dec 5, 2023 ▶ 46:55 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Assertion Not checkable as stated
Royzen: GPT-4 was trained on HumanEval, proving data contamination
“GPT-IV itself has been trained on human eval, and we know this because GPT-IV is able to predict the exact doc string in many of the problems. I've seen it predict, like, the specific example values in the doc string, which is extremely improbable for it to ju…”
Michael Royzen Nov 3, 2023 ▶ 41:31 Beating GPT-4 with Open Source Models - with Michael Royzen of Phind
Assertion Not publicly verifiable
Howard: Alec Radford Built OpenAI's GPT After Reading ULMFiT
“I organized a chat for both of us with Kate Metz in the New York Times, and Kate Metz answered, sorry, and Alec answered this question for Kate, and Kate just like, so how did, you know, GPT come about? And he said, well, I was pretty sure that pre-training on…”
Jeremy Howard Oct 20, 2023 ▶ 15:41 The End of Finetuning — with Jeremy Howard of Fast.ai
Assertion Supported
Lie: Cerebras runs OpenAI's flagship model 14x faster than GPUs
“We're running you know, frontier level, one of the most intelligent models, right? OpenAI's largest, most capable, most intelligent model right now at 14 times faster than their normal, you know, GPU speeds.”
Sean Lie Sep 2, 2026 ▶ 8:37 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Assertion Supported
Kant: Major AI labs did not prioritize RL for LLMs three years ago
“And the second was that reinforcement learning was going to be the biggest driver for LLM capabilities. Today, very obvious three years ago was not an opinion held or direction held at either OpenAI or Google or Anthropic or others.”
Eiso Kant Jul 22, 2026 ▶ 6:03 The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
Opinion
Chen: Meta poaching has calmed down and OpenAI came out on top
“I think that met us calmed down a little bit. I think we came out on top”
Mark Chen Jun 25, 2026 ▶ 0:43 Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen
Prediction Not checkable as stated
Mark Chen: AI scaling laws will continue to hold
“And so I think it's just more and more of the same, right? Like more careful research engineering, more careful data engineering, more careful scaling, and it always unlocks that next ability to scale further. So I mean, it's held for You know, almost 10 order…”
Mark Chen Jun 25, 2026 ▶ 9:58 Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen
Opinion
Chen: Pre-training is not dead and remains underrated in AI research
“Well, I think if you still have a pre-training is dead view of the world I think pre-training is definitely yeah, yeah, not, not dead. It's underrated.”
Mark Chen Jun 25, 2026 ▶ 38:01 Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen
Assertion Not checkable as stated
OpenAI's Chen: AI models already discover novel theorems and advance sciences
“The initial direction we took was you should move it to real world research, right? And we've seen that the models, they've gotten a lot better at just kind of discovering novel theorems and pushing the frontiers of hard sciences. Even today, right, that's no …”
Mark Chen Jun 25, 2026 ▶ 7:27 Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen
Disclosure
Chen: OpenAI favors unifying modalities in as few architectures as possible
“For a research lab, I think there are a lot of advantages for it to being under one. So you just have to maintain one infrastructure stack, for instance. I think the cost to, like, maintaining and scaling many infrastructure stacks at once I think that's somet…”
Mark Chen Jun 25, 2026 ▶ 31:58 Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen
Assertion Not checkable as stated
ChatGPT Pro Generated Lupsasca's Exact Top Three Follow-Up Physics Questions
“You can take this page of this paper and you can feed it to ChatGPT Pro, say, like the best model we have out right now, and you can ask it, what should I do next? Give me the top three follow-up questions to ask based on this paper. I've done this experiment …”
Alex Lupsasca May 5, 2026 ▶ 1:18:37 🔬How GPT‑5 derived new results in theoretical physics and quantum gravity — Alex Lupsasca, OpenAI
Assertion Not checkable as stated
Terry Tao Says AI Math Proofs Merely Cite Obscure References
“I talked to Terry Tao a couple of weeks ago at UCLA. We had an OpenAI event with IPAM, which is this Institute of Mathematics there. And I talked to Terry Tao and he said that in his view, all of the proofs that he's seen AI come up with in math, even the ones…”
Alex Lupsasca May 5, 2026 ▶ 1:09:35 🔬How GPT‑5 derived new results in theoretical physics and quantum gravity — Alex Lupsasca, OpenAI
Assertion Not checkable as stated
Open-source AI demand spiked on hype before reverting to frontier labs
“Like all the open source models, I think what happened was they got like very hyped and people were very interested in using them. But I think like over time, like there was a spike in usage for these models. And then it goes back to open AI, Anthropic and Goo…”
Yasser Elsaid May 2, 2026 ▶ 16:45 ⚡️ Competing with ChatGPT and Sierra, building a $10M ARR company — Yasser Elsaid, Founder, Chatbase
Disclosure
Lopopolo: OpenAI Frontier team operates with post-merge or zero human code review
“You know, we, we've moved beyond even the humans reviewing the code as well. Most of the human review is post merge at this point, but it's not even reviewed.”
Ryan Lopopolo Apr 7, 2026 ▶ 9:21 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Insight
Lopopolo: AI-amplified small teams require extreme package decomposition and strict boundaries
“The structure of the repository is like, 500 NPM packages. It's like architecture to the access for what you would consider, I think, normal for a seven person team. But if every person is actually, like, 10 to 50. Then the, like, numbers on, like, being super…”
Ryan Lopopolo Apr 7, 2026 ▶ 39:56 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Insight
Lopopolo: Converting UI Images to ASCII Art Improves AI Agent Layout Perception
“If we want to actually, like, make it see the layout, it's almost easier to rasterize that image to ASCII arc and feed it in to the agent.”
Ryan Lopopolo Apr 7, 2026 ▶ 47:57 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Insight
Lopopolo: AI models can in-house 2,000-line dependencies in an afternoon
“The level of complexity of the dependencies that we can internalize is I would say low medium right now, right? Just based on model capability. What is medium? I would say like a couple thousand line dependency is a thing that we could in house no problem in a…”
Ryan Lopopolo Apr 7, 2026 ▶ 28:25 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Disclosure
Lopopolo: Symphony discards failed PRs entirely to regenerate from scratch
“In Symfony, there's this like rework state where once the PR is proposed and it's escalated to the human for review, it should be a cheap review, right? It is either mergeable or it is not. And if it's not, you move it to rework. The Elixir service will comple…”
Ryan Lopopolo Apr 7, 2026 ▶ 36:53 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Assertion Not checkable as stated
Lopopolo: Zero-code harness was 10x slower initially before outperforming any single engineer
“Honestly, the first month and a half was 10 times slower than I would be. But because we paid that cost, we ended up getting to something much more productive than any one engineer could be, because we built the tools, the assembly station for the agent to do …”
Ryan Lopopolo Apr 7, 2026 ▶ 5:17 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Insight
Lopopolo: Autonomous Coding Removes Human Language Familiarity Constraints
“No humans in the loop here. So like my, Own personal ability to write or not write Elixir doesn't really have to bias us away from using the right tool for the job, which is just wild.”
Ryan Lopopolo Apr 7, 2026 ▶ 47:38 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Opinion
Manning: OpenAI's Sora cannot produce compelling gameplay or persistent mechanics
“Don't think you can take Sora and produce compelling gameplay, right? If you want to have a world that you can wander around in a bit, you're good, but what are your abilities to have gameplay mechanics implemented the way you'd like them to be, and to have th…”
Chris Manning Apr 2, 2026 ▶ 52:02 Moonlake: Interactive, Multimodal World Models — with Chris Manning and Fan-yun Sun
Assertion Supported
Reddy: Voxtral speech model is much stronger than Whisper
“And I think a big people, I think there's a big rich ecosystem of people finding whisper and people want the same thing with Voxer. It's much stronger than whisper.”
Pavan Kumar Reddy Mar 30, 2026 ▶ 26:34 Mistral: Voxtral TTS, Forge, Leanstral, & Mistral 4 — w/ Pavan Kumar Reddy & Guillaume Lample
Disclosure
Pydantic built a VIP issue scorer after closing an OpenAI founder's ticket
“Basically this started off because one of the OpenAI co-founders created an issue on Pydantic. And we just closed it and said it was wrong. And so we have this that, like, injects itself and tries to summarize someone and it gives them a, like, brutal score of…”
Samuel Colvin Mar 14, 2026 ▶ 13:45 ⚡️Monty: the ultrafast Python interpreter by Agents for Agents — Samuel Colvin, Pydantic
Assertion Not checkable as stated
Glaese: OpenAI no longer trusts further score improvements on SWE-bench Verified
“Issues with the benchmark that means that now that we're at like 80%, we don't really trust like further improvements on it, but like it does measure something that is like a real like capability of models.”
Mia Glaese Feb 23, 2026 ▶ 14:34 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
Assertion Supported
Watkins: Over half of SWE-bench problems investigated by OpenAI had test flaws
“In over half of the problems that were investigated in that deep dive, there was one problem or the other. I think the most common problem are, like, overly narrow tests where there's some particular implementation detail that the tests were looking for but wa…”
Olivia Watkins Feb 23, 2026 ▶ 7:26 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
Assertion Not checkable as stated
Watkins: SWE-bench Verified is contaminated across OpenAI, Claude, and Gemini models
“And in SweetBenchVerified, we found many instances of contamination across like, across OpenEye models, across, like, Quad Opus, 4.5, Gemini Flash, and all of these, we saw things like regurgitating the ground truth solutions, things like in some cases giving,…”
Olivia Watkins Feb 23, 2026 ▶ 11:54 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.