The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 8 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Prediction Not checkable as stated
Future Llama models will match frontier GPT models if progress slows
“As AI progress slows down, so if we get like Llama-IV, Llama-V for example, maybe it's a comparable at that point, like GPT-V or GPT-VI, like, It made it to the point where it was like, look, I just want to use Lama. Like, it's, you know, safe for me to, you k…”
David Hsu Feb 7, 2024 ▶ 50:30 The State of AI in production — with David Hsu of Retool
Prediction Not checkable as stated
Royzen: The leap from GPT-4 to GPT-5 will be smaller
“I think that GPT-IV, my hypothesis is that the jump from four to 4.5, or four to five, will be smaller than the jump from Three to four.”
Michael Royzen Nov 3, 2023 ▶ 37:38 Beating GPT-4 with Open Source Models - with Michael Royzen of Phind
Assertion Contradicted
Frontier models are cheaper for agentic tasks because they require fewer turns
“Interestingly, in Tau Tau Two Bench Telecom, it's cheaper to run, you know, on a per token basis, more expensive models, like a GBD five, compared to some smaller open source models, because the some of the GBD five, for instance got to the answer faster. And …”
George Cameron Jan 9, 2026 ▶ 1:09:14 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Assertion Partly supported
Early GPT-5 testers report noticeable gains across coding, science, and writing
“And the story that we wrote, we kind of talked about how at least the people that we've talked to who tested it so far have been pretty impressed. They seem to think that it's been, you know, there's been improvements in a number of domains and both like scien…”
Stephanie Palazzolo Aug 6, 2025 ▶ 28:49 The AI Agenda: GPT5 leaks and the business of AI News — Steph Palazzolo, The Information
Insight
Byun: Foundational models will not commoditize Elicit due to deep workflow specialization
“I think about this a lot in the context of moats. People are like, oh, what's your moat? What happens if GPT-V comes out? It's like, if GPT-V comes out, there's still like all of this other space that we can go into. And so I think being really obsessed with t…”
Jungwon Byun Apr 11, 2024 ▶ 28:57 Supervise the Process of AI Research — with Jungwon Byun and Andreas Stuhlmüller of Elicit
Assertion Partly supported
McGrath: GPT-5.1 dramatically reduced token usage over GPT-5 while boosting evals
“Yeah, and so you can see, like, from five to 5.1, our overall evals, you know, we bumped some. But if you look at a two D plot of how many tokens it takes for us to get that, it went way down.”
Josh McGrath Dec 31, 2025 ▶ 13:58 [State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI
Disclosure
Fioca: OpenAI evaluates GPT-5 coding models on behavioral software engineering practices
“And so these are just best software engineering practices that turn out to be behavior characteristics, and we can measure the model's performance on those behaviors and grade it that way.”
Brian Fioca Dec 26, 2025 ▶ 4:00 ⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust — Brian Fioca + Bill Chen, OpenAI
Assertion Supported
Fioca: GPT-5 matches Codex coding capability but adds step-by-step preambles
“With the five series, because it's more general, and it's just about as good as coding as codex for a lot of things. We've taught it to be more communicative. And so it has preambles before tool calls. It'll say things like, I'm about to go look for this.”
Brian Fioca Dec 26, 2025 ▶ 10:39 ⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust — Brian Fioca + Bill Chen, OpenAI
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.