The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 4 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Opinion
Morcos: GPT-4.5 and Llama 4 show limits of naive mega-model scaling
“And I think that's what we've seen to some extent with the failure of the mega models, right? With 4.5 and Lama four and others. I think that there is a challenge of just continuing to do that naively and you have to figure out how to break it.”
Ari Morcos Aug 29, 2025 ▶ 27:07 Better Data is All You Need — Ari Morcos, Datology
Opinion
Swyx: Frontier Models Exist Primarily to Distill Smaller, Usable Models
“Even GPT 4.5 is too expensive. Normally it's really gonna use it in, in any reasonable quantity. Like, you know, Claude 3.5 Opus, like if it does exist, still not like, you know, the thing that we actually use is Sonnet, right? So like, it's almost like a depl…”
Shawn Wang Mar 23, 2025 ▶ 3:58 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Assertion Supported
Noam Brown: GPT-4.5 makes Tic-Tac-Toe mistakes without System 2 reasoning
“With Tic-Tac-Toe, we see that, like, GPD-Four .5 falls over. You know, it plays decently well. I shouldn't say it falls over. It does reasonably well. You can draw the board. It can make legal moves, but it will make mistakes sometimes, and if you really need …”
Noam Brown Jun 19, 2025 ▶ 12:06 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Disclosure
OpenAI is working to inject GPT-4.5's humor and nuance into future models
“We're working on incorporating kind of those improvements into the models more generally. People loved about 4.5 is like the humor, the green text, the nuance. So we've heard that feedback and I know, yeah, there's lots of folks working on that and trying to b…”
Michelle Pokrass Apr 15, 2025 ▶ 41:01 GPT 4.1: The New OpenAI Workhorse
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.