The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 4 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Supported
OpenAI o3 model completes 40% of internal research engineer pull requests
“That's another data point, by the way, from this was from the O three system card. They showed a jump from like low to mid single digits to roughly 40% of PRs actually checked in by Research engineers at OpenAI that the model could do. So prior to O three, not…”
Nathan Labenz Oct 14, 2025 ▶ 38:26 Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question
Assertion Supported
Labenz: GPT-4.5 achieved 65% accuracy on SimpleQA versus o3's 50%
“The O-three class of models got about a 50% on that benchmark, and GPT 4.5 popped up to like 65%. So, in other words, it basically, of the things that were not known to the previous generation of models, it picked up a third of them.”
Nathan Labenz Oct 14, 2025 ▶ 8:58 Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question
Assertion Not checkable as stated
Chen: Earlier Codex models spent too little time on hard problems
“What we found is the latest, the previous generation of the codex models, they were spending too little time solving the hardest problems and too much time solving the easy, easy problems. And I think that, that is actually just probably out of the box what yo…”
Mark Chen Sep 25, 2025 ▶ 17:25 From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki
Disclosure
Sherman Wu: OpenAI's o3 model stands out for diligent tool execution
“One of my favorite models is actually O three. Cause it was like one of the most diligent models. It would just like do all these tool calls and it's like really the intelligence itself trying to like do the, you know, tool calls or reg or anything like that o…”
Sherman Wu Nov 28, 2025 ▶ 26:38 How OpenAI Builds for 800 Million Weekly Users: Model Specialization and Fine-Tuning
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.