The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 8 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Supported
Scialom: Llama 3 405B is the best open-source model ever released
“At a high level, it's the best open source model ever. It's Better than GPT-IV. I mean, what version? But, by far, compared to the version originally released even now, I think there's maybe the last cloud Sonya FF-V and GPT-IV-Zero that are performing it.”
Thomas Scialom Jul 23, 2024 ▶ 37:43 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Assertion Supported
Molmo 72B Beats Proprietary Models on Academic Benchmarks
“Academic benchmarks wise, Their big one is the best state-of-the-art everything, better than proprietary, but ELO-wise, it sits behind four-oh.”
Vibhu Sapra Oct 13, 2024 ▶ 45:42 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Assertion Supported
GPT-4.1 reduces extraneous edit rate to 2%, down from GPT-4o's 9%
“And we found that from four O, which got nine percent, which is pretty crazy, nine percent of the time making an extraneous edit is a lot. 4.1 is at two percent, so it's a pretty big improvement.”
Michelle Pokrass Apr 15, 2025 ▶ 23:43 GPT 4.1: The New OpenAI Workhorse
Assertion Supported
Nikunj Handa: OpenAI distilled o-series models into GPT-4o search
“They use, like, synthetic data techniques. They've done, like, O-series model distillation to, like, make these four or fine tunes really good.”
Nikunj Handa Mar 11, 2025 ▶ 9:26 The new OpenAI Agents Platform: CUA, Web Search, Responses API, Agents SDK!!
Assertion Supported
Cosine's fine-tuned GPT-4o scored higher on SWE-bench than OpenAI's original o1
“And then we also on this podcast, we interviewed Cosign that actually fine tune four O to, on three bench, sorry to achieve data on three bench. And that score was actually higher than O one when it came out.”
Shawn Wang Jan 24, 2025 ▶ 19:35 The Unreasonable Effectiveness of Reasoning Distillation: using DeepSeek R1 to beat OpenAI o1
Assertion Supported
Swix: Multi-sampling GPT-4o mini before GPT-4o judging yields net savings
“If I call a GP for a mini 10 times and I do a number of drafts or summaries, and then I have four, oh, judge the summaries that actually is net savings and like a good enough savings then running four, oh, on everything, which given the hundreds and thousands …”
Shawn Wang Sep 20, 2024 ▶ 44:23 The Ultimate Guide to Prompting - with Sander Schulhoff from LearnPrompting.org
Assertion Supported
Structured response format is limited to GPT-4o and GPT-4o mini
“Actually, the new response format is only available on two models. It's Foro Mini and the new Foro. So the old Foro doesn't have the new response format. However, for function calling, we were able to enable it for all models that support function calling, and…”
Michelle Pokrass Sep 17, 2024 ▶ 30:23 Building AGI with OpenAI's Structured Outputs API
Assertion Partly supported
Alessio Fanelli: GPT-4o Search jumps to 90% accuracy on simple QA
“On simple QA, GPT four O is 30% accuracy. Four O search is 90%.”
Alessio Fanelli Mar 11, 2025 ▶ 8:12 The new OpenAI Agents Platform: CUA, Web Search, Responses API, Agents SDK!!
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.