The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 8 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Opinion
GPT-4o Mini cannot accurately evaluate complex Claude 3.7 outputs
“GPD for a mini can't find, ah, cannot evaluate the hard outputs that's 3.7 might be doing correctly, right?”
Pratik Bhavsar Jul 14, 2025 ▶ 23:12 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
Opinion
OpenAI models remain the industry best for chained function calling
“And so we find like the open AI models function calling wise for our use case. So for the best ones and like, yeah, basically verifying the others. We had a beta group and testing different models. So basically from all providers and yeah, find basically the b…”
Thomas Paul Mann Feb 26, 2025 ▶ 10:21 Raycast: Your AI Automation Assistant
Disclosure
Raycast built its Ray 1 models by fine-tuning GPT-4o and mini
“And so we looked into all the various models we had and then we picked, at the moment, it's gbd-for-o and gbd-for-o-mini, Which we basically did a fine tune to really optimize for our use case, and then basically shipping that in the app as Ray one and Ray one…”
Thomas Paul Mann Feb 26, 2025 ▶ 8:15 Raycast: Your AI Automation Assistant
Assertion Supported
Swyx: Distilling GPT-4 to Mini cut costs 15x with 2% hit
“Yeah, I sat in the distillation session just now, and they showed how they distilled from four to four mini, and it was like only like a two percent hit in the performance, and 15 X cheaper.”
Shawn Wang Oct 4, 2024 ▶ 48:14 Building AGI in Real Time (OpenAI Dev Day 2024)
Assertion Supported
Swix: Multi-sampling GPT-4o mini before GPT-4o judging yields net savings
“If I call a GP for a mini 10 times and I do a number of drafts or summaries, and then I have four, oh, judge the summaries that actually is net savings and like a good enough savings then running four, oh, on everything, which given the hundreds and thousands …”
Shawn Wang Sep 20, 2024 ▶ 44:23 The Ultimate Guide to Prompting - with Sander Schulhoff from LearnPrompting.org
Assertion Supported
Structured response format is limited to GPT-4o and GPT-4o mini
“Actually, the new response format is only available on two models. It's Foro Mini and the new Foro. So the old Foro doesn't have the new response format. However, for function calling, we were able to enable it for all models that support function calling, and…”
Michelle Pokrass Sep 17, 2024 ▶ 30:23 Building AGI with OpenAI's Structured Outputs API
Assertion Not checkable as stated
GPT-4.1 Mini significantly outperforms 4o Mini, nearing original GPT-4o performance
“4.1 mini is actually quite significantly better than four o mini and not that far away from the old four o.”
Michelle Pokrass Apr 15, 2025 ▶ 32:15 GPT 4.1: The New OpenAI Workhorse
What-if
Shreya Shankar: GPT-4o Mini reduces DocETL optimization cost by 90%
“The reason it was a hundred dollars, if I ran the optimizer with GPT-Foro mini as the LLMs, it would be 10 dollars. But we use GPT four. Oh, just because I think we did this at a time where many hadn't come out yet.”
Shreya Shankar Nov 29, 2024 ▶ 45:58 [Paper Club] DocETL: Agentic Query Rewriting + Eval for Complex Document Processing w Shreya Shankar
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.