The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 7 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Supported
Chollet: 50,000x LLM scaling yielded flat progress on ARC benchmark
“Because between, like, GPT-II and GPT-IV. There's been this 50,000 X scale up of base models that has resulted in, in, in basically a flat curve. On Arc.”
Francois Chollet Apr 3, 2025 ▶ 11:44 Chasing Real AGI: Inside ARC Prize 2025 with Chollet & Knoop
Prediction Held up
Ratner: Private data models will exceed closed models in specialized tasks
“Closed source models, like a GPT-IV, five, six, seven, whatever comes, are going to be very hard to match in terms of generalist capability for, say, consumer use cases that are reflected in the web data they're trained on and the flywheels that get powered by…”
Alex Ratner May 31, 2023 ▶ 8:11 Entering the Data-Centric Era of Foundation Models with Alex Ratner, Co-Founder & CEO of Snorkel AI
Assertion Supported
Azhar: Mistral matched GPT-4 quality far more computationally efficiently than US firms
“Mistral, which is this Parisian company has been doing some, you know, remarkable things, had done a couple, two things that I thought were really interesting. One was that they were able to get close to GPT-IV quality much more computationally efficiently tha…”
Azeem Azhar Jul 2, 2024 ▶ 43:49 From Business to Warfare: How AI Affects the Modern World | Azeem Azhar
Assertion Supported
Shah: Hippocratic AI outperformed GPT-4 on 105 of 114 healthcare exams
“And then we took it, and then we had GPT-IV take it, and we had all the other language models take it, and we beat them all. And we beat them on a 105 of a 114 for GPT-IV, for example.”
Munjal Shah Aug 23, 2023 ▶ 21:10 Hippocratic AI’s Munjal Shah: Building the First Safety-First LLM for Healthcare
Assertion Supported
Nvidia released an AI model that outperformed GPT-4
“Just a couple of weeks ago, Nvidia announced that In addition to having the best chips, they just released a model that was actually better than GPT-IV.”
Matt Turck Nov 8, 2024 ▶ 39:14 Superintelligence, Bubbles And Big Bets: AI Investing in 2024 | Matt Turck & Aman Kabeer, FirstMark
Assertion Supported
GPT-4 price per token dropped roughly 90% in one year
“The, I think the price per token of GPT-IV dropped something like 90%. Over the last year”
Matt Turck Nov 8, 2024 ▶ 29:26 Superintelligence, Bubbles And Big Bets: AI Investing in 2024 | Matt Turck & Aman Kabeer, FirstMark
Assertion Supported
Traynor: GPT-4 crossed the hallucination threshold required for customer support bots
“So Finn's built on GPT-IV, by the way, we've, we tried, we wanted to build on a three, 3.5, but it didn't, it's still, I remember back when we used to talk about hallucinations, like four was the sort of the perceptual change for us in terms of trust and relia…”
Des Traynor Feb 29, 2024 ▶ 18:40 How Intercom transitioned to being AI-first | Des Traynor, Co-Founder of Intercom
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.