The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 13 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Partly supported
Sivulka: GPT-4 beat BloombergGPT at every financial task
“So then GPT-IV was released, I think, like a few weeks later. I don't know exactly any of the right timeline, but it just destroyed Bloomberg GPT at every single finance task.”
George Sivulka Jan 22, 2025 ▶ 39:13 George Sivulka, Co-Founder & CEO @Hebbia: The Future of Foundation Models | E1250 · 20VC with Harry Stebbings
Prediction Held up
Altman: Competing AI labs will successfully replicate OpenAI's o1 model
“After, after a research lab does something, even if you don't know exactly how they did it, it's, I won't say easy, but it's doable to go off and copy it, and you can see this in the replications of GPT-IV, and I'm sure you'll see this in replications of O-one…”
Sam Altman Nov 4, 2024 ▶ 15:48 Sam Altman: What Startups Will be Steamrolled by OpenAI & Where is Opportunity | E1223 · 20VC with Harry Stebbings
Assertion Partly supported
Gomez: 13-billion-parameter models now outperform original 1.7-trillion GPT-4
“GPT-IV, if it's true what they say, and it's 1.7 trillion parameters, this big MOE, we have models that are better than that model that are like, thirteen billion parameters.”
Aidan Gomez Aug 19, 2024 ▶ 5:38 Aidan Gomez: What No One Understands About Foundation Models | E1191 · 20VC with Harry Stebbings
Assertion Partly supported
Wang: JP Morgan's internal data is 150 petabytes versus GPT-4's sub-petabyte dataset
“JP Morgan's proprietary internal data set is a 150 petabytes. The GPT-IV was trained on an internet data set that was less than one petabyte.”
Alex Wang Jun 12, 2024 ▶ 8:18 Alex Wang: Why Data Not Compute is the Bottleneck to Foundation Model Performance | E1164 · 20VC with Harry Stebbings
Assertion Partly supported
Srinivas: GPT-4 convincingly beats BloombergGPT on finance benchmarks
“Bloomberg spent a lot of money training Bloomberg GPT... And that model is, is beaten convincingly by, like, a GPT-IV on all the finance benchmarks.”
Aravind Srinivas Jun 5, 2024 ▶ 8:47 Aravind Srinivas:Will Foundation Models Commoditise & Diminishing Returns in Model Performance|E1161 · 20VC with Harry Stebbings
Prediction Didn’t hold up
Socher: An open-source GPT-4 equivalent will arrive by end of 2023
“I predicted that we'll have a GPT-IV equivalent model Before the end of the year, that's open source.”
Richard Socher Nov 24, 2023 ▶ 8:42 The Ultimate AI Roundtable: What Happens Now in AI, Why Google are Vulnerable | E1085 · 20VC with Harry Stebbings
Prediction Didn’t hold up
Socher: Open-source GPT-4 equivalent model will launch before end of 2023
“I predicted that we'll have a GBD four equivalent model before the end of the year. That's open source. Of course, GBD four keeps getting better and better. So my prediction was for the version we had like a few months ago,”
Richard Socher Aug 18, 2023 ▶ 28:00 Richard Socher: The 3 Biggest Barriers to Building AGI; AI Startups vs Incumbents | E1050 · 20VC with Harry Stebbings
Assertion Supported
Acharya: Token cost for GPT-4 has dropped 100x since release
“The cost of actually a token on GPT-IV has, you know, gone down a hundred X since the model was released.”
Anish Acharya Feb 9, 2026 ▶ 58:52 a16z, Anish Acharya: Is SaaS Dead? Do Margins Still Matter? Why We Are Not in an AI Bubble? · 20VC with Harry Stebbings
Assertion Supported
Mollick: Unprompted GPT-4 math tutoring led to lower test scores
“The first randomized control trial we have, I have some of my colleagues at Wharton was giving GPT-IV people for math tutoring in Turkey. Now, they didn't do a huge amount of, like, you know, it was an assigned class, and they used the system, but it turns out…”
Ethan Mollick Jul 31, 2024 ▶ 48:39 Ethan Mollick: Why OpenAl Abandons Products, The Biggest Opportunities They Have Not Taken | E1184 · 20VC with Harry Stebbings
Assertion Supported
Hankes: OpenAI Delayed GPT-4 Release by Five Months for Safety Testing
“On GPT four, we saw the demo in the fall. They didn't just release the product then they took, I think four or five months To test and learn about safety in the edges of the model and then ultimately released it to the world.”
Vince Hankes May 3, 2023 ▶ 34:26 Vince Hankes: Why We Put $300M into OpenAI; Sam Altman's Pitch; Lessons from Josh Kushner | E1009 · 20VC with Harry Stebbings
Assertion Supported
Rauch: Llama is nowhere near as capable as GPT-4
“Right now, Llama is not as good as GPT-IV. Not even close.”
Guillermo Rauch Oct 6, 2023 ▶ 1:04:02 Guillermo Rauch: Why Great Companies are Defined by How Many Things They Say No To | E1069 · 20VC with Harry Stebbings
Assertion Supported
Lebrun: LIMA fine-tuned on 1,000 examples beats GPT-3, rivals GPT-4
“Three weeks ago, there was a paper about Lima. So, so this Lima paper shows that with only 1000 question and answer examples, so very, very small data sets they get something for, use for fine tuning, so the second stage, they get something that performs bette…”
Alex Lebrun Jun 19, 2023 ▶ 19:30 Alex Lebrun: Why the EU's AI Regulation is a Disaster; How Zuck Prepares for Meetings | E1027 · 20VC with Harry Stebbings
Assertion Supported
Duolingo was an launch partner for OpenAI's GPT-4 release
“We were one of the launch partners of OpenAI when they first launched GPT-IV”
Severin Hacker May 19, 2025 ▶ 3:19 Duolingo Co-Founder, Severin Hacker: How AI Impacts the Future of Work and Education · 20VC with Harry Stebbings
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.