The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 7 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Supported
Marcus: Models like o1 are not systematically better than GPT-4
“Models like O-I are not systematically better than GPT-IV. There's, they're better in certain use cases. Especially ones where you can create data in advance.”
Gary Marcus May 7, 2025 ▶ 11:46 Are We at the End of Ai Progress? — With Gary Marcus
Assertion Partly supported
DeepSeek used model distillation to match OpenAI's o1 at lower cost
“And, you know, they've just used the process of distillation to, you know, effectively bring those bigger versions of the sort of state-of-the-art models and distill them down into you know, smaller models, which eventually led to this R-one, you know, the equ…”
M.G. Siegler Jan 27, 2025 ▶ 12:00 How DeepSeek Changes AI Research & Silicon Valley w/ M.G. Siegler
Assertion Supported
Wang: China's DeepSeek produced the first replication of OpenAI's o1 model
“OpenAI released O-one and released the O-one preview a number of months ago... Yeah, this is OpenAI's advanced reasoning model, which is great at sort of scientific reasoning and mathematical reasoning and reasoning and code, et cetera. And the very first repl…”
Alexandr Wang Dec 11, 2024 ▶ 13:21 AI Predictions for 2025: Geopolitics, Agents, and Data Scaling — With Alexandr Wang
Prediction Partly held up
Patel: GPT-5 will simultaneously scale pre-training and post-training reasoning
“And so now GPT-Five, as Sam calls it, is, is gonna be a model that has huge pre-training scale, right? Like GPT-Five, but also huge post-training scale, Like O-one and O-three and continuing to scale that up, right? This would be the first time we see a model …”
Dylan Patel Apr 23, 2025 ▶ 36:50 Generative AI 101: Tokens, Pre-training, Fine-tuning, Reasoning — With SemiAnalysis CEO Dylan Patel
Assertion Supported
OpenAI o1 Benchmarks Show Major Gains Across Math, Coding, and Science
“So if you think about, ah, competition math, ah, GPT-IV-O was getting a 13.4 accuracy score on competition math. But this thing and the way that it can think through the different problems is getting 83.3, ah, score on accuracy. Ah, with competition math. So i…”
Alex Kantrowitz Sep 13, 2024 ▶ 1:41 Is OpenAI’s New “o1” Model The Big Step Forward We’ve Been Waiting For?
Assertion Supported
OpenAI Notably Avoided Using the Word 'Agent' in o1 Announcement
“Open AI hasn't actually used the word agent once in its announcement, and it doesn't sound like any of its scientists have talked about that either.”
Parmy Olson Sep 13, 2024 ▶ 21:49 Is OpenAI’s New “o1” Model The Big Step Forward We’ve Been Waiting For?
Assertion Supported
Kantrowitz: OpenAI's Blog Post States o1 Does Not Solve Hallucinations
“Yeah, and OpenAI even in its blog post says that it does not solve, this does not solve hallucinations, and then you can”
Alex Kantrowitz Sep 13, 2024 ▶ 24:49 Is OpenAI’s New “o1” Model The Big Step Forward We’ve Been Waiting For?
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.