The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 8 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Supported
Chollet: Scaling LLMs 50,000x Only Lifted ARC Accuracy to 10%
“And from at the time, back in 2019 to now, with a model like GPT 4.5, for instance, there's been a roughly 50,000 X scale up of basal alarms. And we went from zero percent accuracy on that benchmark to roughly 10%, which is not a lot.”
François Chollet Jul 3, 2025 ▶ 2:09 François Chollet: How We Get To AGI · Y Combinator
Assertion Supported
Chollet: Base LLMs score 0% and static reasoning scores 1-2% on ARC-2
“Well, if you take Bazel Alums, model Slack, GPT-IV-IV-V, LAMA-IV, it's simple, they get zero percent. There is simply no way to do these tasks simply via memorization. Next, if you look at static reasoning systems, so systems that use a single chain of tasks t…”
François Chollet Jul 3, 2025 ▶ 17:36 François Chollet: How We Get To AGI · Y Combinator
Assertion Supported
Chollet: SOTA test-time adaptation takes thousands in compute to solve ARC-1
“Even the latest set of the art CTA techniques they still need thousands of dollars of compute to solve arc one at human level. And that doesn't even scale to arc two.”
François Chollet Jul 3, 2025 ▶ 24:28 François Chollet: How We Get To AGI · Y Combinator
Assertion Supported
Chollet: Base LLMs Score Under 10% on ARC-AGI-1
“So basal alarms were scoring extremely low on V-one, like sub-ten percent, basically. And, I mean, it was true of the original, like, GPT-III actually scoring zero, but that's even true of the latest basal alarms today, you know, as of March.”
François Chollet Mar 27, 2026 ▶ 17:57 François Chollet: Why Scaling Alone Isn’t Enough for AGI · Y Combinator
Assertion Supported
Chollet: Fine-Tuned OpenAI o3 Reached Human-Level Performance on ARC
“So in particular, in December last year, OpenAI previewed its, ah, all three model, and they used a version of it that was, ah, fine-tuned specifically on Arc, and that showed human-level performance on that benchmark versus time.”
François Chollet Jul 3, 2025 ▶ 3:29 François Chollet: How We Get To AGI · Y Combinator
Assertion Supported
Chollet: Every High-Performing ARC AI Method Uses Test-Time Adaptation
“And today, every single AI approach that performs well on Arc is using one of these techniques.”
François Chollet Jul 3, 2025 ▶ 4:13 François Chollet: How We Get To AGI · Y Combinator
Assertion Supported
Chollet: All ARC-3 environments are solvable by untrained humans
“All of these test environments in Arc three Are solvable by humans with no prior training because we actually tested them on, on regular people.”
François Chollet Mar 27, 2026 ▶ 27:52 François Chollet: Why Scaling Alone Isn’t Enough for AGI · Y Combinator
Assertion Partly supported
Chollet: Compute costs have fallen two orders of magnitude per decade since 1940
“The cost of compute has been consistently falling by two orders of magnitude every decade since 1940.”
François Chollet Jul 3, 2025 ▶ 0:13 François Chollet: How We Get To AGI · Y Combinator
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.