The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 14 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Supported
Chollet: 50,000x LLM scaling yielded flat progress on ARC benchmark
“Because between, like, GPT-II and GPT-IV. There's been this 50,000 X scale up of base models that has resulted in, in, in basically a flat curve. On Arc.”
Francois Chollet Apr 3, 2025 ▶ 11:44 Chasing Real AGI: Inside ARC Prize 2025 with Chollet & Knoop
Assertion Not checkable as stated
Chollet: GPT-4 lacks fluid intelligence, but OpenAI's o3 model has it
“GPT-IV does not have fluid intelligence, for instance, but O-III does.”
Francois Chollet Apr 3, 2025 ▶ 5:00 Chasing Real AGI: Inside ARC Prize 2025 with Chollet & Knoop
Prediction Not checkable as stated
Chollet: Commercial AI models will increasingly adopt test-time search architectures
“Increasingly, you're gonna see commercial models that use test-time search, where instead of just trying to generate one single COT to adapt to the task, they're actually gonna run through this, you know, search.”
Francois Chollet Apr 3, 2025 ▶ 9:47 Chasing Real AGI: Inside ARC Prize 2025 with Chollet & Knoop
Assertion Supported
Chollet: Latest base LLMs score zero percent on ARC-AGI-2
“Today the latest base alarms, they're doing something like 10% on ARK-I. But on Arc two, they are doing zero percent.”
Francois Chollet Apr 3, 2025 ▶ 10:50 Chasing Real AGI: Inside ARC Prize 2025 with Chollet & Knoop
Insight
Chollet: Intelligence should be defined as skill acquisition efficiency
“And yeah, so I, to summarize that, you know, I see intelligence as skill acquisition efficiency. So it's not the fact that you can acquire skills, it's how efficiently You can do it. That's a measure of your intelligence.”
Francois Chollet Apr 3, 2025 ▶ 16:06 Chasing Real AGI: Inside ARC Prize 2025 with Chollet & Knoop
Opinion
Chollet: OpenAI o3 is the most advanced test-time adaptation model
“And OSTRI best I can tell is the most advanced the most successful test and adaptation model out there at this time.”
Francois Chollet Apr 3, 2025 ▶ 17:53 Chasing Real AGI: Inside ARC Prize 2025 with Chollet & Knoop
Insight
Chollet: AGI means humans can no longer easily create tasks AI fails
“You have AGI when it's no longer possible to easily come up with tasks that, you know, you and I can do naturally, but no AI system can do.”
Francois Chollet Apr 3, 2025 ▶ 49:37 Chasing Real AGI: Inside ARC Prize 2025 with Chollet & Knoop
Prediction Not checkable as stated
Chollet: NDEA will solve problems previously unsolved by humans in verifiable domains
“The kind of technology we're building on, it's differential advantage that it's going to be capable of solving problems that have never been solved by humans before. That's, you know, that's a very different deal than LLMs, for instance, but effectively only i…”
Francois Chollet Apr 3, 2025 ▶ 56:57 Chasing Real AGI: Inside ARC Prize 2025 with Chollet & Knoop
Assertion Not checkable as stated
Chollet: OpenAI o3 cost $10k–$20k per ARC puzzle on maximum compute
“For instance OpenAI O.S. On the highest compute settings that we tried it on for Arc, it was consuming somewhere between, like, 10,000 dollars to 20,000 dollars per task, like, for one little puzzle, which you could normally solve with a base of an API for a f…”
Francois Chollet Apr 3, 2025 ▶ 20:00 Chasing Real AGI: Inside ARC Prize 2025 with Chollet & Knoop
Assertion Not checkable as stated
Chollet: OpenAI used about 75% of ARC training tasks to adapt o3
“So they told us that they were using a significant fraction, I think they said something like 75%, of the training tasks to, you know, to adapt the model in some way.”
Francois Chollet Apr 3, 2025 ▶ 20:48 Chasing Real AGI: Inside ARC Prize 2025 with Chollet & Knoop
Assertion Supported
Chollet: MindAI dropped out of ARC Prize over open-source requirements
“They ended up dropping out because they did not want to open source their solution. And of course, that meant that they were not eligible for the prize.”
Francois Chollet Apr 3, 2025 ▶ 42:03 Chasing Real AGI: Inside ARC Prize 2025 with Chollet & Knoop
Assertion Supported
Chollet: Average human test score on ARC-AGI-2 is about 60%
“Based on our own testing, an average person in our test sample would score about 60%.”
Francois Chollet Apr 3, 2025 ▶ 46:17 Chasing Real AGI: Inside ARC Prize 2025 with Chollet & Knoop
Disclosure
Chollet: Work has begun on ARC-AGI-3 featuring a brand-new format
“We're already starting to work on version three which you have a brand new format.”
Francois Chollet Apr 3, 2025 ▶ 50:31 Chasing Real AGI: Inside ARC Prize 2025 with Chollet & Knoop
Insight
Chollet: Program synthesis bottleneck is the costly million-point search space
“The main bottleneck is that this search process takes a very, very long time. It's a very large search space, and to evaluate all these points you know, which point takes you some amount of competition to evaluate. You're gonna have to evaluate millions of poi…”
Francois Chollet Apr 3, 2025 ▶ 27:25 Chasing Real AGI: Inside ARC Prize 2025 with Chollet & Knoop
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.