why aren't all 8 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Supported
Chollet: Scaling LLMs 50,000x Only Lifted ARC Accuracy to 10%
“And from at the time, back in 2019 to now, with a model like GPT 4.5, for instance, there's been a roughly 50,000 X scale up of basal alarms. And we went from zero percent accuracy on that benchmark to roughly 10%, which is not a lot.”
Assertion Supported
Chollet: Base LLMs score 0% and static reasoning scores 1-2% on ARC-2
“Well, if you take Bazel Alums, model Slack, GPT-IV-IV-V, LAMA-IV, it's simple, they get zero percent. There is simply no way to do these tasks simply via memorization. Next, if you look at static reasoning systems, so systems that use a single chain of tasks t…”
Assertion Supported
Chollet: SOTA test-time adaptation takes thousands in compute to solve ARC-1
“Even the latest set of the art CTA techniques they still need thousands of dollars of compute to solve arc one at human level. And that doesn't even scale to arc two.”
Assertion Supported
Chollet: Base LLMs Score Under 10% on ARC-AGI-1
“So basal alarms were scoring extremely low on V-one, like sub-ten percent, basically. And, I mean, it was true of the original, like, GPT-III actually scoring zero, but that's even true of the latest basal alarms today, you know, as of March.”
Assertion Supported
Chollet: Fine-Tuned OpenAI o3 Reached Human-Level Performance on ARC
“So in particular, in December last year, OpenAI previewed its, ah, all three model, and they used a version of it that was, ah, fine-tuned specifically on Arc, and that showed human-level performance on that benchmark versus time.”
Assertion Supported
Chollet: Every High-Performing ARC AI Method Uses Test-Time Adaptation
“And today, every single AI approach that performs well on Arc is using one of these techniques.”
Assertion Supported
Chollet: All ARC-3 environments are solvable by untrained humans
“All of these test environments in Arc three Are solvable by humans with no prior training because we actually tested them on, on regular people.”
Assertion Partly supported
Chollet: Compute costs have fallen two orders of magnitude per decade since 1940
“The cost of compute has been consistently falling by two orders of magnitude every decade since 1940.”