Knoop: ARC Saw No Progress Despite 50,000x Model Scaling
“Surprise that it basically hadn't, and not only hadn't been beaten, there'd basically been no progress in it which I thought was really fascinating given the fact that we've like scaled up these language model systems by almost like 50,000 times over the last,…”
Chollet: Commercial AI models will increasingly adopt test-time search architectures
“Increasingly, you're gonna see commercial models that use test-time search, where instead of just trying to generate one single COT to adapt to the task, they're actually gonna run through this, you know, search.”
Chollet: Latest base LLMs score zero percent on ARC-AGI-2
“Today the latest base alarms, they're doing something like 10% on ARK-I. But on Arc two, they are doing zero percent.”
Chollet: 50,000x LLM scaling yielded flat progress on ARC benchmark
“Because between, like, GPT-II and GPT-IV. There's been this 50,000 X scale up of base models that has resulted in, in, in basically a flat curve. On Arc.”
Chollet: Average human test score on ARC-AGI-2 is about 60%
“Based on our own testing, an average person in our test sample would score about 60%.”
Coogan: Keras creator François Chollet is leaving Google to start a company
“Francois Chollet is out at Google, entering free agency. He's gonna start a new company.”
Knoop: AGI is properly defined as efficient skill acquisition
“Francois definition, which is the one that I think is the right one is this definition that general intelligence is a system that can effectively, efficiently acquire new skill. That's it efficiently acquiring new skill and being able to solve these open-ended…”