Mike Knoop is discussing early evaluation scores on the newly released ARC-AGI-2 benchmark designed to test novel visual reasoning.
Prediction Not checkable as stated
Knoop: Scaling Test-Time Compute Will Not Get Us to AGI
“And then there's a new story that's emerged over the last like five months, which is, oh, we're going to scale up this test time compute and that's going to get us to AGI. And I think what V two shows is that that's not quite either. We still need some structu…”
Opinion
Knoop: Language models operate by memorization rather than solving novel patterns
“Language models. Generally working like a memorization style regime where they're right. Learning lots of data. They're able to apply it to very similar types of patterns that they've seen before, but not novel patterns. That's what RKGI shows.”
Assertion Supported
Knoop: ARC Saw No Progress Despite 50,000x Model Scaling
“Surprise that it basically hadn't, and not only hadn't been beaten, there'd basically been no progress in it which I thought was really fascinating given the fact that we've like scaled up these language model systems by almost like 50,000 times over the last,…”
Prediction Not checkable as stated
Knoop: AI agents will start working in 2025 due to ARC progress
“I actually think we're going to start to see agents start to work this year specifically because of progress on arc.”
Insight
Knoop: Startups just doing model training are lighting money on fire
“I think anyone who's like Just doing model training at this point is like lighting money on fire. If you really want to make a unique difference, especially if you're a small startup, like a founder, like you gotta go take an orthogonal approach. You gotta try…”
Opinion
Knoop: Gary Marcus has been more right than wrong on deep learning
“I generally think he's been more right than wrong. I think if you like just take a limited five year view on this from 20, 20 up until 20, 20, end of 20, 24, you know, I think it was a generally right. Like he was making the right ideas.”