Kantrowitz: AI benchmark success without economic impact reveals spiky intelligence
“If it can solve arc AGI and it's not necessarily crushing on these economic factors and these just kind of general rote work things that we would like it to do it shows that instead of being general, it's very spiky intelligence and hence much less useful.”
Hu: Anthropic Opus 5 achieved 30% on ARC-AGI
“You guys got, took Arc AGI three to 30%, which is incredible.”
Chollet: Base LLMs Score Under 10% on ARC-AGI-1
“So basal alarms were scoring extremely low on V-one, like sub-ten percent, basically. And, I mean, it was true of the original, like, GPT-III actually scoring zero, but that's even true of the latest basal alarms today, you know, as of March.”
Poetiq's autonomous prompt generation system produced unexpected, non-human prompt structures
“It was pretty interesting to look at the prompt outputs in particular, I'd say, for ArcGi in that you know, I think you can read those and say, well, that's not what a human would have written. Pretty clearly. And it's, you know, there's some unexpected stuff …”
Kantrowitz: Gemini 3 crushed the ARC-AGI benchmark and topped Chatbot Arena
“Gemini three smashes the benchmarks crushed on the Arc AGI test. It's currently at the top of the LM arena leaderboards”
Frosst: Enterprise clients will never demand Arc AGI pixel manipulation features
“Stuff like the Arc AGI challenge is a benchmark that people talk about, but that's like a pixel manipulation challenge. It's like, you know, taking in like a grid of pixels and based on rules, predicting the next one. That's not a thing any of our customers ha…”
Lambert: AI benchmarks like ARC-AGI should prioritize testing without harnesses
“Harnesses are cool, but they're gonna, they're,
They're a handicap that's changing the learning dynamics substantially. So it's good. It's good demos, but I feel like the core thrust has to be no harnesses.”
Kamradt: Mike Knoop Put Up $1M for ARC Prize Bounty
“Mike actually put he put up a million dollars of his own money and said, Hey, I'm going to put a bounty. So for anybody who can beat this benchmark, they're going to get a million dollars.”
Kantrowitz: Grok 4 beats all models on ARC-AGI by significant margin
“And then of course in the Arc AGI test, it outperforms Every model by a significant margin.”
Douglas: Frontier Labs Avoid RL Training Directly on ARC-AGI
“And I mean, I think if you are old on Arc AGI, then it would, you'd probably get superhuman at it pretty fast. But I think we're all trying not to RL on it so that it functions as like an interesting held out.”
Knoop: Language models operate by memorization rather than solving novel patterns
“Language models. Generally working like a memorization style regime where they're right. Learning lots of data. They're able to apply it to very similar types of patterns that they've seen before, but not novel patterns. That's what RKGI shows.”
Chollet: OpenAI o3 cost $10k–$20k per ARC puzzle on maximum compute
“For instance OpenAI O.S. On the highest compute settings that we tried it on for Arc, it was consuming somewhere between, like, 10,000 dollars to 20,000 dollars per task, like, for one little puzzle, which you could normally solve with a base of an API for a f…”
Coogan: OpenAI's o3 High-Compute Mode Spent $3,000 to Solve a Benchmark Task
“Oh three, which isn't out yet, but is even more advanced in terms of reasoning. They have a high compute model. That spends almost 3000 dollars per task. And it just thinks for hours and hours and hours basically, and it was able to break arc that that AI, AGI…”
Coogan: OpenAI o3 high-compute configuration costs $2,000 per solve
“O-three is a reasoning model. The high version costs 2000 dollars per solve.”
Knoop: ARC-AGI is the only true AGI evaluation that exists
“Arc AGI to best of my knowledge is the only true. AGI eval that actually exists in the world and measures a actually good definition, correct definition of what AGI is, which we can talk about.”
Knoop: ARC-AGI benchmark performance only moved from 20% to 34% in four years
“There's an AI lab called lab 42 out of Switzerland that's been running a small annual contest over the last four years to try and beat this eval and state of the art today is. 34% state of the art four years ago when it was first introduced was 20%. So we've m…”
Knoop: ARC solution will likely need under 10k code lines, not massive LLMs
“It's quite likely actually that the solution it can be like written in like 10,000 lines of code or less. And it's not gonna require these like, you know, gigantic You know, two hundred billion large parameter models in order to solve it.”