ARC-AGI, every mention
17 scenes, the whole family · ← back to ARC-AGI
tap a year for its mentions
every year anyone Shawn Wang 12Greg Kamradt 11Alessio Fanelli 7Nathan Lambert 5Ashvin Nair 1
Verbatim, from the transcripts: the passages where ARC-AGI comes up
[State of AI Papers 2025] Fixing Research with Social Signals, OCR & Implementation — Team AlphaXiv
- ▶ 11:58 unnamed speaker So, like, Sudoku, Arc AGI, I believe I guess, like, 45, 50% on Arc AGI one. 2 times in the scene
[State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
- ▶ 32:08 Ashvin Nair The anthropic, uh, models, like the, uh, Opus two, 4.5, it has this kind of like, uh, there's this like RKGI two plot that looks exactly like the API ones, right?
Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders
- ▶ 29:34 unnamed speaker I just, I can't be, um, you know, I can't have Greg here and not make comparisons with terminal bench two and RKGI two and like that developments, you know, something that, uh, Greg is sort of pushing is.
The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
- ▶ 33:16 Nathan Lambert Good feedback for the ArcGGI people for the V three benchmark is like have things where the language model needs to learn to use new, new actuators in the world after a certain threshold. 5 times in the scene
- ▶ 37:52 Shawn Wang I think the Arc AGI definition of AGI
⚡️ARC-AGI-3: The Interactive Reasoning Benchmark
- ▶ 0:27 Shawn Wang That is the sponsor of the Arc AGI Challenge Benchmark. 5 times in the scene
- ▶ 3:19 Shawn Wang I, I think also, I think there's a little bit, there's a, there's a fit between the reasoning paradigm and RKGI.
- ▶ 14:20 Shawn Wang I, I, this is me using some prior knowledge of the solutions to ArcGi that apparently most people say that adding multimodal vision doesn't actually help. 2 times in the scene
- ▶ 17:03 Alessio Fanelli I know in the B two, you also had the dollar per task thing. 2 times in the scene
- ▶ 24:43 Shawn Wang I think there's always the, the, the question about like the future roadmap for our AGI, presumably, you know, you, you went to two and three really quickly. 2 times in the scene
- ▶ 25:39 Greg Kamradt So the way I think about it is if you had like a linear spectrum, RKGI one and two, it's a static list of like three or four JSON grids static. 2 times in the scene
- ▶ 27:10 Alessio Fanelli Uh, timeline on when you think B two will get close to like 50%, then a hundred percent, because I know graph for yesterday really 16%, which everybody was going crazy over, but it's still 16%. 5 times in the scene
- ▶ 29:04 Greg Kamradt And like, I'm even hesitant to say the words like beat Arc AGI. 3 times in the scene
- ▶ 29:19 Greg Kamradt Because our hypothesis is the thing that actually does beat Arc AGI too. 2 times in the scene
- ▶ 33:55 Greg Kamradt Um, cause that really gives kind of a public boost and a public, like, vote of confidence that Arc AGI, like, has the merit to, like, they're saying that, yes, we trust Arc AGI to help describe the intelligence of these models. 4 times in the scene
2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
- ▶ 1:07:48 Shawn Wang From Andy Kaminsky, which is the guy that we talked to yesterday, who has launched a similar sort of Arc AGI attempt on a suite bench type metric, which arguably is a bit more useful.
State of the Art: Training 70B LLMs on 10,000 H100 clusters
- ▶ 1:11:22 unnamed speaker So, um, Arc AGI, uh, Francois Cholet's, um, uh, hot new thing. 3 times in the scene