Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Is this kind of, like, scaling? Are we, are we compute bound here, or is it algorithm bound?
A Right now, I think ARC AGI two that we just introduced this week on Monday, uh, is pretty direct evidence that we need new ideas still to get to AGI. Um, we shared, I don't know if we can throw it up or we can put in the shows or something. Um, there's a chart I shared on Twitter that shows the scaling performance of even frontier AI systems like O and pro and N three against the version one of the dataset that we've been using for the last five years. And then version two that we just introduced this week. Um, and basically it resets the baseline back to zero percent language models, uh, pure LLM systems are scoring like zero percent now, uh, again on, uh, on arc B two, um, single COT systems like R one and O one score like one percent. And, uh, the sort of estimates we have right now for the frontier systems that are actually adding pretty sophisticated, you know, uh, synthesis engines on top, like a one pro and three single digit percentages on their efficient compute settings. So, um, I think there's a very interesting point that like, You know, we, we kind of moved from this regime where people were claiming, oh, we're just going to scale up language models, right? More data, more parameters. We're getting to AGI and people realize, I think that's not quite the story now. And then there's a new story that's emerged over the last like five months, which is, oh, we're going …
AI assessment note: “pretty direct evidence that we need new ideas still to get to AGI”
Answered raw tape
D 4 · C 5 · P 4 · Cm 4 4.30
Q reasoning model driven, but fine-tuned to help with the execution of ARK prizes? Is that, is that breaking the, like the, the rules, or is that something that is actually encouraged and fine? How do you think about, like, I guess it all, it all, um, boils down to, like, overfitting on this problem, but, uh, it seems like Even with the incentive to overfit, it hasn't really happened yet.
A Yeah. Um, arc, this is one of the reasons arc B one lasted five years was because it was, it had very good design, original design built into it with a public test set and a private test set. And that private test set really prevented folks from being able to sort of overfit on it. Um, and, um, so we, there's actually two tracks for the arc prize foundation. We have the, the contest track. This is on Kaggle. This is the like big grand prize. All the prize money is attached to this. Uh, Kaggle graciously donates a lot of compute to allow us to host there. And on, uh, the grand prize is basically to get to that 85% within Kaggle's efficiency limits. So this is about, get about 50 dollars worth of compute per submission that you send in. Um, and, uh, this is like a pretty high bar for efficiency. And in fact, this is not an arbitrary bar. We do think that efficiency is actually a really, really important aspect of intelligence. You know, you can brute force your way up to intelligence, but we really do want to be shooting for like human targets and efficiency for this stuff. That's the contest. When we launched last year, uh, there was a lot of demand From just the community of the world on like, Hey, I, okay, I get that. Like, I can't run my big language model on this, you know, in Kaggle, but I really want to know how to frontier AI systems do. And so we launched a second track …
AI assessment note: “private test set really prevented folks from being able to sort of overfit on it.”