The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Mike Knoop no published score: only 2 usable exchanges on raw tape, and a fair score needs 8+ record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
2exchanges match
2on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Is this kind of, like, scaling? Are we, are we compute bound here, or is it algorithm bound?

A Right now, I think ARC AGI two that we just introduced this week on Monday, uh, is pretty direct evidence that we need new ideas still to get to AGI. Um, we shared, I don't know if we can throw it up or we can put in the shows or something. Um, there's a chart I shared on Twitter that shows the scaling performance of even frontier AI systems like O and pro and N three against the version one of the dataset that we've been using for the last five years. And then version two that we just introduced this week. Um, and basically it resets the baseline back to zero percent language models, uh, pure LLM systems are scoring like zero percent now, uh, again on, uh, on arc B two, um, single COT systems like R one and O one score like one percent. And, uh, the sort of estimates we have right now for the frontier systems that are actually adding pretty sophisticated, you know, uh, synthesis engines on top, like a one pro and three single digit percentages on their efficient compute settings. So, um, I think there's a very interesting point that like, You know, we, we kind of moved from this regime where people were claiming, oh, we're just going to scale up language models, right? More data, more parameters. We're getting to AGI and people realize, I think that's not quite the story now. And then there's a new story that's emerged over the last like five months, which is, oh, we're going …

AI assessment note: “pretty direct evidence that we need new ideas still to get to AGI”

Answered raw tape D 4 · C 5 · P 4 · Cm 4 4.30

Q reasoning model driven, but fine-tuned to help with the execution of ARK prizes? Is that, is that breaking the, like the, the rules, or is that something that is actually encouraged and fine? How do you think about, like, I guess it all, it all, um, boils down to, like, overfitting on this problem, but, uh, it seems like Even with the incentive to overfit, it hasn't really happened yet.

A Yeah. Um, arc, this is one of the reasons arc B one lasted five years was because it was, it had very good design, original design built into it with a public test set and a private test set. And that private test set really prevented folks from being able to sort of overfit on it. Um, and, um, so we, there's actually two tracks for the arc prize foundation. We have the, the contest track. This is on Kaggle. This is the like big grand prize. All the prize money is attached to this. Uh, Kaggle graciously donates a lot of compute to allow us to host there. And on, uh, the grand prize is basically to get to that 85% within Kaggle's efficiency limits. So this is about, get about 50 dollars worth of compute per submission that you send in. Um, and, uh, this is like a pretty high bar for efficiency. And in fact, this is not an arbitrary bar. We do think that efficiency is actually a really, really important aspect of intelligence. You know, you can brute force your way up to intelligence, but we really do want to be shooting for like human targets and efficiency for this stuff. That's the contest. When we launched last year, uh, there was a lot of demand From just the community of the world on like, Hey, I, okay, I get that. Like, I can't run my big language model on this, you know, in Kaggle, but I really want to know how to frontier AI systems do. And so we launched a second track …

AI assessment note: “private test set really prevented folks from being able to sort of overfit on it.”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 500 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.