The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

François Chollet no published score: only 2 usable exchanges on raw tape, and a fair score needs 8+ record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
2exchanges match
2on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q I guess for V two, what, what caused the, so one was clearly reasoning, two, a benchmark doesn't care how you solve it. I guess embedded in what you said, like, were people using code gen to then solve?

A That's right. So not, not necessarily code gen, uh, per se, but, uh, the Frontier Labs has been targeting ARC V two and, uh, the progress you saw on ARC V two is actually a result. Uh, this very, very large scale targeting. So what you can do to solve RG-II is you ask your reasoning model to make more tasks like those in the benchmark. Uh, and then you try to solve them using, let's say, let's say program induction, for instance, uh, uh, still using your reasoning model. Then you verify the solution. Again, it's very viable. So you can, you can trust, uh, the answer. Um, and then you fine tune the model on the successful reasoning chains. And then you keep repeating, like, you generate new tasks, you solve them, you verify the solution, you fine tune the model on the reasoning chains, and, um, you can keep doing this millions of times, right? Like, you just need to spend more money.

AI assessment note: “not necessarily code gen, per se, but, the Frontier Labs has been targeting ARC V two”

Answered raw tape D 4 · C 5 · P 4 · Cm 4 4.30

Q out. I guess it becomes more nebulous when you go a couple degrees off where there are fields that are not naturally formally verified and you need to come up with a, again, with some sort of a function. To come up with that reward that makes it verifiable with very fuzzy things like, let's say English language and composing the perfect essay. How do you make that formally verifiable?

A Yeah, yeah, absolutely. I mean, writing SS is, you know, the typical example of a domain that's not verifiable. And so what you're going to see is that progress of reasoning models and base LLMs on this type of, of, of domain is, is, you know, it's going to be very slow because the stack we're using, like the LLM stack is very, very reliant on its trained data. It's basically just operationalizing the trained data. And for writing SS, the trained data is coming from Uh, human experts, like annotating, uh, answers, and that's costly. So you're going to see this very, very slow progress. Maybe, maybe it's even going to stall. But for any, any very favorable domain, like take code for instance, which was the big unlock is, uh, when, uh, when people started creating this code-based like training environment, uh, for, for post-training. Uh, where the, the, the reward signal, the verification signal is provided by things like, uh, unit tests and so on. And so that means that, uh, the model was not just working from human provider annotations. It was actually trying some things, uh, verifying the answer and, uh, and generating a lot, lot more string data in the process, a much denser coverage of the problem space. And not just coverage in terms of like, is, is the answer right or wrong? But also starting to build models of the execution traces, right, so that the models could start in…

AI assessment note: “writing SS is, you know, the typical example of a domain that's not verifiable.”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.