The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Will Brown no published score: only 2 usable exchanges on raw tape, and a fair score needs 8+ record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
2exchanges match
2on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q Do we have any, this is like already veering off from Claude directly into speculation, but do we have any idea, um, if there are any material differences between how Claude extended thinking works versus like the old series models? Do we know?

A The biggest difference seems to be, at least, and this is kind of a thing that's been, I mean, I don't know, this is all speculation, of course, but from the start, Anthropic had always kind of had this, like, little thinking thing where you could, Sometimes even like cloud 3.5 would do like a tiny bit of thinking, and it was really just like deciding which tool to use for the most part. Like if it was doing, um, an artifact in the cloud UI, it would have this little thing where it would think for like two sentences about which tool to use. And it seemed like Anthropik's kind of attitude has been that extended thinking is an instance of tool use and that it's the kind of thing you want to equip the model with the ability to do. But it's not like, oh, it's a thinking model. It's just a sync for the model to, like, brain vomit, because that brain vomiting will help it, like, find a nice thing to do next. In the same way that doing search or doing code execution are, like, ways to kind of get more information on the path towards, like, finishing a problem.

AI assessment note: “The biggest difference seems to be... extended thinking is an instance of tool use”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q Uh, and this was obviously before the Jason Wei chain of thought paper, but it's all the same sort of method, uh, general family of techniques. I think the question for me is also like, is there some model routing going on? Like, are they different models that the thinking, non-thinking, or are they the same models with like, just like you turn off the end of turn token generation?

A I mean, I think these models should be the same model, and Anthropic knows what they're doing. Like, it's not that hard to, like, Quen did it in a very kind of, like, simple way, and they kind of talked about how they did it a little bit. But it's not, like, too difficult to, um, like, have whether or not a model thinks, like, be the sort of thing. I mean, like, obviously all this stuff is, like, hard at, like, serious scale, but, like, conceptually at least, um, it's not, like, a big problem to solve about how would you ever do it. It's like, no, we have reinforcement learning. We can kind of, like, or we just SFT on, like, different things. We can kind of teach models skills like that pretty.

AI assessment note: “I think these models should be the same model”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.