The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Nathan Labenz no published score: only 4 usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.0/5 from 4 raw tape exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
4exchanges match
4on raw tape
2redirected or not addressed
Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q And, and so, I mean, one of the things that you mentioned that, that Cal, um, you know, the analysis missed was, was that it way underestimated the value of, of, of, of, of extended reasoning. Right. Um, and so what, what would it mean to, to fully sort of appreciate that?

A Well, I mean, a big one from just the last few weeks was that we had an IMO gold medal with pure reasoning models, uh, with no access to tools from multiple companies. And, you know, that is night and day compared to what GPT-IV could do with math, right? We, and these things are really weird. Like it's nothing I say here should be, uh, intended to suggest that people won't be able to find weaknesses in the models. I, I still use a tic-tac-toe puzzle to this day where I take a picture of a tic-tac-toe board Where some, the, one of the players has made a wrong move, um, that is not optimal, and thus allows the other player to force a win, and I ask the models if somebody can force a win from this position. Only very recently, only the last generation of models are starting to get that right some of the time. Almost always before, they were like, tic-tac-toe's a solved game, you know, you can always get a draw, there's, and they would wrongly assess my board position as the player can still get a draw. So there's a lot of weird stuff, right? The jagged, uh, capabilities frontier remains a real issue and people are going to find, you know, peaks and valleys for sure. But GPT-IV, when it first came out, couldn't do anything approaching IMO gold problems. It was still struggling on like high school math. And since then we've seen this high school math progression all the way up thro…

AI assessment note: “we had an IMO gold medal with pure reasoning models”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q And, and so, I mean, one of the things that you mentioned that, that Cal, um, you know, the analysis missed was, was that it way underestimated the value of, of, of, of, of extended reasoning. Right. Um, and so what, what would it mean to, to fully sort of appreciate that?

A Well, I mean, a big one from just the last few weeks was that we had an IMO gold medal with pure reasoning models, uh, with no access to tools from multiple companies. And, you know, that is night and day compared to what GPT-IV could do with math, right? We, and these things are really weird. Like it's nothing I say here should be, uh, intended to suggest that people won't be able to find weaknesses in the models. I, I still use a tic-tac-toe puzzle to this day where I take a picture of a tic-tac-toe board Where some, the, one of the players has made a wrong move, um, that is not optimal, and thus allows the other player to force a win, and I ask the models if somebody can force a win from this position. Only very recently, only the last generation of models are starting to get that right some of the time. Almost always before, they were like, tic-tac-toe's a solved game, you know, you can always get a draw, there's, and they would wrongly assess my board position as the player can still get a draw. So there's a lot of weird stuff, right? The jagged, uh, capabilities frontier remains a real issue and people are going to find, you know, peaks and valleys for sure. But GPT-IV, when it first came out, couldn't do anything approaching IMO gold problems. It was still struggling on like high school math. And since then we've seen this high school math progression all the way up thro…

AI assessment note: “we had an IMO gold medal with pure reasoning models”

Redirected raw tape D 1 · C 4 · P 4 · Cm 4 3.10

Q maybe there's, there's less of a, uh, sort of concern around, um, you know, people being replaced in the next, next few years in, in, in, in, in mass. I think we spoke maybe a year ago about this, or I think you said something like 50% of 50% of jobs. Um, I'm curious if that's still your, your, uh, your litmus test, or how do you think about it?

A Well, for one thing, I think that meter paper is worth unpacking a little bit more because this was one of those things that was, and I am a big fan of meter and I have no, um, you know, no shade on them. Cause I do think do science, publish your results. Like that's good. You don't have to, uh, make every experimental result and everything you put out conform to a narrative. But I do think it was a little bit. It was a little bit too easy for people who wanted to say that, oh, this is all nonsense to latch onto that. And, you know, again, there's, there's something there that I would kind of put in the Cal Newport category too, where for me, maybe the most interesting thing was the users thought that they were faster when in fact they seem to be slower. So that sort of misperception of oneself, I think is really interesting. Personally, I think there's some explanations for that that include like, Hit and go on the agent, going to social media and scrolling around for a while and then coming back. The thing might have been done for quite a while by the time I get back. So honestly, one like really simple, and we're starting to see this in products. One really simple thing that the products can do to address those concerns is just provide notifications. Like the thing is done now. So, you know, stop scrolling and come back and check its work. That, in terms of just clock time, …

AI assessment note: “I think that meter paper is worth unpacking a little bit more”

Not addressed raw tape D 1 · C 3 · P 3 · Cm 3 2.40

Q maybe there's, there's less of a, uh, sort of concern around, um, you know, people being replaced in the next, next few years in, in, in, in, in mass. I think we spoke maybe a year ago about this, or I think you said something like 50% of 50% of jobs. Um, I'm curious if that's still your, your, uh, your litmus test, or how do you think about it?

A Well, for one thing, I think that meter paper is worth unpacking a little bit more because this was one of those things that was, and I am a big fan of meter and I have no, um, you know, no shade on them. Cause I do think do science, publish your results. Like that's good. You don't have to, uh, make every experimental result and everything you put out conform to a narrative. But I do think it was a little bit. It was a little bit too easy for people who wanted to say that, oh, this is all nonsense to latch onto that. And, you know, again, there's, there's something there that I would kind of put in the Cal Newport category too, where for me, maybe the most interesting thing was the users thought that they were faster when in fact they seem to be slower. So that sort of misperception of oneself, I think is really interesting. Personally, I think there's some explanations for that that include like, Hit and go on the agent, going to social media and scrolling around for a while and then coming back. The thing might have been done for quite a while by the time I get back. So honestly, one like really simple, and we're starting to see this in products. One really simple thing that the products can do to address those concerns is just provide notifications. Like the thing is done now. So, you know, stop scrolling and come back and check its work. That, in terms of just clock time, …

AI assessment note: “Well, for one thing, I think that meter paper is worth unpacking”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.