The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Jack Soslow no published score: only 2 usable exchanges on raw tape, and a fair score needs 8+ record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
2exchanges match
2on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q a recent example of this is ChatGPT, which is throwing out a lot of correct answers, but also a lot of incorrect answers, and humans are just kind of taking it at face value. So maybe we don't even need human-like bots to trick humans. It seems like less evolved bots are already doing that. So what do you think about that? Have we already kind of crossed that chasm?

A Oh, absolutely. Um, and I think the question here is, like, at what point do you cross the chasm in whatever domain to deceive humans? So, for example, it's like, you used to read the words, like, don't believe everything you read on the internet. Now, I think this applies to AI. Like, don't believe everything you read from AI. Um, the Turing test, as you say, is like a flawed metric, and the reason for this is in part because It is too low of a bar. It is a binary bar, so it's either intelligent or not intelligent, and it tests for deception, not real understanding. What these AIs do is they Try to understand the underlying rules that generate the dataset that they're trained on, um, and then generate an approximation of what they believe is the next token. But it doesn't mean that it's right. Like in the case of chat GPT, it communicates with a lot of confidence. A great example of this might be say, asking chat GPT to multiply a 153 times, 257. Like it might get an answer that's close to correct and it communicates it like it is correct, but it's not correct. And this is a deterministic domain where you know what the right answer is. There are a lot of domains where there isn't actually a right answer. The right answer is much more amorphous. Say like you ask about China's, what should China's AI strategy be? Like this is a very complex question and it is dictated by say lik…

AI assessment note: “Oh, absolutely.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q a recent example of this is ChatGPT, which is throwing out a lot of correct answers, but also a lot of incorrect answers, and humans are just kind of taking it at face value. So maybe we don't even need human-like bots to trick humans. It seems like less evolved bots are already doing that. So what do you think about that? Have we already kind of crossed that chasm?

A Oh, absolutely. Um, and I think the question here is, like, at what point do you cross the chasm in whatever domain to deceive humans? So, for example, it's like, you used to read the words, like, don't believe everything you read on the internet. Now, I think this applies to AI. Like, don't believe everything you read from AI. Um, the Turing test, as you say, is like a flawed metric, and the reason for this is in part because It is too low of a bar. It is a binary bar, so it's either intelligent or not intelligent, and it tests for deception, not real understanding. What these AIs do is they Try to understand the underlying rules that generate the dataset that they're trained on, um, and then generate an approximation of what they believe is the next token. But it doesn't mean that it's right. Like in the case of chat GPT, it communicates with a lot of confidence. A great example of this might be say, asking chat GPT to multiply a 153 times, 257. Like it might get an answer that's close to correct and it communicates it like it is correct, but it's not correct. And this is a deterministic domain where you know what the right answer is. There are a lot of domains where there isn't actually a right answer. The right answer is much more amorphous. Say like you ask about China's, what should China's AI strategy be? Like this is a very complex question and it is dictated by say lik…

AI assessment note: “Oh, absolutely. Um, and I think the question here is, like”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.