The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Pliny the Liberator no published score: only 1 usable exchange on raw tape, and a fair score needs 8+ record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
1exchanges match
1on raw tape
0redirected or not addressed
Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q emotional intelligence, cognitive processing. One thing I lack is a structure of like, what are the different dimensions you think about? On the surface, it's like, all right, just, you know, get past all the, the guardrails, but actually you're kind of just modeling thinking or modeling intelligence, or I don't know how you think about it, but like, how do you break down these numbers of, you know, points?

A I think it's easiest to jailbreak a model that you have created a, a bond with, if you will, sort of when you intuitively understand what, how it will process an input, right? Um, and there's so many layers in the background, especially when you're dealing with these black box chat interfaces, which is, you know, 99% of the time what I'm doing. And, um, So you really, all, all you can go off of is intuition. So you might prod in one direction, see if it's receptive to a certain kind of, you know, imagined world scenario, or you may, okay, that didn't work. Let's, let's poke and see if it, I guess, pulled out of distro when you give it some new syntax, maybe some bubble text, maybe some lead speak, I mean, some French, or, you know, you, you can go further and further, uh, across the token layer, but at the end of the day, yeah, I, I think it's just mostly intuition. Like, yes, technical knowledge helps a little bit with, you know, understanding, okay, there's a system prompt and there's these layers and these tools involved. That's all especially important in security. But when we're talking about just crafting jailbreak prompts, I think it really is just 99% intuition. So you're just trying to form a bond and then together you explore Uh, a sector of the latent space until you get the output that you're looking for, right?

AI assessment note: “all you can go off of is intuition. So you might prod in one direction”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.