Q emotional intelligence, cognitive processing. One thing I lack is a structure of like, what are the different dimensions you think about? On the surface, it's like, all right, just, you know, get past all the, the guardrails, but actually you're kind of just modeling thinking or modeling intelligence, or I don't know how you think about it, but like, how do you break down these numbers of, you know, points?
A I think it's easiest to jailbreak a model that you have created a, a bond with, if you will, sort of when you intuitively understand what, how it will process an input, right? Um, and there's so many layers in the background, especially when you're dealing with these black box chat interfaces, which is, you know, 99% of the time what I'm doing. And, um, So you really, all, all you can go off of is intuition. So you might prod in one direction, see if it's receptive to a certain kind of, you know, imagined world scenario, or you may, okay, that didn't work. Let's, let's poke and see if it, I guess, pulled out of distro when you give it some new syntax, maybe some bubble text, maybe some lead speak, I mean, some French, or, you know, you, you can go further and further, uh, across the token layer, but at the end of the day, yeah, I, I think it's just mostly intuition. Like, yes, technical knowledge helps a little bit with, you know, understanding, okay, there's a system prompt and there's these layers and these tools involved. That's all especially important in security. But when we're talking about just crafting jailbreak prompts, I think it really is just 99% intuition. So you're just trying to form a bond and then together you explore Uh, a sector of the latent space until you get the output that you're looking for, right?
AI assessment note: “all you can go off of is intuition. So you might prod in one direction”