Jailbreaking
topic on 3 shows · 5 statements across 4 episodes
Latent Space
Lenny's Podcast
No Priors
5 statements about Jailbreaking, every show
Bissell: Steering Experiments Can Predict Examples Needed for Jailbreaks
“What's in this in context learning and activation steering equivalence paper is you can like predict the number of examples that you will need to put in there in order to jailbreak the model. By doing steering experiments and using this sort of like equivalenc…”
Schulhoff: Jailbreaking targets models directly; prompt injection overrides developer prompts
“So the difference is in jailbreaking. It's just a malicious user and a model. In prompt injection, it's a malicious user, a model, and some developer prompt that the malicious user is trying to get the model to ignore.”
Schulhoff: All Transformer-Based Chatbots Are Vulnerable to Adversarial Attacks
“And because all I guess for the most part, all currently deployed chatbots are based on transformers or transformer adjacent technologies. They're all vulnerable to Prompt injection, jailbreaking, forms of adversarial attacks.”
Schulhoff: Prompt injection overrides developer instructions; jailbreaking bypasses model directly
“Basically prompt injection is something that occurs when there is developer input, In the prompt, as well as user input in the prompt. So the developer instructions will say to do one thing, the user input will say to do something else. Jailbreaking is when it…”
Liang: Interconnected AI accepting external inputs risks cascading jailbreak exploits
“If these models start interacting with the world and accepting external inputs, now you can not only just sort of jailbreak your own model, but you can jailbreak other people's model and get them to do various things. And then, so that could lead to sort of a …”