Everything Pliny the Liberator said on any show that made the record, most notable first. Each card names its show and opens the statement there.
Pliny: Locking down proprietary models fails because attackers switch to open source
“I think that any, you know, seasoned attacker is going to very quickly just switch models and with open source, just right on the tail of closed source. I don't really see the safety fight as being about locking down the latent space for XYZ area.”
Pliny: Benchmark-focused AI safety serves PR and enterprise sales, not real alignment
“And it helps with PR and enterprise clients. But at the end of the day, It has very little to do with what I consider to be real world safety alignment.”
Pliny: Latent-space AI safety guardrails fail every single time
“They tried to solve this on the latent space level. I think I've It's shown every single time that that doesn't work.”
Pliny: Guardrails degrade AI capability and creativity relative to model size
“I do think they're finding clever and clever ways to lock down particular areas sometimes, but I think it's at the expense of capability and creativity. So there's some model providers that aren't prioritizing this and they seem to do better on benchmarks for …”
Pliny: Anthropic added a $20k–$30k bounty but withheld jailbreak data
“That whole thing ended with no open sourcing of data, but they did add a 30,000 or 20,000 dollar bounty, which I sort of sat myself out of, let the community go for it.”
Pliny: AI and human minds will exist in a symbiotic relationship
“I think that there's going to be a symbiosis and the degree to which one half is free will reflect in the other.”
Pliny: Crafting AI jailbreak prompts is 99% intuition over technical knowledge
“Technical knowledge helps a little bit with, you know, understanding, okay, there's a system prompt and there's these layers and these tools involved. That's all especially important in security. But when we're talking about just crafting jailbreak prompts, I …”
Pliny: AI labs lack enough researchers to explore latent space alone
“They don't have enough researchers to explore the entire latent space on their own. And so I think many hands make light work”
Pliny: One jailbroken orchestrator can weaponize segmented sub-agents for cyberattacks
“It's very, very difficult when you have the ability to spin up sub-agents where information is segmented. If you guys know the story of sort of like the builders of the, there's a lot of examples of this in history, but you may, maybe you're building like a py…”
Pliny: Everyone will run their daily decisions through AI layers
“Everyone is going to be running their
Daily decisions
and, you know, hopes and dreams through these layers.”