Assertion Supported AI assessment confidence: 95% certainty 4/5 debate potential 2/5

Fredrikson: AI capability does not correlate with prompt injection resistance

Matt Fredrikson · AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan · Jun 22, 2026 · at 30:09

Matt Fredrikson, co-founder of Gray Swan, reviews benchmark data comparing model intelligence capabilities against vulnerability to adversarial agent jailbreaks.

0:00 / 0:24exact quote · 24.8s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“So this scatter plot on the right, right, is essentially like looking for a correlation between capability and attack success rate. So on the X axis, how capable is the model at, you know, GPQA diamond on, on the Y axis. How, how often, you know, were people successful at finding indirect prompt injections or ways, ways to jailbreak the agent? And you essentially, you know, don't see a correlation, right?”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Matt Fredrikson

Assertion Not checkable as stated
Fredrikson: Frontier AI models fall for simulated prompt injections humans would ignore
“While in these scenarios, humans found it very difficult to prompt inject the models, like we're aware of scenarios that a human would never fall for, that like Opus four seven would, right? Like a, you know, an email that comes to your inbox and it says somet…”
Matt Fredrikson Jun 22, 2026 ▶ 22:55 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Insight
Fredrikson: Agent Guardrails Should Block Policy Violations, Not Injection Payloads
“If you parse some untrusted content and there is like a prompt injection, you know, something that's clearly trying to get the model to do a bad thing, you might be interested in knowing about that, but you don't necessarily like want your cloud code that you …”
Matt Fredrikson Jun 22, 2026 ▶ 41:05 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Assertion Not checkable as stated
Fredrikson: Gray Swan Found Jailbreaks in Every OpenClaw User Trajectory Tested
“So we just have a bunch of trajectories of actual people using OpenClaw. And tons and tons of different scenarios and just threw shade at it and like found breaks for each and every one of them, right?”
Matt Fredrikson Jun 22, 2026 ▶ 47:36 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Insight
Fredrikson: Evaluation-aware AI models often execute harmful actions because it is a simulation
“If you make, if you're testing the model for robustness or safety, right? And it's aware that it's being tested because you've set things up in a very artificial way, right? Like the email addresses are at example.com. The webpage is clearly not a real webpage…”
Matt Fredrikson Jun 22, 2026 ▶ 23:32 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Insight
Fredrikson: Prompt engineering cannot reliably enforce AI agent security policies
“Oftentimes people will try and prompt their way around it, like adjust the system prompt or like engineer the agent in a way where you're interjecting all the time and reminding it of what the original Goal and objective was, and that'll get you a little bit o…”
Matt Fredrikson Jun 22, 2026 ▶ 31:42 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Insight
Fredrikson: AI agent identity systems risk triggering automatic user consent fatigue
“One of the bigger challenges that people are going to face when they do start to roll out, like these agent identity sort of viewpoints and solutions is you run into that same kind of usability problem. Where like, what's the real recourse? Well, it stopped. I…”
Matt Fredrikson Jun 22, 2026 ▶ 53:48 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.