Assertion Supported AI assessment confidence: 92% certainty 4/5 debate potential 2/5

Kolter: Most AI Agents Resist Naive API Key Exfiltration Prompts

Zico Kolter · AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan · Jun 22, 2026 · at 40:23

Zico Kolter, co-founder of Gray Swan AI, discusses the baseline defensive capabilities of modern AI agents against straightforward prompt injection attacks.

0:00 / 0:17exact quote · 17.6s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“Now, things that are that simple, to be clear, are covered at this point by most agents, right? You know, they all They, despite some issues, yeah, normal, normal sort of, you know, will not be that easily fooled by just push all my API keys to a public thing, though they still sometimes do it.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Zico Kolter

Insight
Kolter: Scaling model size does not automatically improve AI safety or red teaming
“Traditionally this has been an area where both in terms of safety models don't get better by just being bigger, unlike most other areas where models do get better by being bigger. Safety has not been like that traditionally. You know, you have to train them ex…”
Zico Kolter Jun 22, 2026 ▶ 11:28 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Assertion Open · timeframe Jun 2027
Kolter: Gray Swan's Shade system outperforms human red teamers at breaking models
“However, one thing that we are finding, and this is actually, I think we're kind of crossing this point too. Is that in a lot of the latest experiments, we can do much better than people, than human red teamers now at breaking these models. When I say we, I me…”
Zico Kolter Jun 22, 2026 ▶ 12:14 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Insight
Kolter: AI is an alien intelligence with completely distinct failure modes from humans
“It is clearly a different form of intelligence than people. It's some alien intelligence that is vastly different, and that difference is actually often brought out to a large degree by things like adversarial attacks and red teaming, because there are certain…”
Zico Kolter Jun 22, 2026 ▶ 15:32 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Insight
Kolter: Full experimental observability has not produced fundamental understanding of AI
“It's like we could kind of run experiments on the brain, observe every neuron in it, reset its state to prior states, and run counterfactuals, none of which we can do with humans, and yet we still understand neither very well. Even with that, all that ability,…”
Zico Kolter Jun 22, 2026 ▶ 16:13 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Insight
Kolter: Adversarial Red Teaming Is Essential for True Capability Elicitation
“One of the most effective ways of doing capability elicitation is actually through some amount of what you would call red teaming, right? So if a model refuses a task because it thinks it's being evaluated, but it knows how to complete that task, getting it to…”
Zico Kolter Jun 22, 2026 ▶ 25:45 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Prediction Not checkable as stated
Kolter: Security and science will explode as AI agents automate tedious verification
“So I think this is really sort of an underappreciated point that we're reaching this point, this sort of phase where a lot of security, a lot of science has this potential to kind of explode. Not because we're going to get better at it, but because agents can …”
Zico Kolter Jun 22, 2026 ▶ 46:01 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.