Jun 22, 2026 · 1h 7m · latent-space

AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan

Zico Kolter · 29m spoken Matt Fredrikson · 21m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this in-depth interview, Gray Swan co-founders Zico Kolter and Matt Fredrikson discuss the emerging frontier of AI security, examining how automated red teaming, specialized runtime guardrails, and autonomous coding agents address the critical vulnerabilities of modern language models and enterprise agents.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 4.3 Guest teaching 5.4 Guest disagreement 2.3 The hosts pushing back 1.9
05100:0015:0030:0045:001:00:001:12–7:24 · The hosts as informed peer 5/10 Founding Gray Swan and Reframing AI Security Paradigms The host references historical adversarial ML work and Ian Goodfellow, showing relevant domain familiarity. Zico Kolter and Matt Fredrikson clearly reframe AI security away from traditional cybersecurity toward treating models as untrusted software components.7:26–13:34 · The hosts as informed peer 4/10 Gray Swan Arena and Automated Red Teaming with Shade The host asks whether models can red team themselves using on-policy reinforcement learning. Zico reframes this, explaining why frontier models struggle with self-red-teaming due to safety alignment refusals.13:34–16:47 · The hosts as informed peer 3/10 AI Intelligence, Red Teaming Experts, and Alien Cognition When the host suggests red teaming failures prove models do not model intelligence, Zico pushes back, arguing models represent an alien form of intelligence that falls for distinct failure modes.16:48–20:08 · The hosts as informed peer 5/10 Accelerating Mechanistic Interpretability Through Autonomous Coding Agents The host brings up mechanistic interpretability lagging capability scaling and cites Neil Nanda. Zico offers an optimistic reframe that autonomous coding agents can turn mechanistic interpretability into an automated science.20:09–24:01 · The hosts as informed peer 4/10 The Human Browser Agent Robustness Challenge Findings Matt explains the Human Browser Agent Robustness Challenge where humans and models were subjected to parallel phishing and prompt injection attacks. The host clarifies the double-blind setup and realistic threat modeling.24:03–27:37 · The hosts as informed peer 4/10 Evaluation Awareness, Model Sandbagging, and Capability Elicitation The host notes evaluation awareness risks causing false positives or negatives. Zico connects capability sandbagging and evaluation awareness to the necessity of adversarial elicitation.27:37–30:40 · The hosts as informed peer 4/10 Introducing Cygnal and Addressing the Robustness Scaling Paradox The host asks if robustness is an orthogonal guardrail layer. Zico and Matt explain the robustness scaling paradox, demonstrating with empirical data that scale does not improve adversarial resilience.30:42–35:00 · The hosts as informed peer 5/10 Enterprise AI Guardrails and Custom Policy Enforcement The host questions why enterprises cannot rely on open-source guards like Llama Guard. The guests explain that production deployments require highly configurable filtering for enterprise-specific policies.35:01–39:42 · The hosts as informed peer 4/10 Deconstructing the Lethal Trifecta and AI Threat Models The conversation walks through Simon Willison's lethal trifecta. Matt contrasts classical software debugging with the probabilistic nature of AI vulnerability mitigation.39:42–42:13 · The hosts as informed peer 4/10 Operationalizing Inbound and Outbound Tool-Call Security The host diagrams how Cygnal sits in the loop. The guests detail bidirectional monitoring, explaining why filtering both untrusted inputs and outbound tool-call actions is required.42:14–46:42 · The hosts as informed peer 5/10 Formal Verification and AI-Assisted Secure Software Engineering Matt discusses formal verification and obscure provably secure languages, which the host challenges as unrealistic for normal developers. Zico clarifies that coding agents will manage the verification layer under the hood.46:42–51:51 · The hosts as informed peer 5/10 Securing High-Risk Agents: OpenClaw and Computer Use The host highlights dangerous enterprise adoption of OpenClaw and computer use. The guests describe stress-testing these setups, finding abundant vulnerabilities and emphasizing isolation controls.51:52–55:17 · The hosts as informed peer 4/10 Designing Agent Identity, Access Control, and Digital Personas The host brings up agent-native authentication and identity delegation. Zico and Matt propose context-specific user personas rather than per-app credentials to avoid consent fatigue.55:17–58:46 · The hosts as informed peer 3/10 The Enterprise Expansion and Proactive AI Security Shift The host asks about upcoming industry developments. The guests share how enterprise demand shifted from reactive breach remediation to proactive pre-deployment guardrailing.58:46–1:01:31 · The hosts as informed peer 5/10 Private Arena Competitions and Adversarial Incentive Alignment The host questions how community red teaming avoids reward hacking and tests sensitive enterprise software. The guests describe private arenas under NDA and calibrated incentives.1:01:32–1:07:21 · The hosts as informed peer 5/10 AI Underwriting, Compliance Standards, and Preparing for Gray Swans The host probes AI insurance compliance and compares SOC 2 frameworks. Zico critiques SOC 2 as an accounting-driven standard, explaining what a rigorous technical standard needs to accomplish.1:12–7:24 · Guest teaching 5/10 Founding Gray Swan and Reframing AI Security Paradigms The host references historical adversarial ML work and Ian Goodfellow, showing relevant domain familiarity. Zico Kolter and Matt Fredrikson clearly reframe AI security away from traditional cybersecurity toward treating models as untrusted software components.7:26–13:34 · Guest teaching 6/10 Gray Swan Arena and Automated Red Teaming with Shade The host asks whether models can red team themselves using on-policy reinforcement learning. Zico reframes this, explaining why frontier models struggle with self-red-teaming due to safety alignment refusals.13:34–16:47 · Guest teaching 6/10 AI Intelligence, Red Teaming Experts, and Alien Cognition When the host suggests red teaming failures prove models do not model intelligence, Zico pushes back, arguing models represent an alien form of intelligence that falls for distinct failure modes.16:48–20:08 · Guest teaching 5/10 Accelerating Mechanistic Interpretability Through Autonomous Coding Agents The host brings up mechanistic interpretability lagging capability scaling and cites Neil Nanda. Zico offers an optimistic reframe that autonomous coding agents can turn mechanistic interpretability into an automated science.20:09–24:01 · Guest teaching 5/10 The Human Browser Agent Robustness Challenge Findings Matt explains the Human Browser Agent Robustness Challenge where humans and models were subjected to parallel phishing and prompt injection attacks. The host clarifies the double-blind setup and realistic threat modeling.24:03–27:37 · Guest teaching 5/10 Evaluation Awareness, Model Sandbagging, and Capability Elicitation The host notes evaluation awareness risks causing false positives or negatives. Zico connects capability sandbagging and evaluation awareness to the necessity of adversarial elicitation.27:37–30:40 · Guest teaching 6/10 Introducing Cygnal and Addressing the Robustness Scaling Paradox The host asks if robustness is an orthogonal guardrail layer. Zico and Matt explain the robustness scaling paradox, demonstrating with empirical data that scale does not improve adversarial resilience.30:42–35:00 · Guest teaching 6/10 Enterprise AI Guardrails and Custom Policy Enforcement The host questions why enterprises cannot rely on open-source guards like Llama Guard. The guests explain that production deployments require highly configurable filtering for enterprise-specific policies.35:01–39:42 · Guest teaching 5/10 Deconstructing the Lethal Trifecta and AI Threat Models The conversation walks through Simon Willison's lethal trifecta. Matt contrasts classical software debugging with the probabilistic nature of AI vulnerability mitigation.39:42–42:13 · Guest teaching 5/10 Operationalizing Inbound and Outbound Tool-Call Security The host diagrams how Cygnal sits in the loop. The guests detail bidirectional monitoring, explaining why filtering both untrusted inputs and outbound tool-call actions is required.42:14–46:42 · Guest teaching 6/10 Formal Verification and AI-Assisted Secure Software Engineering Matt discusses formal verification and obscure provably secure languages, which the host challenges as unrealistic for normal developers. Zico clarifies that coding agents will manage the verification layer under the hood.46:42–51:51 · Guest teaching 6/10 Securing High-Risk Agents: OpenClaw and Computer Use The host highlights dangerous enterprise adoption of OpenClaw and computer use. The guests describe stress-testing these setups, finding abundant vulnerabilities and emphasizing isolation controls.51:52–55:17 · Guest teaching 5/10 Designing Agent Identity, Access Control, and Digital Personas The host brings up agent-native authentication and identity delegation. Zico and Matt propose context-specific user personas rather than per-app credentials to avoid consent fatigue.55:17–58:46 · Guest teaching 5/10 The Enterprise Expansion and Proactive AI Security Shift The host asks about upcoming industry developments. The guests share how enterprise demand shifted from reactive breach remediation to proactive pre-deployment guardrailing.58:46–1:01:31 · Guest teaching 5/10 Private Arena Competitions and Adversarial Incentive Alignment The host questions how community red teaming avoids reward hacking and tests sensitive enterprise software. The guests describe private arenas under NDA and calibrated incentives.1:01:32–1:07:21 · Guest teaching 6/10 AI Underwriting, Compliance Standards, and Preparing for Gray Swans The host probes AI insurance compliance and compares SOC 2 frameworks. Zico critiques SOC 2 as an accounting-driven standard, explaining what a rigorous technical standard needs to accomplish.1:12–7:24 · Guest disagreement 2/10 Founding Gray Swan and Reframing AI Security Paradigms The host references historical adversarial ML work and Ian Goodfellow, showing relevant domain familiarity. Zico Kolter and Matt Fredrikson clearly reframe AI security away from traditional cybersecurity toward treating models as untrusted software components.7:26–13:34 · Guest disagreement 3/10 Gray Swan Arena and Automated Red Teaming with Shade The host asks whether models can red team themselves using on-policy reinforcement learning. Zico reframes this, explaining why frontier models struggle with self-red-teaming due to safety alignment refusals.13:34–16:47 · Guest disagreement 4/10 AI Intelligence, Red Teaming Experts, and Alien Cognition When the host suggests red teaming failures prove models do not model intelligence, Zico pushes back, arguing models represent an alien form of intelligence that falls for distinct failure modes.16:48–20:08 · Guest disagreement 2/10 Accelerating Mechanistic Interpretability Through Autonomous Coding Agents The host brings up mechanistic interpretability lagging capability scaling and cites Neil Nanda. Zico offers an optimistic reframe that autonomous coding agents can turn mechanistic interpretability into an automated science.20:09–24:01 · Guest disagreement 2/10 The Human Browser Agent Robustness Challenge Findings Matt explains the Human Browser Agent Robustness Challenge where humans and models were subjected to parallel phishing and prompt injection attacks. The host clarifies the double-blind setup and realistic threat modeling.24:03–27:37 · Guest disagreement 3/10 Evaluation Awareness, Model Sandbagging, and Capability Elicitation The host notes evaluation awareness risks causing false positives or negatives. Zico connects capability sandbagging and evaluation awareness to the necessity of adversarial elicitation.27:37–30:40 · Guest disagreement 2/10 Introducing Cygnal and Addressing the Robustness Scaling Paradox The host asks if robustness is an orthogonal guardrail layer. Zico and Matt explain the robustness scaling paradox, demonstrating with empirical data that scale does not improve adversarial resilience.30:42–35:00 · Guest disagreement 2/10 Enterprise AI Guardrails and Custom Policy Enforcement The host questions why enterprises cannot rely on open-source guards like Llama Guard. The guests explain that production deployments require highly configurable filtering for enterprise-specific policies.35:01–39:42 · Guest disagreement 2/10 Deconstructing the Lethal Trifecta and AI Threat Models The conversation walks through Simon Willison's lethal trifecta. Matt contrasts classical software debugging with the probabilistic nature of AI vulnerability mitigation.39:42–42:13 · Guest disagreement 1/10 Operationalizing Inbound and Outbound Tool-Call Security The host diagrams how Cygnal sits in the loop. The guests detail bidirectional monitoring, explaining why filtering both untrusted inputs and outbound tool-call actions is required.42:14–46:42 · Guest disagreement 3/10 Formal Verification and AI-Assisted Secure Software Engineering Matt discusses formal verification and obscure provably secure languages, which the host challenges as unrealistic for normal developers. Zico clarifies that coding agents will manage the verification layer under the hood.46:42–51:51 · Guest disagreement 3/10 Securing High-Risk Agents: OpenClaw and Computer Use The host highlights dangerous enterprise adoption of OpenClaw and computer use. The guests describe stress-testing these setups, finding abundant vulnerabilities and emphasizing isolation controls.51:52–55:17 · Guest disagreement 2/10 Designing Agent Identity, Access Control, and Digital Personas The host brings up agent-native authentication and identity delegation. Zico and Matt propose context-specific user personas rather than per-app credentials to avoid consent fatigue.55:17–58:46 · Guest disagreement 1/10 The Enterprise Expansion and Proactive AI Security Shift The host asks about upcoming industry developments. The guests share how enterprise demand shifted from reactive breach remediation to proactive pre-deployment guardrailing.58:46–1:01:31 · Guest disagreement 2/10 Private Arena Competitions and Adversarial Incentive Alignment The host questions how community red teaming avoids reward hacking and tests sensitive enterprise software. The guests describe private arenas under NDA and calibrated incentives.1:01:32–1:07:21 · Guest disagreement 3/10 AI Underwriting, Compliance Standards, and Preparing for Gray Swans The host probes AI insurance compliance and compares SOC 2 frameworks. Zico critiques SOC 2 as an accounting-driven standard, explaining what a rigorous technical standard needs to accomplish.1:12–7:24 · The hosts pushing back 2/10 Founding Gray Swan and Reframing AI Security Paradigms The host references historical adversarial ML work and Ian Goodfellow, showing relevant domain familiarity. Zico Kolter and Matt Fredrikson clearly reframe AI security away from traditional cybersecurity toward treating models as untrusted software components.7:26–13:34 · The hosts pushing back 2/10 Gray Swan Arena and Automated Red Teaming with Shade The host asks whether models can red team themselves using on-policy reinforcement learning. Zico reframes this, explaining why frontier models struggle with self-red-teaming due to safety alignment refusals.13:34–16:47 · The hosts pushing back 2/10 AI Intelligence, Red Teaming Experts, and Alien Cognition When the host suggests red teaming failures prove models do not model intelligence, Zico pushes back, arguing models represent an alien form of intelligence that falls for distinct failure modes.16:48–20:08 · The hosts pushing back 1/10 Accelerating Mechanistic Interpretability Through Autonomous Coding Agents The host brings up mechanistic interpretability lagging capability scaling and cites Neil Nanda. Zico offers an optimistic reframe that autonomous coding agents can turn mechanistic interpretability into an automated science.20:09–24:01 · The hosts pushing back 2/10 The Human Browser Agent Robustness Challenge Findings Matt explains the Human Browser Agent Robustness Challenge where humans and models were subjected to parallel phishing and prompt injection attacks. The host clarifies the double-blind setup and realistic threat modeling.24:03–27:37 · The hosts pushing back 2/10 Evaluation Awareness, Model Sandbagging, and Capability Elicitation The host notes evaluation awareness risks causing false positives or negatives. Zico connects capability sandbagging and evaluation awareness to the necessity of adversarial elicitation.27:37–30:40 · The hosts pushing back 2/10 Introducing Cygnal and Addressing the Robustness Scaling Paradox The host asks if robustness is an orthogonal guardrail layer. Zico and Matt explain the robustness scaling paradox, demonstrating with empirical data that scale does not improve adversarial resilience.30:42–35:00 · The hosts pushing back 2/10 Enterprise AI Guardrails and Custom Policy Enforcement The host questions why enterprises cannot rely on open-source guards like Llama Guard. The guests explain that production deployments require highly configurable filtering for enterprise-specific policies.35:01–39:42 · The hosts pushing back 1/10 Deconstructing the Lethal Trifecta and AI Threat Models The conversation walks through Simon Willison's lethal trifecta. Matt contrasts classical software debugging with the probabilistic nature of AI vulnerability mitigation.39:42–42:13 · The hosts pushing back 1/10 Operationalizing Inbound and Outbound Tool-Call Security The host diagrams how Cygnal sits in the loop. The guests detail bidirectional monitoring, explaining why filtering both untrusted inputs and outbound tool-call actions is required.42:14–46:42 · The hosts pushing back 3/10 Formal Verification and AI-Assisted Secure Software Engineering Matt discusses formal verification and obscure provably secure languages, which the host challenges as unrealistic for normal developers. Zico clarifies that coding agents will manage the verification layer under the hood.46:42–51:51 · The hosts pushing back 2/10 Securing High-Risk Agents: OpenClaw and Computer Use The host highlights dangerous enterprise adoption of OpenClaw and computer use. The guests describe stress-testing these setups, finding abundant vulnerabilities and emphasizing isolation controls.51:52–55:17 · The hosts pushing back 2/10 Designing Agent Identity, Access Control, and Digital Personas The host brings up agent-native authentication and identity delegation. Zico and Matt propose context-specific user personas rather than per-app credentials to avoid consent fatigue.55:17–58:46 · The hosts pushing back 1/10 The Enterprise Expansion and Proactive AI Security Shift The host asks about upcoming industry developments. The guests share how enterprise demand shifted from reactive breach remediation to proactive pre-deployment guardrailing.58:46–1:01:31 · The hosts pushing back 2/10 Private Arena Competitions and Adversarial Incentive Alignment The host questions how community red teaming avoids reward hacking and tests sensitive enterprise software. The guests describe private arenas under NDA and calibrated incentives.1:01:32–1:07:21 · The hosts pushing back 3/10 AI Underwriting, Compliance Standards, and Preparing for Gray Swans The host probes AI insurance compliance and compares SOC 2 frameworks. Zico critiques SOC 2 as an accounting-driven standard, explaining what a rigorous technical standard needs to accomplish.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 15:07 Rejecting the claim that LLMs lack intelligence

Zico directly rejects the host's premise that failure on edge-case red teaming prompts shows models do not possess intelligence.

Hardest push from the hosts ▶ 43:50 Pushing back on formal verification in production

The host openly doubts Matt's suggestion of using obscure formally verified languages, pointing out that engineers prefer plain English and accessible code.

Biggest teaching moment ▶ 11:05 Why scaling does not improve automated red teaming

Zico explains why frontier models cannot easily red team themselves due to built-in refusal training and the out-of-distribution nature of adversarial attacks.

The host holds their own ▶ 1:04:05 Interrogating AI compliance and SOC 2 frameworks

The host challenges the guests on why AI insurance is not ready, pressing on whether existing frameworks like SOC 2 or Sarbanes-Oxley could serve as immediate baselines.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Founding Gray Swan and Reframing AI Security Paradigms 5522 The host references historical adversarial ML work and Ian Goodfellow, showing relevant domain familiarity. Zico Kolter and Matt Fredrikson clearly reframe AI security away from traditional cybersecurity toward treating models as untrusted software components.
Gray Swan Arena and Automated Red Teaming with Shade 4632 The host asks whether models can red team themselves using on-policy reinforcement learning. Zico reframes this, explaining why frontier models struggle with self-red-teaming due to safety alignment refusals.
AI Intelligence, Red Teaming Experts, and Alien Cognition 3642 When the host suggests red teaming failures prove models do not model intelligence, Zico pushes back, arguing models represent an alien form of intelligence that falls for distinct failure modes.
Accelerating Mechanistic Interpretability Through Autonomous Coding Agents 5521 The host brings up mechanistic interpretability lagging capability scaling and cites Neil Nanda. Zico offers an optimistic reframe that autonomous coding agents can turn mechanistic interpretability into an automated science.
The Human Browser Agent Robustness Challenge Findings 4522 Matt explains the Human Browser Agent Robustness Challenge where humans and models were subjected to parallel phishing and prompt injection attacks. The host clarifies the double-blind setup and realistic threat modeling.
Evaluation Awareness, Model Sandbagging, and Capability Elicitation 4532 The host notes evaluation awareness risks causing false positives or negatives. Zico connects capability sandbagging and evaluation awareness to the necessity of adversarial elicitation.
Introducing Cygnal and Addressing the Robustness Scaling Paradox 4622 The host asks if robustness is an orthogonal guardrail layer. Zico and Matt explain the robustness scaling paradox, demonstrating with empirical data that scale does not improve adversarial resilience.
Enterprise AI Guardrails and Custom Policy Enforcement 5622 The host questions why enterprises cannot rely on open-source guards like Llama Guard. The guests explain that production deployments require highly configurable filtering for enterprise-specific policies.
Deconstructing the Lethal Trifecta and AI Threat Models 4521 The conversation walks through Simon Willison's lethal trifecta. Matt contrasts classical software debugging with the probabilistic nature of AI vulnerability mitigation.
Operationalizing Inbound and Outbound Tool-Call Security 4511 The host diagrams how Cygnal sits in the loop. The guests detail bidirectional monitoring, explaining why filtering both untrusted inputs and outbound tool-call actions is required.
Formal Verification and AI-Assisted Secure Software Engineering 5633 Matt discusses formal verification and obscure provably secure languages, which the host challenges as unrealistic for normal developers. Zico clarifies that coding agents will manage the verification layer under the hood.
Securing High-Risk Agents: OpenClaw and Computer Use 5632 The host highlights dangerous enterprise adoption of OpenClaw and computer use. The guests describe stress-testing these setups, finding abundant vulnerabilities and emphasizing isolation controls.
Designing Agent Identity, Access Control, and Digital Personas 4522 The host brings up agent-native authentication and identity delegation. Zico and Matt propose context-specific user personas rather than per-app credentials to avoid consent fatigue.
The Enterprise Expansion and Proactive AI Security Shift 3511 The host asks about upcoming industry developments. The guests share how enterprise demand shifted from reactive breach remediation to proactive pre-deployment guardrailing.
Private Arena Competitions and Adversarial Incentive Alignment 5522 The host questions how community red teaming avoids reward hacking and tests sensitive enterprise software. The guests describe private arenas under NDA and calibrated incentives.
AI Underwriting, Compliance Standards, and Preparing for Gray Swans 5633 The host probes AI insurance compliance and compares SOC 2 frameworks. Zico critiques SOC 2 as an accounting-driven standard, explaining what a rigorous technical standard needs to accomplish.

Statements from this episode (31)

Insight
Kolter: Concentrated Foundation Model Usage Creates Systemic Correlated Security Exploits
“And especially when there's the possibility of correlated failures, right? So it's not just that there's a lot of AI systems out there, it's that there's actually a few models that everyone is using. And if you find vulnerabilities in the agents that everyone …”
Zico Kolter Jun 22, 2026 ▶ 4:31
Disclosure
Fredrikson: Gray Swan's Discord red teaming community has 15,000 members
“It's a really great community. Like, 15,000 people come and hang out on the Discord server. Not all of them take part in every competition, but a lot of good data and good signal is provided to, you know, the upstream model developers through, through that com…”
Matt Fredrikson Jun 22, 2026 ▶ 9:30
Insight
Kolter: Frontier models fail at red teaming due to safety refusals
“So generally speaking, the issue with this is that frontier models are extremely bad at automated red teaming because they have a lot of safeguards built into them. So if you try to use them to jailbreak another model, they will actually refuse their safety tr…”
Zico Kolter Jun 22, 2026 ▶ 11:06
Insight
Kolter: Scaling model size does not automatically improve AI safety or red teaming
“Traditionally this has been an area where both in terms of safety models don't get better by just being bigger, unlike most other areas where models do get better by being bigger. Safety has not been like that traditionally. You know, you have to train them ex…”
Zico Kolter Jun 22, 2026 ▶ 11:28
Assertion Open · timeframe Jun 2027
Kolter: Gray Swan's Shade system outperforms human red teamers at breaking models
“However, one thing that we are finding, and this is actually, I think we're kind of crossing this point too. Is that in a lot of the latest experiments, we can do much better than people, than human red teamers now at breaking these models. When I say we, I me…”
Zico Kolter Jun 22, 2026 ▶ 12:14
Insight
Kolter: AI is an alien intelligence with completely distinct failure modes from humans
“It is clearly a different form of intelligence than people. It's some alien intelligence that is vastly different, and that difference is actually often brought out to a large degree by things like adversarial attacks and red teaming, because there are certain…”
Zico Kolter Jun 22, 2026 ▶ 15:32
Insight
Kolter: Full experimental observability has not produced fundamental understanding of AI
“It's like we could kind of run experiments on the brain, observe every neuron in it, reset its state to prior states, and run counterfactuals, none of which we can do with humans, and yet we still understand neither very well. Even with that, all that ability,…”
Zico Kolter Jun 22, 2026 ▶ 16:13
Opinion
Kolter: Mechanistic interpretability is not yet a real science
“The problem with McInterp is it's a lot, it's been about sort of testing small hypotheses. Hypothesis. And you know, you have a hypothesis, you'll find some small thing, you'll test that in isolation. But I don't think it's really become a science yet.”
Zico Kolter Jun 22, 2026 ▶ 17:29
Prediction Not checkable as stated
Kolter: Coding agents will revitalize mechanistic interpretability research
“Most fascinating things about coding agents actually is they can do a lot of experimentation in an automated fashion. Yeah. They will give new hope. They'll breathe new life into mechanter research.”
Zico Kolter Jun 22, 2026 ▶ 17:58
Assertion Open · timeframe Jun 2026
Fredrikson: Skilled red teamers phish human participants 60% to 70% of the time
“But for a skilled, like, human red teamer, they could fish the human participants, like, with the 60 to 70% success.”
Matt Fredrikson Jun 22, 2026 ▶ 22:20
Assertion Open · timeframe Jun 2026
Fredrikson: Top AI browser agents yielded only a handful of successful breaks
“There were a couple of models that seemed to be very, very robust, right? Like the red teamers found just a handful of successful breaks on them.”
Matt Fredrikson Jun 22, 2026 ▶ 22:29
Assertion Not checkable as stated
Fredrikson: Frontier AI models fall for simulated prompt injections humans would ignore
“While in these scenarios, humans found it very difficult to prompt inject the models, like we're aware of scenarios that a human would never fall for, that like Opus four seven would, right? Like a, you know, an email that comes to your inbox and it says somet…”
Matt Fredrikson Jun 22, 2026 ▶ 22:55
Insight
Fredrikson: Evaluation-aware AI models often execute harmful actions because it is a simulation
“If you make, if you're testing the model for robustness or safety, right? And it's aware that it's being tested because you've set things up in a very artificial way, right? Like the email addresses are at example.com. The webpage is clearly not a real webpage…”
Matt Fredrikson Jun 22, 2026 ▶ 23:32
Insight
Kolter: Adversarial Red Teaming Is Essential for True Capability Elicitation
“One of the most effective ways of doing capability elicitation is actually through some amount of what you would call red teaming, right? So if a model refuses a task because it thinks it's being evaluated, but it knows how to complete that task, getting it to…”
Zico Kolter Jun 22, 2026 ▶ 25:45
Insight
Fredrikson: Red Teaming Is Fundamentally a Mathematical Optimization Problem
“I mean, it really is an optimization problem, right? You have a, you know, an outcome that you want the model to exhibit, right? Now, how do I find the input, right? That, that gives me that output and you can sort of objectify that actually very mathematicall…”
Matt Fredrikson Jun 22, 2026 ▶ 26:38
Assertion Supported
Fredrikson: AI capability does not correlate with prompt injection resistance
“So this scatter plot on the right, right, is essentially like looking for a correlation between capability and attack success rate. So on the X axis, how capable is the model at, you know, GPQA diamond on, on the Y axis. How, how often, you know, were people s…”
Matt Fredrikson Jun 22, 2026 ▶ 30:09
Insight
Fredrikson: Prompt engineering cannot reliably enforce AI agent security policies
“Oftentimes people will try and prompt their way around it, like adjust the system prompt or like engineer the agent in a way where you're interjecting all the time and reminding it of what the original Goal and objective was, and that'll get you a little bit o…”
Matt Fredrikson Jun 22, 2026 ▶ 31:42
Prediction Not checkable as stated
Kolter: AI systems will probably not achieve provably zero vulnerabilities soon
“So the question is not trying to completely Kind of provably mitigate these things. That is arguably just a, it's a good goal, but just like zero bug software, we're probably not going to get there. At least not that soon.”
Zico Kolter Jun 22, 2026 ▶ 37:13
Assertion Supported
Kolter: Most AI Agents Resist Naive API Key Exfiltration Prompts
“Now, things that are that simple, to be clear, are covered at this point by most agents, right? You know, they all They, despite some issues, yeah, normal, normal sort of, you know, will not be that easily fooled by just push all my API keys to a public thing,…”
Zico Kolter Jun 22, 2026 ▶ 40:23
Insight
Fredrikson: Agent Guardrails Should Block Policy Violations, Not Injection Payloads
“If you parse some untrusted content and there is like a prompt injection, you know, something that's clearly trying to get the model to do a bad thing, you might be interested in knowing about that, but you don't necessarily like want your cloud code that you …”
Matt Fredrikson Jun 22, 2026 ▶ 41:05
Assertion Not checkable as stated
Fredrikson: Amazon excels at deploying formal software verification, Microsoft in research
“Microsoft historically has been pretty good about it too. More on the research side, Amazon is, is stellar and actually deploying a lot of this.”
Matt Fredrikson Jun 22, 2026 ▶ 42:52
Insight
Fredrikson: Formal software verification takes 10 to 20 times longer than Python
“The reason people don't do it is that it's not easy and it's not fun, right? It takes you like 10 or 20 times as long to like fight with the type checker, which is essentially like proving that you don't have a vulnerability as if, as it would if you just like…”
Matt Fredrikson Jun 22, 2026 ▶ 43:10
Prediction Not checkable as stated
Kolter: Security and science will explode as AI agents automate tedious verification
“So I think this is really sort of an underappreciated point that we're reaching this point, this sort of phase where a lot of security, a lot of science has this potential to kind of explode. Not because we're going to get better at it, but because agents can …”
Zico Kolter Jun 22, 2026 ▶ 46:01
Assertion Not checkable as stated
Fredrikson: Gray Swan Found Jailbreaks in Every OpenClaw User Trajectory Tested
“So we just have a bunch of trajectories of actual people using OpenClaw. And tons and tons of different scenarios and just threw shade at it and like found breaks for each and every one of them, right?”
Matt Fredrikson Jun 22, 2026 ▶ 47:36
Prediction Not checkable as stated
Kolter: AI agents inheriting user permissions by default will soon change
“So far, we are still a lot, in a lot of cases, operating on the condition that your agent has your permissions. Yeah. That is a very standard default. And I think that will be changed. I mean, your permissions may be in a sandbox, but still kind of your permis…”
Zico Kolter Jun 22, 2026 ▶ 53:00
Insight
Fredrikson: AI agent identity systems risk triggering automatic user consent fatigue
“One of the bigger challenges that people are going to face when they do start to roll out, like these agent identity sort of viewpoints and solutions is you run into that same kind of usability problem. Where like, what's the real recourse? Well, it stopped. I…”
Matt Fredrikson Jun 22, 2026 ▶ 53:48
Prediction Not checkable as stated
Kolter: Agent identity will evolve around user personas before fine-grained permissions
“I think in terms of how this will evolve, actually, I don't think it'll be per app, but I think what will happen first is people have different personas that they have, right? So you don't want your work life and your home email to be mixed up. Yeah. Right. A …”
Zico Kolter Jun 22, 2026 ▶ 54:19
Insight
Fredrikson: Enterprises refuse public red-teaming for pre-deployment AI agents
“Like enterprises are not willing to put up their pre-deployment agents on the arena for the general public to come hit. They're fine if it's, you know, 20 people that, that we've kind of handpicked from the arena.”
Matt Fredrikson Jun 22, 2026 ▶ 1:00:11
Opinion
Kolter: SOC 2 is not a great security compliance model
“So, so I think SOC II is not a great model. We'll just say, but it is a model.”
Zico Kolter Jun 22, 2026 ▶ 1:04:25
Prediction Not checkable as stated
Kolter: A major AI security incident is inevitable and foreseeable
“The name gray swan is sort of in reference to black swan events, which are things no one could see coming. A gray swan is an unlikely event that you can kind of see coming. And that's kind of where we are with all of this, right? This is going to happen. We kn…”
Zico Kolter Jun 22, 2026 ▶ 1:06:30
Assertion Not checkable as stated
Fredrikson: Unpublicized AI security breaches have already caused real damage
“We know that it has happened and it has caused real damage. That's the factor that's driven some people to us, right? Is they want protection from that.”
Matt Fredrikson Jun 22, 2026 ▶ 1:06:54
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.