Jun 22, 2026 · 1h 7m · latent-space
AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this in-depth interview, Gray Swan co-founders Zico Kolter and Matt Fredrikson discuss the emerging frontier of AI security, examining how automated red teaming, specialized runtime guardrails, and autonomous coding agents address the critical vulnerabilities of modern language models and enterprise agents.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Zico directly rejects the host's premise that failure on edge-case red teaming prompts shows models do not possess intelligence.
Hardest push from the hosts ▶ 43:50 Pushing back on formal verification in productionThe host openly doubts Matt's suggestion of using obscure formally verified languages, pointing out that engineers prefer plain English and accessible code.
Biggest teaching moment ▶ 11:05 Why scaling does not improve automated red teamingZico explains why frontier models cannot easily red team themselves due to built-in refusal training and the out-of-distribution nature of adversarial attacks.
The host holds their own ▶ 1:04:05 Interrogating AI compliance and SOC 2 frameworksThe host challenges the guests on why AI insurance is not ready, pressing on whether existing frameworks like SOC 2 or Sarbanes-Oxley could serve as immediate baselines.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Founding Gray Swan and Reframing AI Security Paradigms | 5 | 5 | 2 | 2 | The host references historical adversarial ML work and Ian Goodfellow, showing relevant domain familiarity. Zico Kolter and Matt Fredrikson clearly reframe AI security away from traditional cybersecurity toward treating models as untrusted software components. | |
| Gray Swan Arena and Automated Red Teaming with Shade | 4 | 6 | 3 | 2 | The host asks whether models can red team themselves using on-policy reinforcement learning. Zico reframes this, explaining why frontier models struggle with self-red-teaming due to safety alignment refusals. | |
| AI Intelligence, Red Teaming Experts, and Alien Cognition | 3 | 6 | 4 | 2 | When the host suggests red teaming failures prove models do not model intelligence, Zico pushes back, arguing models represent an alien form of intelligence that falls for distinct failure modes. | |
| Accelerating Mechanistic Interpretability Through Autonomous Coding Agents | 5 | 5 | 2 | 1 | The host brings up mechanistic interpretability lagging capability scaling and cites Neil Nanda. Zico offers an optimistic reframe that autonomous coding agents can turn mechanistic interpretability into an automated science. | |
| The Human Browser Agent Robustness Challenge Findings | 4 | 5 | 2 | 2 | Matt explains the Human Browser Agent Robustness Challenge where humans and models were subjected to parallel phishing and prompt injection attacks. The host clarifies the double-blind setup and realistic threat modeling. | |
| Evaluation Awareness, Model Sandbagging, and Capability Elicitation | 4 | 5 | 3 | 2 | The host notes evaluation awareness risks causing false positives or negatives. Zico connects capability sandbagging and evaluation awareness to the necessity of adversarial elicitation. | |
| Introducing Cygnal and Addressing the Robustness Scaling Paradox | 4 | 6 | 2 | 2 | The host asks if robustness is an orthogonal guardrail layer. Zico and Matt explain the robustness scaling paradox, demonstrating with empirical data that scale does not improve adversarial resilience. | |
| Enterprise AI Guardrails and Custom Policy Enforcement | 5 | 6 | 2 | 2 | The host questions why enterprises cannot rely on open-source guards like Llama Guard. The guests explain that production deployments require highly configurable filtering for enterprise-specific policies. | |
| Deconstructing the Lethal Trifecta and AI Threat Models | 4 | 5 | 2 | 1 | The conversation walks through Simon Willison's lethal trifecta. Matt contrasts classical software debugging with the probabilistic nature of AI vulnerability mitigation. | |
| Operationalizing Inbound and Outbound Tool-Call Security | 4 | 5 | 1 | 1 | The host diagrams how Cygnal sits in the loop. The guests detail bidirectional monitoring, explaining why filtering both untrusted inputs and outbound tool-call actions is required. | |
| Formal Verification and AI-Assisted Secure Software Engineering | 5 | 6 | 3 | 3 | Matt discusses formal verification and obscure provably secure languages, which the host challenges as unrealistic for normal developers. Zico clarifies that coding agents will manage the verification layer under the hood. | |
| Securing High-Risk Agents: OpenClaw and Computer Use | 5 | 6 | 3 | 2 | The host highlights dangerous enterprise adoption of OpenClaw and computer use. The guests describe stress-testing these setups, finding abundant vulnerabilities and emphasizing isolation controls. | |
| Designing Agent Identity, Access Control, and Digital Personas | 4 | 5 | 2 | 2 | The host brings up agent-native authentication and identity delegation. Zico and Matt propose context-specific user personas rather than per-app credentials to avoid consent fatigue. | |
| The Enterprise Expansion and Proactive AI Security Shift | 3 | 5 | 1 | 1 | The host asks about upcoming industry developments. The guests share how enterprise demand shifted from reactive breach remediation to proactive pre-deployment guardrailing. | |
| Private Arena Competitions and Adversarial Incentive Alignment | 5 | 5 | 2 | 2 | The host questions how community red teaming avoids reward hacking and tests sensitive enterprise software. The guests describe private arenas under NDA and calibrated incentives. | |
| AI Underwriting, Compliance Standards, and Preparing for Gray Swans | 5 | 6 | 3 | 3 | The host probes AI insurance compliance and compares SOC 2 frameworks. Zico critiques SOC 2 as an accounting-driven standard, explaining what a rigorous technical standard needs to accomplish. |