May 7, 2026 · 1h 16m · mad

OpenAI Board Member Zico Kolter: Modern AI Is Just 200 Lines of Code

Zico Kolter · 1h 1m spoken Matt Turck · 8m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The MAD Podcast, Carnegie Mellon University professor and OpenAI Board Member Zico Kolter joins Matt Turck to discuss AI safety governance, adversarial robustness and red-teaming, reinforcement learning frontiers, and the underlying architectural simplicity of modern AI models.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 12.2% of the talking time here. How this is scored →

Matt as informed peer 2.7 Guest teaching 4.3 Guest disagreement 0.7 Matt pushing back 0.8
05100:0020:0040:001:00:001:27–5:50 · Matt as informed peer 2/10 Podcast Title Bumper Matt asks introductory questions about how OpenAI's Safety and Security Committee operates in practice and mentions sitting on corporate boards. Zico explains the governance oversight role and how model release reviews work.5:50–9:46 · Matt as informed peer 1/10 The Preparedness Framework and Emerging AI Risks Matt asks for internal organizational details. Zico educates the host on OpenAI's Preparedness Framework, detailing dual-use risk categories like biological and cyber hazards.9:46–12:38 · Matt as informed peer 2/10 Industry Progress in AI Safety vs. Expanding Control Surfaces Matt asks whether safety progress is keeping pace with core capability progress. Zico details how models are measurably safer yet face exponentially expanding control surfaces as agents gain real-world autonomy.12:38–15:23 · Matt as informed peer 3/10 Model Scale, Robustness, and Red-Teaming Benchmarks Matt references Gray Swan's massive red-teaming competition to ask about capability versus vulnerability. Zico explains that model scale alone does not deliver robustness, requiring explicit safety training and system-level monitoring.15:23–19:39 · Matt as informed peer 1/10 The Four Pillars of AI Safety Risks Matt asks where safety issues originate. Zico lays out a clear four-pillar taxonomy ranging from simple model mistakes to existential loss of control.19:39–26:20 · Matt as informed peer 4/10 Debunking Accelerationism vs. Doomerism and Historical AI Pauses Matt probes the accelerationism versus doomerism debate and references the historic 6-month pause letter. Zico rejects both polar labels as pejorative and corrects Matt's timeline on which model era spawned the letter.26:20–34:33 · Matt as informed peer 3/10 Global AI Safety Institutes and International Cooperation Matt asks about global AI safety institutes and pivots to Zico's background. Zico recounts attending OpenAI's 2015 launch party and their early, contrarian bet on scale.34:33–38:44 · Matt as informed peer 4/10 CMU's AI Legacy and Academia's Role in the AI Era Matt demonstrates historical context by citing notable CMU leaders like Andrew Moore and Tom Mitchell while questioning how academia competes with tech giants. Zico explains academia's need to pivot toward agentic research and fundamental science.38:44–40:44 · Matt as informed peer 2/10 Founding Gray Swan AI and Enterprise Security Offerings Matt asks Zico to introduce his startup Gray Swan AI. Zico explains automated red teaming tools and enterprise firewalls for agentic workflows.40:44–43:15 · Matt as informed peer 2/10 Defining AI Security vs. AI for Security Matt asks for clarification between safety and security. Zico delineates AI security from AI for security, emphasizing worst-case adversarial resilience over average-case benchmarks.43:15–49:19 · Matt as informed peer 3/10 The Breakthrough GCG Paper and Universal Transferable Jailbreaks Matt asks about Zico's landmark 2023 GCG paper. Zico educates the host on how automated suffix optimization revealed universal, transferable jailbreaks across closed commercial models.49:19–52:32 · Matt as informed peer 3/10 Lab Responses, Safety Classifiers, and Reasoning Model Resilience Matt asks how frontier labs responded to GCG attacks. Zico explains multi-layered defense stacks and why reasoning models with chain-of-thought traces are inherently harder to manipulate.52:32–58:38 · Matt as informed peer 3/10 Advanced Multi-Query Attack Strategies Matt asks how builders should approach agent security. Zico explains complex attack vectors like indirect prompt injection when agents process external untrusted web or email inputs.58:38–1:02:32 · Matt as informed peer 3/10 Deploying Agents Safely in Production Matt challenges Zico directly on whether agents should be deployed in production given security vulnerabilities. Zico defends deployment readiness subject to proper sandboxing before pivoting to mechanistic interpretability.1:02:32–1:06:49 · Matt as informed peer 3/10 Two-Year Outlook on AI Safety and Autonomous Capabilities Matt asks for a multi-year outlook on safety and frontier developments. Zico explains why reinforcement learning on model-generated outputs refutes the idea that synthetic data causes model collapse.1:06:49–1:09:29 · Matt as informed peer 3/10 Continual Learning and AI Breakthroughs Matt asks about continual learning and post-transformer architectures. Zico offers a contrarian view that specific network architectures matter much less than the core discovery of scaling sequence prediction.1:09:29–1:12:54 · Matt as informed peer 3/10 Career Advice for PhD Students in AI Matt asks what advice Zico gives his PhD students. Zico emphasizes taking risks and encourages young researchers to ignore established academic orthodoxy.1:12:54–1:16:07 · Matt as informed peer 3/10 Building Modern AI in 200 Lines of Code Matt asks whether Zico's 200-line LLM codebase includes pre-training and RL. Zico details how the core mathematical framework of modern AI is shockingly minimal, whereas real-world complexity lies in data and GPU infrastructure.1:27–5:50 · Guest teaching 3/10 Podcast Title Bumper Matt asks introductory questions about how OpenAI's Safety and Security Committee operates in practice and mentions sitting on corporate boards. Zico explains the governance oversight role and how model release reviews work.5:50–9:46 · Guest teaching 4/10 The Preparedness Framework and Emerging AI Risks Matt asks for internal organizational details. Zico educates the host on OpenAI's Preparedness Framework, detailing dual-use risk categories like biological and cyber hazards.9:46–12:38 · Guest teaching 4/10 Industry Progress in AI Safety vs. Expanding Control Surfaces Matt asks whether safety progress is keeping pace with core capability progress. Zico details how models are measurably safer yet face exponentially expanding control surfaces as agents gain real-world autonomy.12:38–15:23 · Guest teaching 5/10 Model Scale, Robustness, and Red-Teaming Benchmarks Matt references Gray Swan's massive red-teaming competition to ask about capability versus vulnerability. Zico explains that model scale alone does not deliver robustness, requiring explicit safety training and system-level monitoring.15:23–19:39 · Guest teaching 5/10 The Four Pillars of AI Safety Risks Matt asks where safety issues originate. Zico lays out a clear four-pillar taxonomy ranging from simple model mistakes to existential loss of control.19:39–26:20 · Guest teaching 3/10 Debunking Accelerationism vs. Doomerism and Historical AI Pauses Matt probes the accelerationism versus doomerism debate and references the historic 6-month pause letter. Zico rejects both polar labels as pejorative and corrects Matt's timeline on which model era spawned the letter.26:20–34:33 · Guest teaching 3/10 Global AI Safety Institutes and International Cooperation Matt asks about global AI safety institutes and pivots to Zico's background. Zico recounts attending OpenAI's 2015 launch party and their early, contrarian bet on scale.34:33–38:44 · Guest teaching 4/10 CMU's AI Legacy and Academia's Role in the AI Era Matt demonstrates historical context by citing notable CMU leaders like Andrew Moore and Tom Mitchell while questioning how academia competes with tech giants. Zico explains academia's need to pivot toward agentic research and fundamental science.38:44–40:44 · Guest teaching 3/10 Founding Gray Swan AI and Enterprise Security Offerings Matt asks Zico to introduce his startup Gray Swan AI. Zico explains automated red teaming tools and enterprise firewalls for agentic workflows.40:44–43:15 · Guest teaching 5/10 Defining AI Security vs. AI for Security Matt asks for clarification between safety and security. Zico delineates AI security from AI for security, emphasizing worst-case adversarial resilience over average-case benchmarks.43:15–49:19 · Guest teaching 6/10 The Breakthrough GCG Paper and Universal Transferable Jailbreaks Matt asks about Zico's landmark 2023 GCG paper. Zico educates the host on how automated suffix optimization revealed universal, transferable jailbreaks across closed commercial models.49:19–52:32 · Guest teaching 5/10 Lab Responses, Safety Classifiers, and Reasoning Model Resilience Matt asks how frontier labs responded to GCG attacks. Zico explains multi-layered defense stacks and why reasoning models with chain-of-thought traces are inherently harder to manipulate.52:32–58:38 · Guest teaching 6/10 Advanced Multi-Query Attack Strategies Matt asks how builders should approach agent security. Zico explains complex attack vectors like indirect prompt injection when agents process external untrusted web or email inputs.58:38–1:02:32 · Guest teaching 4/10 Deploying Agents Safely in Production Matt challenges Zico directly on whether agents should be deployed in production given security vulnerabilities. Zico defends deployment readiness subject to proper sandboxing before pivoting to mechanistic interpretability.1:02:32–1:06:49 · Guest teaching 5/10 Two-Year Outlook on AI Safety and Autonomous Capabilities Matt asks for a multi-year outlook on safety and frontier developments. Zico explains why reinforcement learning on model-generated outputs refutes the idea that synthetic data causes model collapse.1:06:49–1:09:29 · Guest teaching 5/10 Continual Learning and AI Breakthroughs Matt asks about continual learning and post-transformer architectures. Zico offers a contrarian view that specific network architectures matter much less than the core discovery of scaling sequence prediction.1:09:29–1:12:54 · Guest teaching 3/10 Career Advice for PhD Students in AI Matt asks what advice Zico gives his PhD students. Zico emphasizes taking risks and encourages young researchers to ignore established academic orthodoxy.1:12:54–1:16:07 · Guest teaching 5/10 Building Modern AI in 200 Lines of Code Matt asks whether Zico's 200-line LLM codebase includes pre-training and RL. Zico details how the core mathematical framework of modern AI is shockingly minimal, whereas real-world complexity lies in data and GPU infrastructure.1:27–5:50 · Guest disagreement 1/10 Podcast Title Bumper Matt asks introductory questions about how OpenAI's Safety and Security Committee operates in practice and mentions sitting on corporate boards. Zico explains the governance oversight role and how model release reviews work.5:50–9:46 · Guest disagreement 0/10 The Preparedness Framework and Emerging AI Risks Matt asks for internal organizational details. Zico educates the host on OpenAI's Preparedness Framework, detailing dual-use risk categories like biological and cyber hazards.9:46–12:38 · Guest disagreement 1/10 Industry Progress in AI Safety vs. Expanding Control Surfaces Matt asks whether safety progress is keeping pace with core capability progress. Zico details how models are measurably safer yet face exponentially expanding control surfaces as agents gain real-world autonomy.12:38–15:23 · Guest disagreement 1/10 Model Scale, Robustness, and Red-Teaming Benchmarks Matt references Gray Swan's massive red-teaming competition to ask about capability versus vulnerability. Zico explains that model scale alone does not deliver robustness, requiring explicit safety training and system-level monitoring.15:23–19:39 · Guest disagreement 1/10 The Four Pillars of AI Safety Risks Matt asks where safety issues originate. Zico lays out a clear four-pillar taxonomy ranging from simple model mistakes to existential loss of control.19:39–26:20 · Guest disagreement 2/10 Debunking Accelerationism vs. Doomerism and Historical AI Pauses Matt probes the accelerationism versus doomerism debate and references the historic 6-month pause letter. Zico rejects both polar labels as pejorative and corrects Matt's timeline on which model era spawned the letter.26:20–34:33 · Guest disagreement 0/10 Global AI Safety Institutes and International Cooperation Matt asks about global AI safety institutes and pivots to Zico's background. Zico recounts attending OpenAI's 2015 launch party and their early, contrarian bet on scale.34:33–38:44 · Guest disagreement 1/10 CMU's AI Legacy and Academia's Role in the AI Era Matt demonstrates historical context by citing notable CMU leaders like Andrew Moore and Tom Mitchell while questioning how academia competes with tech giants. Zico explains academia's need to pivot toward agentic research and fundamental science.38:44–40:44 · Guest disagreement 0/10 Founding Gray Swan AI and Enterprise Security Offerings Matt asks Zico to introduce his startup Gray Swan AI. Zico explains automated red teaming tools and enterprise firewalls for agentic workflows.40:44–43:15 · Guest disagreement 1/10 Defining AI Security vs. AI for Security Matt asks for clarification between safety and security. Zico delineates AI security from AI for security, emphasizing worst-case adversarial resilience over average-case benchmarks.43:15–49:19 · Guest disagreement 0/10 The Breakthrough GCG Paper and Universal Transferable Jailbreaks Matt asks about Zico's landmark 2023 GCG paper. Zico educates the host on how automated suffix optimization revealed universal, transferable jailbreaks across closed commercial models.49:19–52:32 · Guest disagreement 0/10 Lab Responses, Safety Classifiers, and Reasoning Model Resilience Matt asks how frontier labs responded to GCG attacks. Zico explains multi-layered defense stacks and why reasoning models with chain-of-thought traces are inherently harder to manipulate.52:32–58:38 · Guest disagreement 0/10 Advanced Multi-Query Attack Strategies Matt asks how builders should approach agent security. Zico explains complex attack vectors like indirect prompt injection when agents process external untrusted web or email inputs.58:38–1:02:32 · Guest disagreement 2/10 Deploying Agents Safely in Production Matt challenges Zico directly on whether agents should be deployed in production given security vulnerabilities. Zico defends deployment readiness subject to proper sandboxing before pivoting to mechanistic interpretability.1:02:32–1:06:49 · Guest disagreement 1/10 Two-Year Outlook on AI Safety and Autonomous Capabilities Matt asks for a multi-year outlook on safety and frontier developments. Zico explains why reinforcement learning on model-generated outputs refutes the idea that synthetic data causes model collapse.1:06:49–1:09:29 · Guest disagreement 2/10 Continual Learning and AI Breakthroughs Matt asks about continual learning and post-transformer architectures. Zico offers a contrarian view that specific network architectures matter much less than the core discovery of scaling sequence prediction.1:09:29–1:12:54 · Guest disagreement 0/10 Career Advice for PhD Students in AI Matt asks what advice Zico gives his PhD students. Zico emphasizes taking risks and encourages young researchers to ignore established academic orthodoxy.1:12:54–1:16:07 · Guest disagreement 0/10 Building Modern AI in 200 Lines of Code Matt asks whether Zico's 200-line LLM codebase includes pre-training and RL. Zico details how the core mathematical framework of modern AI is shockingly minimal, whereas real-world complexity lies in data and GPU infrastructure.1:27–5:50 · Matt pushing back 1/10 Podcast Title Bumper Matt asks introductory questions about how OpenAI's Safety and Security Committee operates in practice and mentions sitting on corporate boards. Zico explains the governance oversight role and how model release reviews work.5:50–9:46 · Matt pushing back 0/10 The Preparedness Framework and Emerging AI Risks Matt asks for internal organizational details. Zico educates the host on OpenAI's Preparedness Framework, detailing dual-use risk categories like biological and cyber hazards.9:46–12:38 · Matt pushing back 1/10 Industry Progress in AI Safety vs. Expanding Control Surfaces Matt asks whether safety progress is keeping pace with core capability progress. Zico details how models are measurably safer yet face exponentially expanding control surfaces as agents gain real-world autonomy.12:38–15:23 · Matt pushing back 1/10 Model Scale, Robustness, and Red-Teaming Benchmarks Matt references Gray Swan's massive red-teaming competition to ask about capability versus vulnerability. Zico explains that model scale alone does not deliver robustness, requiring explicit safety training and system-level monitoring.15:23–19:39 · Matt pushing back 0/10 The Four Pillars of AI Safety Risks Matt asks where safety issues originate. Zico lays out a clear four-pillar taxonomy ranging from simple model mistakes to existential loss of control.19:39–26:20 · Matt pushing back 3/10 Debunking Accelerationism vs. Doomerism and Historical AI Pauses Matt probes the accelerationism versus doomerism debate and references the historic 6-month pause letter. Zico rejects both polar labels as pejorative and corrects Matt's timeline on which model era spawned the letter.26:20–34:33 · Matt pushing back 0/10 Global AI Safety Institutes and International Cooperation Matt asks about global AI safety institutes and pivots to Zico's background. Zico recounts attending OpenAI's 2015 launch party and their early, contrarian bet on scale.34:33–38:44 · Matt pushing back 1/10 CMU's AI Legacy and Academia's Role in the AI Era Matt demonstrates historical context by citing notable CMU leaders like Andrew Moore and Tom Mitchell while questioning how academia competes with tech giants. Zico explains academia's need to pivot toward agentic research and fundamental science.38:44–40:44 · Matt pushing back 0/10 Founding Gray Swan AI and Enterprise Security Offerings Matt asks Zico to introduce his startup Gray Swan AI. Zico explains automated red teaming tools and enterprise firewalls for agentic workflows.40:44–43:15 · Matt pushing back 0/10 Defining AI Security vs. AI for Security Matt asks for clarification between safety and security. Zico delineates AI security from AI for security, emphasizing worst-case adversarial resilience over average-case benchmarks.43:15–49:19 · Matt pushing back 0/10 The Breakthrough GCG Paper and Universal Transferable Jailbreaks Matt asks about Zico's landmark 2023 GCG paper. Zico educates the host on how automated suffix optimization revealed universal, transferable jailbreaks across closed commercial models.49:19–52:32 · Matt pushing back 1/10 Lab Responses, Safety Classifiers, and Reasoning Model Resilience Matt asks how frontier labs responded to GCG attacks. Zico explains multi-layered defense stacks and why reasoning models with chain-of-thought traces are inherently harder to manipulate.52:32–58:38 · Matt pushing back 0/10 Advanced Multi-Query Attack Strategies Matt asks how builders should approach agent security. Zico explains complex attack vectors like indirect prompt injection when agents process external untrusted web or email inputs.58:38–1:02:32 · Matt pushing back 4/10 Deploying Agents Safely in Production Matt challenges Zico directly on whether agents should be deployed in production given security vulnerabilities. Zico defends deployment readiness subject to proper sandboxing before pivoting to mechanistic interpretability.1:02:32–1:06:49 · Matt pushing back 0/10 Two-Year Outlook on AI Safety and Autonomous Capabilities Matt asks for a multi-year outlook on safety and frontier developments. Zico explains why reinforcement learning on model-generated outputs refutes the idea that synthetic data causes model collapse.1:06:49–1:09:29 · Matt pushing back 1/10 Continual Learning and AI Breakthroughs Matt asks about continual learning and post-transformer architectures. Zico offers a contrarian view that specific network architectures matter much less than the core discovery of scaling sequence prediction.1:09:29–1:12:54 · Matt pushing back 0/10 Career Advice for PhD Students in AI Matt asks what advice Zico gives his PhD students. Zico emphasizes taking risks and encourages young researchers to ignore established academic orthodoxy.1:12:54–1:16:07 · Matt pushing back 1/10 Building Modern AI in 200 Lines of Code Matt asks whether Zico's 200-line LLM codebase includes pre-training and RL. Zico details how the core mathematical framework of modern AI is shockingly minimal, whereas real-world complexity lies in data and GPU infrastructure.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 38.3% · guest 61.7%0:00 · Matt 38.3% · guest 61.7%3:00 · Matt 20.5% · guest 79.5%3:00 · Matt 20.5% · guest 79.5%6:00 · Matt 0.1% · guest 99.9%6:00 · Matt 0.1% · guest 99.9%9:00 · Matt 16.9% · guest 83.1%9:00 · Matt 16.9% · guest 83.1%12:00 · Matt 11.6% · guest 88.4%12:00 · Matt 11.6% · guest 88.4%15:00 · Matt 6.4% · guest 93.6%15:00 · Matt 6.4% · guest 93.6%18:00 · Matt 9.2% · guest 90.8%18:00 · Matt 9.2% · guest 90.8%21:00 · Matt 13.6% · guest 86.4%21:00 · Matt 13.6% · guest 86.4%24:00 · Matt 18.9% · guest 81.1%24:00 · Matt 18.9% · guest 81.1%27:00 · Matt 13.5% · guest 86.5%27:00 · Matt 13.5% · guest 86.5%30:00 · Matt 11.3% · guest 88.7%30:00 · Matt 11.3% · guest 88.7%33:00 · Matt 18.6% · guest 81.4%33:00 · Matt 18.6% · guest 81.4%36:00 · Matt 5.4% · guest 94.6%36:00 · Matt 5.4% · guest 94.6%39:00 · Matt 12.4% · guest 87.6%39:00 · Matt 12.4% · guest 87.6%42:00 · Matt 12.7% · guest 87.3%42:00 · Matt 12.7% · guest 87.3%45:00 · Matt 0% · guest 100%45:00 · Matt 0% · guest 100%48:00 · Matt 8.2% · guest 91.8%48:00 · Matt 8.2% · guest 91.8%51:00 · Matt 7.1% · guest 92.9%51:00 · Matt 7.1% · guest 92.9%54:00 · Matt 11.2% · guest 88.8%54:00 · Matt 11.2% · guest 88.8%57:00 · Matt 14.7% · guest 85.3%57:00 · Matt 14.7% · guest 85.3%1:00:00 · Matt 6.8% · guest 93.2%1:00:00 · Matt 6.8% · guest 93.2%1:03:00 · Matt 14.1% · guest 85.9%1:03:00 · Matt 14.1% · guest 85.9%1:06:00 · Matt 7% · guest 93%1:06:00 · Matt 7% · guest 93%1:09:00 · Matt 13% · guest 87%1:09:00 · Matt 13% · guest 87%1:12:00 · Matt 3.3% · guest 96.7%1:12:00 · Matt 3.3% · guest 96.7%1:15:00 · Matt 31.9% · guest 68.1%1:15:00 · Matt 31.9% · guest 68.1%
Sharpest disagreement ▶ 19:55 Rejection of Doomer vs. Accelerationist Framing

Zico strongly pushes back against the terms doomer and accelerationist, calling them inherently dismissive and pejorative labels used by both sides to avoid nuanced safety discussions.

Hardest push from Matt ▶ 58:48 Challenging Production Readiness of AI Agents

Matt refuses to accept Zico's initial confirmation that agents are in production, interrupting to ask specifically whether they should be deployed from a strict security standpoint.

Biggest teaching moment ▶ 47:40 Discovery of Universal Transferable Jailbreaks

Zico educates Matt on the unexpected mechanics of GCG jailbreaks, explaining how adversarial suffix tokens optimized on open-source models surprisingly transfer to closed commercial models.

Matt holds his own ▶ 34:33 Contextualizing CMU's AI Legacy Against Industry Pull

Matt displays impressive industry knowledge by citing foundational CMU figures like Tom Mitchell and Andrew Moore alongside the Robotics Institute to challenge Zico on academic relevance in the compute era.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Podcast Title Bumper 2311 Matt asks introductory questions about how OpenAI's Safety and Security Committee operates in practice and mentions sitting on corporate boards. Zico explains the governance oversight role and how model release reviews work.
The Preparedness Framework and Emerging AI Risks 1400 Matt asks for internal organizational details. Zico educates the host on OpenAI's Preparedness Framework, detailing dual-use risk categories like biological and cyber hazards.
Industry Progress in AI Safety vs. Expanding Control Surfaces 2411 Matt asks whether safety progress is keeping pace with core capability progress. Zico details how models are measurably safer yet face exponentially expanding control surfaces as agents gain real-world autonomy.
Model Scale, Robustness, and Red-Teaming Benchmarks 3511 Matt references Gray Swan's massive red-teaming competition to ask about capability versus vulnerability. Zico explains that model scale alone does not deliver robustness, requiring explicit safety training and system-level monitoring.
The Four Pillars of AI Safety Risks 1510 Matt asks where safety issues originate. Zico lays out a clear four-pillar taxonomy ranging from simple model mistakes to existential loss of control.
Debunking Accelerationism vs. Doomerism and Historical AI Pauses 4323 Matt probes the accelerationism versus doomerism debate and references the historic 6-month pause letter. Zico rejects both polar labels as pejorative and corrects Matt's timeline on which model era spawned the letter.
Global AI Safety Institutes and International Cooperation 3300 Matt asks about global AI safety institutes and pivots to Zico's background. Zico recounts attending OpenAI's 2015 launch party and their early, contrarian bet on scale.
CMU's AI Legacy and Academia's Role in the AI Era 4411 Matt demonstrates historical context by citing notable CMU leaders like Andrew Moore and Tom Mitchell while questioning how academia competes with tech giants. Zico explains academia's need to pivot toward agentic research and fundamental science.
Founding Gray Swan AI and Enterprise Security Offerings 2300 Matt asks Zico to introduce his startup Gray Swan AI. Zico explains automated red teaming tools and enterprise firewalls for agentic workflows.
Defining AI Security vs. AI for Security 2510 Matt asks for clarification between safety and security. Zico delineates AI security from AI for security, emphasizing worst-case adversarial resilience over average-case benchmarks.
The Breakthrough GCG Paper and Universal Transferable Jailbreaks 3600 Matt asks about Zico's landmark 2023 GCG paper. Zico educates the host on how automated suffix optimization revealed universal, transferable jailbreaks across closed commercial models.
Lab Responses, Safety Classifiers, and Reasoning Model Resilience 3501 Matt asks how frontier labs responded to GCG attacks. Zico explains multi-layered defense stacks and why reasoning models with chain-of-thought traces are inherently harder to manipulate.
Advanced Multi-Query Attack Strategies 3600 Matt asks how builders should approach agent security. Zico explains complex attack vectors like indirect prompt injection when agents process external untrusted web or email inputs.
Deploying Agents Safely in Production 3424 Matt challenges Zico directly on whether agents should be deployed in production given security vulnerabilities. Zico defends deployment readiness subject to proper sandboxing before pivoting to mechanistic interpretability.
Two-Year Outlook on AI Safety and Autonomous Capabilities 3510 Matt asks for a multi-year outlook on safety and frontier developments. Zico explains why reinforcement learning on model-generated outputs refutes the idea that synthetic data causes model collapse.
Continual Learning and AI Breakthroughs 3521 Matt asks about continual learning and post-transformer architectures. Zico offers a contrarian view that specific network architectures matter much less than the core discovery of scaling sequence prediction.
Career Advice for PhD Students in AI 3300 Matt asks what advice Zico gives his PhD students. Zico emphasizes taking risks and encourages young researchers to ignore established academic orthodoxy.
Building Modern AI in 200 Lines of Code 3501 Matt asks whether Zico's 200-line LLM codebase includes pre-training and RL. Zico details how the core mathematical framework of modern AI is shockingly minimal, whereas real-world complexity lies in data and GPU infrastructure.

Statements from this episode (35)

Disclosure
OpenAI's Safety Committee has authority to delay model releases
“In the case essentially where we have more questions, we can delay model release if we feel that we need to understand that, that better.”
Zico Kolter May 7, 2026 ▶ 3:46
Opinion
Kolter: AI companies should establish board-level safety and security committees
“I think it's actually very important that AI companies start to establish similar governance policies, because this is something that requires that level of just oversight and of assurance. It is a becoming You know, a massive industry, and just like there are…”
Zico Kolter May 7, 2026 ▶ 4:54
Insight
Kolter: AI safety focus is shifting from single models to ecosystems
“I think actually one of the big trends we're seeing is a lot of safety is moving from the model level to the ecosystem level and talking about, you know, what's not one model capable of, but what's AI broadly capable of.”
Zico Kolter May 7, 2026 ▶ 9:18
Assertion Not checkable as stated
Zico Kolter: AI models are objectively safer than a year ago
“Models, definitely, objectively, I would say, in a lot of scenarios we can measure, are safer than they were a year ago.”
Zico Kolter May 7, 2026 ▶ 10:21
Assertion Not checkable as stated
Zico Kolter: Agentic AI systems have far more autonomy than last year
“The amount of autonomy granted to agentic systems now is far greater than a year ago.”
Zico Kolter May 7, 2026 ▶ 11:25
Insight
Kolter: AI models do not get safer automatically by scaling up
“You can't just sort of trust models to get safer by getting bigger. You have to put in the work to actually make them safer. And this is, I think what a lot of AI companies are investing in. This is why we in fact do have models that are improving on these dim…”
Zico Kolter May 7, 2026 ▶ 15:03
Disclosure
Kolter: Calculating a P(doom) for AI is a flawed concept
“I have never expressed a P. Doom and things like this. I just think it's a very weird concept as if the world is some stochastic set of dice that you can roll multiple times and that we don't have direct influence over this.”
Zico Kolter May 7, 2026 ▶ 20:28
Insight
Kolter: AI safety is solved through frontier interaction, not pauses
“I think the way you solve things is through, through ongoing exploration of what's happening and through, through interaction with the frontier.”
Zico Kolter May 7, 2026 ▶ 26:11
Insight
Kolter: Renaming AI Safety Summit to AI Action Summit shows political shift
“The, you know, AI safety conference was, or AI safety summit was renamed the AI action summit or something is, has some significance actually in terms of the sort of taking temperature of where the world is politically.”
Zico Kolter May 7, 2026 ▶ 27:06
Disclosure
Kolter tried recruiting Schulman and Karpathy to CMU before OpenAI launched
“I was trying to get both John Schulman and Andre Karpathy to apply for faculty jobs at CMU, and I was trying to understand where they were, if they were going to apply, what they're going to be like, and they said, no, I think I'm going to be doing this startu…”
Zico Kolter May 7, 2026 ▶ 31:34
Assertion Not checkable as stated
OpenAI differentiated early on by prioritizing model scale over new methods
“About opening up early on is that they always had this bet on scale. In a time where I think that was looked upon very suspiciously that, oh, if you, the thought somehow that we had all the methods already, and all you had to do was scale them up that mindset …”
Zico Kolter May 7, 2026 ▶ 32:46
Opinion
Kolter: Robotics AI is not yet ready for pure compute scaling
“Certain fields. I think things like robotics is still one. I don't think we're quite at the, let's just scale it up level with robotics yet. Some companies might argue we are. I don't think we are. I think we're still in the let's explore methods to find the r…”
Zico Kolter May 7, 2026 ▶ 37:38
Prediction Not checkable as stated
Kolter: Universities will drive AI breakthroughs in math and science
“There's going to be a whole lot of breakthroughs happening with AI enablement in math and basic science, those kinds of things. Universities, I think will play a foundational role in shaping that future.”
Zico Kolter May 7, 2026 ▶ 38:27
Insight
Kolter: AI evals test average performance; security tests worst-case performance
“Most evaluations are done kind of in a, They measure expected value, basically. They measure sort of how well does it work on average, and security measures how well does it work in the worst case.”
Zico Kolter May 7, 2026 ▶ 42:20
Assertion Supported
Kolter: Adversarial prompts optimized on open-source LLMs break commercial models
“Once we had done that, we found that when you had these weird terms that you sort of flipped around to optimize one, to optimize the response for one model, you could just take those same exact strings you would optimize, paste them into a commercial model, an…”
Zico Kolter May 7, 2026 ▶ 47:51
Insight
Kolter: Reasoning models are much harder to jailbreak via probability optimization
“Reasoning models were much more effective because you can't really do the same trick of optimizing for a probability with a reasoning model that has a whole trace of reasoning that happens in the middle and kind of reflect a bit more. So it's much harder to br…”
Zico Kolter May 7, 2026 ▶ 49:51
Assertion Not checkable as stated
Kolter: State-of-the-art AI security combines classifiers, safety training, and opsec
“What they look like is basically classifiers on input. So you'll read what the, what a user types in classifiers on things like tool responses to classifiers. And when I say classifier, I just mean things that will read texts and kind of classify whether or no…”
Zico Kolter May 7, 2026 ▶ 51:04
Insight
Kolter: Jailbreaking modern AI safety systems requires complex multi-query attacks
“But they are, they require that degree of complexity to really jailbreak modern systems in a, for information that has this sensitivity to it.”
Zico Kolter May 7, 2026 ▶ 54:13
Insight
Kolter: Prompt injection introduces data exfiltration risks to AI agents
“Things like prompt injection are really a new security vulnerability for AI agents, and they mean that your risk is not just that you could have some, the model says something mean to you or something like that. Or even they could just write bad code. It could…”
Zico Kolter May 7, 2026 ▶ 57:05
Insight
Kolter: AI agent security requires managing permissions alongside manipulation risks
“AI security of agents is this interaction between what can the agent be manipulated into doing? What might it do accidentally? And what credentials or access does it have to really affect change?”
Zico Kolter May 7, 2026 ▶ 58:13
Opinion
Kolter: AI agent benefits outweigh security risks if deployed with proper guardrails
“Yes, I think so, actually. I think if you run with proper guardrails, you know, we release guardrails for coding agents, for example. If you're on proper guardrails with proper sandboxing, and right now, yes, you probably also take some care to be a little bit…”
Zico Kolter May 7, 2026 ▶ 58:56
Disclosure
Kolter no longer writes code manually, relying entirely on AI agents
“I don't write code anymore. I do all my work now, and I do lots of, you know, I still do some research, right? It's entirely telling Codex what to do.”
Zico Kolter May 7, 2026 ▶ 59:29
Insight
Kolter: AI coding agents are extremely good mechanistic interpretability researchers
“Coding agents are extremely good Mechinterp researchers.”
Zico Kolter May 7, 2026 ▶ 1:01:11
Prediction Not checkable as stated
Kolter: AI agents might turn mechanistic interpretability into a science
“I think that we actually might finally be able to make more what I would consider a science of this through essentially leveraging mass research by agents deployed for this problem.”
Zico Kolter May 7, 2026 ▶ 1:02:13
Assertion Not checkable as stated
Zico Kolter: Reinforcement learning is now the foundation of all AI post-training
“RL is now the foundation of really all post training. It's all done by RL.”
Zico Kolter May 7, 2026 ▶ 1:04:27
Insight
The vast majority of modern AI intelligence comes from self-training
“I don't think people have properly internalized the fact that the vast majority of intelligence comes from self training effectively.”
Zico Kolter May 7, 2026 ▶ 1:05:53
Prediction Not checkable as stated
Zico Kolter: Current AI trajectory will yield capable systems without breakthroughs
“I think the current trajectory we're on is going to get us, even if there were no more breakthroughs, I think, you know, with the minor additions that we are doing right now, we will get to incredibly capable systems, even if we were to freeze things right now…”
Zico Kolter May 7, 2026 ▶ 1:06:34
Insight
Kolter: Major AI breakthroughs require both massive scale and luck
“Reasoning models were the next big breakthrough. Those are rare. They do take kind of a, you know, both, both a massive scale and kind of a bit of luck to get there.”
Zico Kolter May 7, 2026 ▶ 1:07:44
Opinion
Kolter: Neural network architectures matter less than commonly believed
“I actually think architectures don't matter as much as everyone else thinks they do.”
Zico Kolter May 7, 2026 ▶ 1:08:15
What-if
AI capabilities would have been achieved even without inventing transformers
“I think if we hadn't invented the transformer, we would have gotten there with whatever LSTM you know, state space model, whatever, anything else people were developing, we would have gotten there.”
Zico Kolter May 7, 2026 ▶ 1:08:19
Opinion
Text model scaling is one of humanity's top scientific discoveries
“The discovery That when you train big enough models on lots of text, and then turn, and then a little bit of additional sort of, you know, fine-tuning text, and then turn them loose to generate, that this generates long-form coherent thought. That was probably…”
Zico Kolter May 7, 2026 ▶ 1:09:06
Insight
Kolter: Scientific progress occurs when young researchers ignore old guard beliefs
“Basically progress happens when The current crop of young researchers ignores the things they've been taught that the old guard believes.”
Zico Kolter May 7, 2026 ▶ 1:10:24
Assertion Supported
CMU undergrad AI course has students build an LLM from scratch
“You build a LLM completely from scratch. You use PyTorch, but you build one from scratch that, you know, can be a chatbot. You train it on data. You RL it to solve math problems with tool calls. You do all of this. And this is a undergrad level course.”
Zico Kolter May 7, 2026 ▶ 1:11:36
Assertion Not checkable as stated
A complete modern LLM takes only 200 to 300 lines of Python
“You have this code this code to build a complete large language model that can train on a large data set and learn to speak, runs on GPUs yes, eventually is trained with RL and tool calls. That entire set of code, probably two to 300 lines of Python code.”
Zico Kolter May 7, 2026 ▶ 1:13:46
Insight
Kolter: Entire complexity of AI systems emerges from training data
“The entire complexity of an AI system evolves from the data they're trained on.”
Zico Kolter May 7, 2026 ▶ 1:14:36
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.