Matt Fredrikson

CEO & Co-Founder, Gray Swan · 1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

founderexecutiveacademicscientistLinkedIn ↗mattfredrikson.com ↗

Matt Fredrikson conducts research in AI security, privacy, and formal verification, pioneering work on model inversion attacks and automated adversarial jailbreaks against large language models. He co-founded Gray Swan to build AI defensive guardrails, automated adversarial testing, and enterprise red-teaming infrastructure.

15statements → 7claims → 1claims resolved → 3.93/5average certainty → 2.2/5average debate potential →

1 supported 0 partly supported 0 contradicted 2 not yet assessed 4 not checkable as stated how the 7 claims stand · each chip opens the sources

7 assertions · 7 insights · 1 disclosure · every statement was checked. The predictions and assertions are the 7 claims: statements the public record can support or contradict. 1 is resolved, 2 are not yet assessed, and 4 name no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Matt argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
Fredrikson: AI capability does not correlate with prompt injection resistance
“So this scatter plot on the right, right, is essentially like looking for a correlation between capability and attack success rate. So on the X axis, how capable is the model at, you know, GPQA diamond on, on the Y axis. How, how often, you know, were people s…”
Matt Fredrikson Jun 22, 2026 ▶ 30:09 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan

Everything Matt Fredrikson said on Latent Space that made the record, most notable first. Filter by type, assessment or year in the ledger →

Assertion Not checkable as stated
Fredrikson: Frontier AI models fall for simulated prompt injections humans would ignore
“While in these scenarios, humans found it very difficult to prompt inject the models, like we're aware of scenarios that a human would never fall for, that like Opus four seven would, right? Like a, you know, an email that comes to your inbox and it says somet…”
Matt Fredrikson Jun 22, 2026 ▶ 22:55 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Insight
Fredrikson: Agent Guardrails Should Block Policy Violations, Not Injection Payloads
“If you parse some untrusted content and there is like a prompt injection, you know, something that's clearly trying to get the model to do a bad thing, you might be interested in knowing about that, but you don't necessarily like want your cloud code that you …”
Matt Fredrikson Jun 22, 2026 ▶ 41:05 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Assertion Not checkable as stated
Fredrikson: Gray Swan Found Jailbreaks in Every OpenClaw User Trajectory Tested
“So we just have a bunch of trajectories of actual people using OpenClaw. And tons and tons of different scenarios and just threw shade at it and like found breaks for each and every one of them, right?”
Matt Fredrikson Jun 22, 2026 ▶ 47:36 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Insight
Fredrikson: Evaluation-aware AI models often execute harmful actions because it is a simulation
“If you make, if you're testing the model for robustness or safety, right? And it's aware that it's being tested because you've set things up in a very artificial way, right? Like the email addresses are at example.com. The webpage is clearly not a real webpage…”
Matt Fredrikson Jun 22, 2026 ▶ 23:32 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Insight
Fredrikson: Prompt engineering cannot reliably enforce AI agent security policies
“Oftentimes people will try and prompt their way around it, like adjust the system prompt or like engineer the agent in a way where you're interjecting all the time and reminding it of what the original Goal and objective was, and that'll get you a little bit o…”
Matt Fredrikson Jun 22, 2026 ▶ 31:42 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Insight
Fredrikson: AI agent identity systems risk triggering automatic user consent fatigue
“One of the bigger challenges that people are going to face when they do start to roll out, like these agent identity sort of viewpoints and solutions is you run into that same kind of usability problem. Where like, what's the real recourse? Well, it stopped. I…”
Matt Fredrikson Jun 22, 2026 ▶ 53:48 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Assertion Not checkable as stated
Fredrikson: Unpublicized AI security breaches have already caused real damage
“We know that it has happened and it has caused real damage. That's the factor that's driven some people to us, right? Is they want protection from that.”
Matt Fredrikson Jun 22, 2026 ▶ 1:06:54 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Assertion Open · timeframe Jun 2026
Fredrikson: Skilled red teamers phish human participants 60% to 70% of the time
“But for a skilled, like, human red teamer, they could fish the human participants, like, with the 60 to 70% success.”
Matt Fredrikson Jun 22, 2026 ▶ 22:20 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Assertion Open · timeframe Jun 2026
Fredrikson: Top AI browser agents yielded only a handful of successful breaks
“There were a couple of models that seemed to be very, very robust, right? Like the red teamers found just a handful of successful breaks on them.”
Matt Fredrikson Jun 22, 2026 ▶ 22:29 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Insight
Fredrikson: Red Teaming Is Fundamentally a Mathematical Optimization Problem
“I mean, it really is an optimization problem, right? You have a, you know, an outcome that you want the model to exhibit, right? Now, how do I find the input, right? That, that gives me that output and you can sort of objectify that actually very mathematicall…”
Matt Fredrikson Jun 22, 2026 ▶ 26:38 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Assertion Supported
Fredrikson: AI capability does not correlate with prompt injection resistance
“So this scatter plot on the right, right, is essentially like looking for a correlation between capability and attack success rate. So on the X axis, how capable is the model at, you know, GPQA diamond on, on the Y axis. How, how often, you know, were people s…”
Matt Fredrikson Jun 22, 2026 ▶ 30:09 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Assertion Not checkable as stated
Fredrikson: Amazon excels at deploying formal software verification, Microsoft in research
“Microsoft historically has been pretty good about it too. More on the research side, Amazon is, is stellar and actually deploying a lot of this.”
Matt Fredrikson Jun 22, 2026 ▶ 42:52 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Insight
Fredrikson: Formal software verification takes 10 to 20 times longer than Python
“The reason people don't do it is that it's not easy and it's not fun, right? It takes you like 10 or 20 times as long to like fight with the type checker, which is essentially like proving that you don't have a vulnerability as if, as it would if you just like…”
Matt Fredrikson Jun 22, 2026 ▶ 43:10 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Insight
Fredrikson: Enterprises refuse public red-teaming for pre-deployment AI agents
“Like enterprises are not willing to put up their pre-deployment agents on the arena for the general public to come hit. They're fine if it's, you know, 20 people that, that we've kind of handpicked from the arena.”
Matt Fredrikson Jun 22, 2026 ▶ 1:00:11 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Disclosure
Fredrikson: Gray Swan's Discord red teaming community has 15,000 members
“It's a really great community. Like, 15,000 people come and hang out on the Discord server. Not all of them take part in every competition, but a lot of good data and good signal is provided to, you know, the upstream model developers through, through that com…”
Matt Fredrikson Jun 22, 2026 ▶ 9:30 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan

Appearances (1)

EpisodeDateSpeaking time
AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan Jun 22, 2026 21m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.