Gray Swan

0 statements across 0 episodes · 5 bullish · 3 bearish · 2 people on the record · first statement Jun 22, 2026 by Zico Kolter · said 21 times in 1 episodes since 2026 · across every show →

Mentions by year

brought up most by Zico Kolter (14), Matt Fredrikson (1)

tap a year for its mentions
001312512026episodesmentions
0112026episodes it came up in
00130.52512026episodesmentions per episode

every mention, scene by scene, with the transcript →

Everything said about Gray Swan, oldest first

Jun 22, 2026 neutral
Prediction Not checkable as stated
Kolter: AI agents inheriting user permissions by default will soon change
“So far, we are still a lot, in a lot of cases, operating on the condition that your agent has your permissions. Yeah. That is a very standard default. And I think that will be changed. I mean, your permissions may be in a sandbox, but still kind of your permis…”
Zico Kolter Jun 22, 2026 ▶ 53:00 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Jun 22, 2026 positive
Prediction Not checkable as stated
Kolter: Agent identity will evolve around user personas before fine-grained permissions
“I think in terms of how this will evolve, actually, I don't think it'll be per app, but I think what will happen first is people have different personas that they have, right? So you don't want your work life and your home email to be mixed up. Yeah. Right. A …”
Zico Kolter Jun 22, 2026 ▶ 54:19 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Jun 22, 2026 negative
Insight
Fredrikson: AI agent identity systems risk triggering automatic user consent fatigue
“One of the bigger challenges that people are going to face when they do start to roll out, like these agent identity sort of viewpoints and solutions is you run into that same kind of usability problem. Where like, what's the real recourse? Well, it stopped. I…”
Matt Fredrikson Jun 22, 2026 ▶ 53:48 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Jun 22, 2026 neutral
Assertion Not checkable as stated
Fredrikson: Unpublicized AI security breaches have already caused real damage
“We know that it has happened and it has caused real damage. That's the factor that's driven some people to us, right? Is they want protection from that.”
Matt Fredrikson Jun 22, 2026 ▶ 1:06:54 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Jun 22, 2026 bearish
Prediction Not checkable as stated
Kolter: A major AI security incident is inevitable and foreseeable
“The name gray swan is sort of in reference to black swan events, which are things no one could see coming. A gray swan is an unlikely event that you can kind of see coming. And that's kind of where we are with all of this, right? This is going to happen. We kn…”
Zico Kolter Jun 22, 2026 ▶ 1:06:30 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Jun 22, 2026 negative
Assertion Not checkable as stated
Fredrikson: Gray Swan Found Jailbreaks in Every OpenClaw User Trajectory Tested
“So we just have a bunch of trajectories of actual people using OpenClaw. And tons and tons of different scenarios and just threw shade at it and like found breaks for each and every one of them, right?”
Matt Fredrikson Jun 22, 2026 ▶ 47:36 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Jun 22, 2026 positive
Insight
Kolter: Adversarial Red Teaming Is Essential for True Capability Elicitation
“One of the most effective ways of doing capability elicitation is actually through some amount of what you would call red teaming, right? So if a model refuses a task because it thinks it's being evaluated, but it knows how to complete that task, getting it to…”
Zico Kolter Jun 22, 2026 ▶ 25:45 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Jun 22, 2026 neutral
Assertion Open · timeframe Jun 2026
Fredrikson: Skilled red teamers phish human participants 60% to 70% of the time
“But for a skilled, like, human red teamer, they could fish the human participants, like, with the 60 to 70% success.”
Matt Fredrikson Jun 22, 2026 ▶ 22:20 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Jun 22, 2026 positive
Disclosure
Fredrikson: Gray Swan's Discord red teaming community has 15,000 members
“It's a really great community. Like, 15,000 people come and hang out on the Discord server. Not all of them take part in every competition, but a lot of good data and good signal is provided to, you know, the upstream model developers through, through that com…”
Matt Fredrikson Jun 22, 2026 ▶ 9:30 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Jun 22, 2026 positive
Assertion Open · timeframe Jun 2026
Fredrikson: Top AI browser agents yielded only a handful of successful breaks
“There were a couple of models that seemed to be very, very robust, right? Like the red teamers found just a handful of successful breaks on them.”
Matt Fredrikson Jun 22, 2026 ▶ 22:29 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Jun 22, 2026 bullish
Assertion Open · timeframe Jun 2027
Kolter: Gray Swan's Shade system outperforms human red teamers at breaking models
“However, one thing that we are finding, and this is actually, I think we're kind of crossing this point too. Is that in a lot of the latest experiments, we can do much better than people, than human red teamers now at breaking these models. When I say we, I me…”
Zico Kolter Jun 22, 2026 ▶ 12:14 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Jun 22, 2026
Insight
Fredrikson: Enterprises refuse public red-teaming for pre-deployment AI agents
“Like enterprises are not willing to put up their pre-deployment agents on the arena for the general public to come hit. They're fine if it's, you know, 20 people that, that we've kind of handpicked from the arena.”
Matt Fredrikson Jun 22, 2026 ▶ 1:00:11 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.