The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 16 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 1 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Insight
Kolter: Scaling model size does not automatically improve AI safety or red teaming
“Traditionally this has been an area where both in terms of safety models don't get better by just being bigger, unlike most other areas where models do get better by being bigger. Safety has not been like that traditionally. You know, you have to train them ex…”
Zico Kolter Jun 22, 2026 ▶ 11:28 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Assertion Open · timeframe Jun 2027
Kolter: Gray Swan's Shade system outperforms human red teamers at breaking models
“However, one thing that we are finding, and this is actually, I think we're kind of crossing this point too. Is that in a lot of the latest experiments, we can do much better than people, than human red teamers now at breaking these models. When I say we, I me…”
Zico Kolter Jun 22, 2026 ▶ 12:14 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Insight
Kolter: AI is an alien intelligence with completely distinct failure modes from humans
“It is clearly a different form of intelligence than people. It's some alien intelligence that is vastly different, and that difference is actually often brought out to a large degree by things like adversarial attacks and red teaming, because there are certain…”
Zico Kolter Jun 22, 2026 ▶ 15:32 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Insight
Kolter: Full experimental observability has not produced fundamental understanding of AI
“It's like we could kind of run experiments on the brain, observe every neuron in it, reset its state to prior states, and run counterfactuals, none of which we can do with humans, and yet we still understand neither very well. Even with that, all that ability,…”
Zico Kolter Jun 22, 2026 ▶ 16:13 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Insight
Kolter: Adversarial Red Teaming Is Essential for True Capability Elicitation
“One of the most effective ways of doing capability elicitation is actually through some amount of what you would call red teaming, right? So if a model refuses a task because it thinks it's being evaluated, but it knows how to complete that task, getting it to…”
Zico Kolter Jun 22, 2026 ▶ 25:45 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Prediction Not checkable as stated
Kolter: Security and science will explode as AI agents automate tedious verification
“So I think this is really sort of an underappreciated point that we're reaching this point, this sort of phase where a lot of security, a lot of science has this potential to kind of explode. Not because we're going to get better at it, but because agents can …”
Zico Kolter Jun 22, 2026 ▶ 46:01 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Insight
Kolter: Concentrated Foundation Model Usage Creates Systemic Correlated Security Exploits
“And especially when there's the possibility of correlated failures, right? So it's not just that there's a lot of AI systems out there, it's that there's actually a few models that everyone is using. And if you find vulnerabilities in the agents that everyone …”
Zico Kolter Jun 22, 2026 ▶ 4:31 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Insight
Kolter: Frontier models fail at red teaming due to safety refusals
“So generally speaking, the issue with this is that frontier models are extremely bad at automated red teaming because they have a lot of safeguards built into them. So if you try to use them to jailbreak another model, they will actually refuse their safety tr…”
Zico Kolter Jun 22, 2026 ▶ 11:06 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Opinion
Kolter: Mechanistic interpretability is not yet a real science
“The problem with McInterp is it's a lot, it's been about sort of testing small hypotheses. Hypothesis. And you know, you have a hypothesis, you'll find some small thing, you'll test that in isolation. But I don't think it's really become a science yet.”
Zico Kolter Jun 22, 2026 ▶ 17:29 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Opinion
Kolter: SOC 2 is not a great security compliance model
“So, so I think SOC II is not a great model. We'll just say, but it is a model.”
Zico Kolter Jun 22, 2026 ▶ 1:04:25 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Prediction Not checkable as stated
Kolter: A major AI security incident is inevitable and foreseeable
“The name gray swan is sort of in reference to black swan events, which are things no one could see coming. A gray swan is an unlikely event that you can kind of see coming. And that's kind of where we are with all of this, right? This is going to happen. We kn…”
Zico Kolter Jun 22, 2026 ▶ 1:06:30 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Prediction Not checkable as stated
Kolter: Coding agents will revitalize mechanistic interpretability research
“Most fascinating things about coding agents actually is they can do a lot of experimentation in an automated fashion. Yeah. They will give new hope. They'll breathe new life into mechanter research.”
Zico Kolter Jun 22, 2026 ▶ 17:58 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Prediction Not checkable as stated
Kolter: AI systems will probably not achieve provably zero vulnerabilities soon
“So the question is not trying to completely Kind of provably mitigate these things. That is arguably just a, it's a good goal, but just like zero bug software, we're probably not going to get there. At least not that soon.”
Zico Kolter Jun 22, 2026 ▶ 37:13 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Assertion Supported
Kolter: Most AI Agents Resist Naive API Key Exfiltration Prompts
“Now, things that are that simple, to be clear, are covered at this point by most agents, right? You know, they all They, despite some issues, yeah, normal, normal sort of, you know, will not be that easily fooled by just push all my API keys to a public thing,…”
Zico Kolter Jun 22, 2026 ▶ 40:23 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Prediction Not checkable as stated
Kolter: AI agents inheriting user permissions by default will soon change
“So far, we are still a lot, in a lot of cases, operating on the condition that your agent has your permissions. Yeah. That is a very standard default. And I think that will be changed. I mean, your permissions may be in a sandbox, but still kind of your permis…”
Zico Kolter Jun 22, 2026 ▶ 53:00 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Prediction Not checkable as stated
Kolter: Agent identity will evolve around user personas before fine-grained permissions
“I think in terms of how this will evolve, actually, I don't think it'll be per app, but I think what will happen first is people have different personas that they have, right? So you don't want your work life and your home email to be mixed up. Yeah. Right. A …”
Zico Kolter Jun 22, 2026 ▶ 54:19 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.