Zico Kolter discusses why AI safety and red teaming capabilities require explicit, specialized training rather than simple model scaling.
Assertion Open · timeframe Jun 2027
Kolter: Gray Swan's Shade system outperforms human red teamers at breaking models
“However, one thing that we are finding, and this is actually, I think we're kind of crossing this point too. Is that in a lot of the latest experiments, we can do much better than people, than human red teamers now at breaking these models. When I say we, I me…”
Insight
Kolter: AI is an alien intelligence with completely distinct failure modes from humans
“It is clearly a different form of intelligence than people. It's some alien intelligence that is vastly different, and that difference is actually often brought out to a large degree by things like adversarial attacks and red teaming, because there are certain…”
Insight
Kolter: Full experimental observability has not produced fundamental understanding of AI
“It's like we could kind of run experiments on the brain, observe every neuron in it, reset its state to prior states, and run counterfactuals, none of which we can do with humans, and yet we still understand neither very well. Even with that, all that ability,…”
Insight
Kolter: Adversarial Red Teaming Is Essential for True Capability Elicitation
“One of the most effective ways of doing capability elicitation is actually through some amount of what you would call red teaming, right? So if a model refuses a task because it thinks it's being evaluated, but it knows how to complete that task, getting it to…”
Prediction Not checkable as stated
Kolter: Security and science will explode as AI agents automate tedious verification
“So I think this is really sort of an underappreciated point that we're reaching this point, this sort of phase where a lot of security, a lot of science has this potential to kind of explode. Not because we're going to get better at it, but because agents can …”
Insight
Kolter: Concentrated Foundation Model Usage Creates Systemic Correlated Security Exploits
“And especially when there's the possibility of correlated failures, right? So it's not just that there's a lot of AI systems out there, it's that there's actually a few models that everyone is using. And if you find vulnerabilities in the agents that everyone …”