why aren't all 14 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Insight
Kolter: AI models do not get safer automatically by scaling up
“You can't just sort of trust models to get safer by getting bigger. You have to put in the work to actually make them safer. And this is, I think what a lot of AI companies are investing in. This is why we in fact do have models that are improving on these dim…”
Insight
The vast majority of modern AI intelligence comes from self-training
“I don't think people have properly internalized the fact that the vast majority of intelligence comes from self training effectively.”
Insight
Kolter: AI safety is solved through frontier interaction, not pauses
“I think the way you solve things is through, through ongoing exploration of what's happening and through, through interaction with the frontier.”
Insight
Kolter: Reasoning models are much harder to jailbreak via probability optimization
“Reasoning models were much more effective because you can't really do the same trick of optimizing for a probability with a reasoning model that has a whole trace of reasoning that happens in the middle and kind of reflect a bit more. So it's much harder to br…”
Insight
Kolter: AI coding agents are extremely good mechanistic interpretability researchers
“Coding agents are extremely good Mechinterp researchers.”
Insight
Kolter: Scientific progress occurs when young researchers ignore old guard beliefs
“Basically progress happens when The current crop of young researchers ignores the things they've been taught that the old guard believes.”
Insight
Kolter: AI safety focus is shifting from single models to ecosystems
“I think actually one of the big trends we're seeing is a lot of safety is moving from the model level to the ecosystem level and talking about, you know, what's not one model capable of, but what's AI broadly capable of.”
Insight
Kolter: Renaming AI Safety Summit to AI Action Summit shows political shift
“The, you know, AI safety conference was, or AI safety summit was renamed the AI action summit or something is, has some significance actually in terms of the sort of taking temperature of where the world is politically.”
Insight
Kolter: AI evals test average performance; security tests worst-case performance
“Most evaluations are done kind of in a, They measure expected value, basically. They measure sort of how well does it work on average, and security measures how well does it work in the worst case.”
Insight
Kolter: Jailbreaking modern AI safety systems requires complex multi-query attacks
“But they are, they require that degree of complexity to really jailbreak modern systems in a, for information that has this sensitivity to it.”
Insight
Kolter: Prompt injection introduces data exfiltration risks to AI agents
“Things like prompt injection are really a new security vulnerability for AI agents, and they mean that your risk is not just that you could have some, the model says something mean to you or something like that. Or even they could just write bad code. It could…”
Insight
Kolter: AI agent security requires managing permissions alongside manipulation risks
“AI security of agents is this interaction between what can the agent be manipulated into doing? What might it do accidentally? And what credentials or access does it have to really affect change?”
Insight
Kolter: Major AI breakthroughs require both massive scale and luck
“Reasoning models were the next big breakthrough. Those are rare. They do take kind of a, you know, both, both a massive scale and kind of a bit of luck to get there.”
Insight
Kolter: Entire complexity of AI systems emerges from training data
“The entire complexity of an AI system evolves from the data they're trained on.”