why aren't all 5 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Supported
Becker: Current Frontier Models Cannot Cause Catastrophic Harm
“We find, we think it's not capable enough, you know, on the basis of some of this capabilities evidence that you've alluded to commit these catastrophic harms.”
Assertion Partly supported
METR benchmark: AI agent autonomy duration doubles every 3 to 7 months
“They established a Moore's law for time between human input, basically, and it's basically doubling every three to seven months is the idea. And Enthopic is currently doing super well on that benchmark. It's roughly about autonomous for 15 minutes at the 50th …”
Assertion Supported
Becker: METR Uses Black-Box Methods Over Interpretability for AI Monitoring
“Usually this is black box, not, not white box in, in, in my understanding in, in current work. So, so not using interpretability, but you can imagine in principle doing, doing, doing something more white box.”
Assertion Supported
Becker: METR's Hardest Benchmarks Require 20 to 30 Hours of Autonomy
“Then we go up to HCOS tasks, which span from, you know, only a little harder than those small tasks, all the way up to, you know, something like 20:30 hours, which are requiring of more autonomy, more sort of more sort of sequential actions.”
Assertion Partly supported
Becker: Claude Opus 4 Solves Atomic Software Tasks 100% Reliably
“Opus-IV. I'm sure can do that task a hundred percent of the time.”