why aren't all 14 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Prediction Not checkable as stated
Becker: Operational Long Tail Will Delay Full AI R&D Automation
“There's this, Very long tail of things potentially involved in in R&D that would perhaps need to be fully automated in order to lead to capabilities explosion. I expect we're measuring, you know, in some ways only, only a small proportion of, only a small prop…”
Insight
Becker: Time Horizon Metric Measures AI Task Difficulty in Human Time
“You know, instead we're just plotting what's the difficulty of tasks they can do over time, and that difficulty is measured in human time.”
Assertion Supported
Becker: Current Frontier Models Cannot Cause Catastrophic Harm
“We find, we think it's not capable enough, you know, on the basis of some of this capabilities evidence that you've alluded to commit these catastrophic harms.”
Insight
Becker: AI Progress Remains Highly Continuous Across Compute Scales
“You know, in some ways, I think the story of Time Horizon is that progress has been remarkably continuous over, over so many years, so many orders of magnitude of compute and effective compute.”
Assertion Not checkable as stated
AI engineering speedups require specific use cases and strong digital hygiene
“AI sped me up a bit, but only in specific cases and only when taking a lot of sort of digital hygiene practices.”
Assertion Partly supported
METR benchmark: AI agent autonomy duration doubles every 3 to 7 months
“They established a Moore's law for time between human input, basically, and it's basically doubling every three to seven months is the idea. And Enthopic is currently doing super well on that benchmark. It's roughly about autonomous for 15 minutes at the 50th …”
Insight
Becker: Low-Context Benchmark Criteria Exclude Real-World Situational Work
“Could a low-context human who was sufficiently skilled at sort of The general skills, but maybe, maybe not the particulars in the background would they be able to achieve success on, on this task? And I think that, that rules out a lot of real work because, yo…”
Opinion
Becker: AI Models Lag on Time Horizons for Vision Tasks
“Tasks that are requiring of vision capabilities, they're probably to take one example, they're probably much less capable today as measured by time horizon, as for these tasks that are typically not requiring vision, vision capabilities that we give them.”
Opinion
Becker: AI Coding Fails at Merge Readiness Despite High SWE-bench Scores
“Maybe one that I'll call out there is this difference between whether models pass unit tests, whether they succeed by, you know, SWE bench-like scoring kind of meter-like scoring, benchmark-style scoring, versus whether their solution would be merged into main…”
Disclosure
METR Shifts Focus from Autonomous Replication to AI R&D Acceleration
“So, something like the autonomous replication threat model, that is being able to set yourself up and control resources, something like that, has been deprioritized relative to AR and D acceleration. That is, you know, the possibility there could be some capab…”
Assertion Supported
Becker: METR Uses Black-Box Methods Over Interpretability for AI Monitoring
“Usually this is black box, not, not white box in, in, in my understanding in, in current work. So, so not using interpretability, but you can imagine in principle doing, doing, doing something more white box.”
Assertion Supported
Becker: METR's Hardest Benchmarks Require 20 to 30 Hours of Autonomy
“Then we go up to HCOS tasks, which span from, you know, only a little harder than those small tasks, all the way up to, you know, something like 20:30 hours, which are requiring of more autonomy, more sort of more sort of sequential actions.”
Disclosure
Becker: METR Is Rerunning Its Developer Productivity Randomized Controlled Trial
“We have been redoing it in the background.”
Assertion Partly supported
Becker: Claude Opus 4 Solves Atomic Software Tasks 100% Reliably
“Opus-IV. I'm sure can do that task a hundred percent of the time.”