why aren't all 19 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Opinion
Becker: Overly Bullish AI Developer Speedup Estimates Are Inflated
“I do think that very bullish estimates of speed up today are, you know, to some extent inflated by what we document in that original paper, that people's expectations of speed up tend to be too optimistic, it seems. They also tend to be inflated, I think, by n…”
Prediction Not checkable as stated
Becker: Operational Long Tail Will Delay Full AI R&D Automation
“There's this, Very long tail of things potentially involved in in R&D that would perhaps need to be fully automated in order to lead to capabilities explosion. I expect we're measuring, you know, in some ways only, only a small proportion of, only a small prop…”
Insight
Becker: Algorithmic Progress in AI Is Strictly a Function of Compute
“The suggestion in this paper is that if you think that algorithmic progress, you know, that, that is coming up with the transformer, coming up with RLHF, you know, MOEs, all of this stuff, better learning rate schedules is, is is itself a function of compute b…”
Prediction Not checkable as stated
Becker: Halving AI Compute Growth Halves Algorithmic Progress and Milestones
“And both of them both of those components half when compute halves sort of trivially, because compute is halving, and algorithmic progress halves because compute is this important input, and compute halves, then you might expect time horizon growth to half. An…”
Insight
Becker: AI Scaffolding Value Does Not Persist Across Model Generations
“Within model generation, it's valuable, and across model generations, it's not so valuable.”
Disclosure
Becker: Stopped Investing in Personal Software Engineering Skills Due to AI
“Intentionally not investing in engineering skills, because the areas are getting so good, maybe that's the wrong decision.”
Assertion Supported
Becker: Current Frontier Models Cannot Cause Catastrophic Harm
“We find, we think it's not capable enough, you know, on the basis of some of this capabilities evidence that you've alluded to commit these catastrophic harms.”
Insight
Becker: Time Horizon Metric Measures AI Task Difficulty in Human Time
“You know, instead we're just plotting what's the difficulty of tasks they can do over time, and that difficulty is measured in human time.”
Insight
Becker: AI Progress Remains Highly Continuous Across Compute Scales
“You know, in some ways, I think the story of Time Horizon is that progress has been remarkably continuous over, over so many years, so many orders of magnitude of compute and effective compute.”
Opinion
Becker: Prediction Markets' Social Value May Not Justify Gambling Harms
“I think gambling like behaviors are socially costly and the value of higher quality information is is, is, is real, but, you know, is it worth that disbenefit of people trading away their money, it's not, you know, it's not so clear to me.”
Disclosure
METR Shifts Focus from Autonomous Replication to AI R&D Acceleration
“So, something like the autonomous replication threat model, that is being able to set yourself up and control resources, something like that, has been deprioritized relative to AR and D acceleration. That is, you know, the possibility there could be some capab…”
Opinion
Becker: AI Models Lag on Time Horizons for Vision Tasks
“Tasks that are requiring of vision capabilities, they're probably to take one example, they're probably much less capable today as measured by time horizon, as for these tasks that are typically not requiring vision, vision capabilities that we give them.”
Insight
Becker: Low-Context Benchmark Criteria Exclude Real-World Situational Work
“Could a low-context human who was sufficiently skilled at sort of The general skills, but maybe, maybe not the particulars in the background would they be able to achieve success on, on this task? And I think that, that rules out a lot of real work because, yo…”
Disclosure
Becker: Donated $5,000 Won on Manifold Markets to Charity
“And then I ended up donating, I can't remember exactly how, how much it was, not, not so much, something like 5000 dollars .”
Opinion
Becker: AI Coding Fails at Merge Readiness Despite High SWE-bench Scores
“Maybe one that I'll call out there is this difference between whether models pass unit tests, whether they succeed by, you know, SWE bench-like scoring kind of meter-like scoring, benchmark-style scoring, versus whether their solution would be merged into main…”
Assertion Supported
Becker: METR Uses Black-Box Methods Over Interpretability for AI Monitoring
“Usually this is black box, not, not white box in, in, in my understanding in, in current work. So, so not using interpretability, but you can imagine in principle doing, doing, doing something more white box.”
Assertion Supported
Becker: METR's Hardest Benchmarks Require 20 to 30 Hours of Autonomy
“Then we go up to HCOS tasks, which span from, you know, only a little harder than those small tasks, all the way up to, you know, something like 20:30 hours, which are requiring of more autonomy, more sort of more sort of sequential actions.”
Disclosure
Becker: METR Is Rerunning Its Developer Productivity Randomized Controlled Trial
“We have been redoing it in the background.”
Assertion Partly supported
Becker: Claude Opus 4 Solves Atomic Software Tasks 100% Reliably
“Opus-IV. I'm sure can do that task a hundred percent of the time.”