Joel Becker of METR discusses how AI evaluation benchmarks work on Latent Space, clarifying a common industry misunderstanding regarding model execution duration versus task difficulty.
Opinion
Becker: Overly Bullish AI Developer Speedup Estimates Are Inflated
“I do think that very bullish estimates of speed up today are, you know, to some extent inflated by what we document in that original paper, that people's expectations of speed up tend to be too optimistic, it seems. They also tend to be inflated, I think, by n…”
Prediction Not checkable as stated
Becker: Operational Long Tail Will Delay Full AI R&D Automation
“There's this, Very long tail of things potentially involved in in R&D that would perhaps need to be fully automated in order to lead to capabilities explosion. I expect we're measuring, you know, in some ways only, only a small proportion of, only a small prop…”
Insight
Becker: Algorithmic Progress in AI Is Strictly a Function of Compute
“The suggestion in this paper is that if you think that algorithmic progress, you know, that, that is coming up with the transformer, coming up with RLHF, you know, MOEs, all of this stuff, better learning rate schedules is, is is itself a function of compute b…”
Prediction Not checkable as stated
Becker: Halving AI Compute Growth Halves Algorithmic Progress and Milestones
“And both of them both of those components half when compute halves sort of trivially, because compute is halving, and algorithmic progress halves because compute is this important input, and compute halves, then you might expect time horizon growth to half. An…”
Insight
Becker: AI Scaffolding Value Does Not Persist Across Model Generations
“Within model generation, it's valuable, and across model generations, it's not so valuable.”
Disclosure
Becker: Stopped Investing in Personal Software Engineering Skills Due to AI
“Intentionally not investing in engineering skills, because the areas are getting so good, maybe that's the wrong decision.”