Becker: AI Coding Fails at Merge Readiness Despite High SWE-bench Scores
Joel Becker · Measuring Exponential Trends Rising (in AI) — Joel Becker, METR · Feb 27, 2026 · at 52:26
Joel Becker of METR discusses limitations in current coding evaluations with hosts Alessio Fanelli and Swyx on the Latent Space podcast.
“Maybe one that I'll call out there is this difference between whether models pass unit tests, whether they succeed by, you know, SWE bench-like scoring kind of meter-like scoring, benchmark-style scoring, versus whether their solution would be merged into main. That is, you know, whether the solution adds tests where it should or doesn't. Whether it follows existing patterns in the code base, whether it makes sure that, that its changes sort of speak to other parts of the code base in, in, in, in appropriate ways. That, that seems very interesting to me. You know, I think model capabilities probably are lagging behind there somewhat versus versus that, which you might see on sweet bench like scoring.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →