SWE-bench co-creator John Yang discusses coding evaluation efficiency with the host, comparing cheap single-turn completion benchmarks with costly agentic evaluations.
Opinion
Yang: Unit tests and independent task instances limit SWE-bench's effectiveness
“I don't like unit tests as a form of verification, and I also think there's an issue with Sweepbench where all of the task instances are independent of each other.”
Assertion Supported
Yang: AI models falsely claim completion on impossible coding tasks
“I think they're all, the models are all kind of attempting and saying, like, oh, I did it, you know, so maybe not great.”
Insight
Yang: TerminalBench enables more creative eval design than SWE-bench
“As we bench, you're confined in some sense to the domain of issues and PRs that already exist which I think has its benefits of being close to reality and natural, but I think with Terminal Bench, there's a lot of creativity that you can infuse into that”
Insight
Yang: Evaluating Human-AI Code Interaction Requires Compelling Products or Simulators
“From an academic standpoint, it feels like there's two difficult approaches to resolving that. Either you build, like, a really compelling product, like Elmarina, that people have people use consistently, which is, I mean, really tricky in and of itself. Or yo…”
Assertion Not checkable as stated
Yang: SWE-bench adoption took off only after Cognition's Devin launch
“You know, we put it out October, 22, 23, and then people didn't really touch it too much. And then, of course, like, Cognition came on the scene, and Devon was an amazing release, and I think after that, it kind of kicked off the arms.”
Disclosure
Yang: Walden Yan emailed benchmark results two weeks before Devin launched
“I got an email about, like, two weeks ago. I think it was from, I think it was from Walden. It was like, hey, you know, we have a good number on it.”