John Yang, co-creator of SWE-bench, compares SWE-bench's PR-based methodology with TerminalBench's interactive terminal environment.
Opinion
Yang: Unit tests and independent task instances limit SWE-bench's effectiveness
“I don't like unit tests as a form of verification, and I also think there's an issue with Sweepbench where all of the task instances are independent of each other.”
Insight
Yang: Models should pass completion benchmarks before expensive multi-turn agentic evals
“Like, you know, you can do well on those first, and then sort of graduate to the multi-turn expensive stuff.”
Assertion Supported
Yang: AI models falsely claim completion on impossible coding tasks
“I think they're all, the models are all kind of attempting and saying, like, oh, I did it, you know, so maybe not great.”
Insight
Yang: Evaluating Human-AI Code Interaction Requires Compelling Products or Simulators
“From an academic standpoint, it feels like there's two difficult approaches to resolving that. Either you build, like, a really compelling product, like Elmarina, that people have people use consistently, which is, I mean, really tricky in and of itself. Or yo…”
Assertion Not checkable as stated
Yang: SWE-bench adoption took off only after Cognition's Devin launch
“You know, we put it out October, 22, 23, and then people didn't really touch it too much. And then, of course, like, Cognition came on the scene, and Devon was an amazing release, and I think after that, it kind of kicked off the arms.”
Disclosure
Yang: Walden Yan emailed benchmark results two weeks before Devin launched
“I got an email about, like, two weeks ago. I think it was from, I think it was from Walden. It was like, hey, you know, we have a good number on it.”