Dec 31, 2025 · 17m · latent-space

[State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang

John Yang · 11m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this NeurIPS interview, SWE-bench co-creator John Yang reviews the current landscape of AI coding benchmarks, introduces the competitive long-horizon framework Code Clash, and discusses the future of autonomous agents versus interactive human-AI developer collaboration.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 6.2 Guest teaching 2.4 Guest disagreement 1.2 The hosts pushing back 3.4
05100:0010:001:07–3:08 · The hosts as informed peer 4/10 SWE-bench Variants and Benchmark Curation Methods The host inquires about SWE-bench spinoffs like SWE-bench Pro and asks about language diversification beyond Django. The guest clarifies that SWE-bench Pro was completely independent and explains their multilingual and multimodal expansion efforts.3:08–5:57 · The hosts as informed peer 7/10 Code Clash: Long-Horizon Tournaments and Game Arenas The guest introduces Code Clash and mentions Halite. The host demonstrates direct domain knowledge by correcting the guest on Halite being from Two Sigma rather than Jane Street, sharing his personal experience coding for it.5:57–9:11 · The hosts as informed peer 6/10 Surveying the SOTA Coding Benchmark Landscape The conversation covers several recent benchmark papers across performance and security. The host demonstrates subject expertise by synthesizing the role of PsyCode versus expensive agentic benchmarks and skeptically critiques single-path user simulators like Vending Bench.9:12–14:31 · The hosts as informed peer 7/10 Evaluating Impossible Tasks and Model Refusals The host explicitly pushes back against the guest's excitement for 24-hour unassisted autonomous execution, arguing that real developer workflows require fast, interactive human-in-the-loop iteration. The guest concedes and reframes long autonomy as suitable for repetitive background tasks.14:32–17:31 · The hosts as informed peer 7/10 Telemetry Needs and Codebase Understanding Frontiers The guest asks for user interaction data and outlines multi-agent evaluation setups. The host details Cognition's upcoming work on codebase understanding and automatic context engineering, also correcting the guest on Google's internal team structure.1:07–3:08 · Guest teaching 3/10 SWE-bench Variants and Benchmark Curation Methods The host inquires about SWE-bench spinoffs like SWE-bench Pro and asks about language diversification beyond Django. The guest clarifies that SWE-bench Pro was completely independent and explains their multilingual and multimodal expansion efforts.3:08–5:57 · Guest teaching 2/10 Code Clash: Long-Horizon Tournaments and Game Arenas The guest introduces Code Clash and mentions Halite. The host demonstrates direct domain knowledge by correcting the guest on Halite being from Two Sigma rather than Jane Street, sharing his personal experience coding for it.5:57–9:11 · Guest teaching 3/10 Surveying the SOTA Coding Benchmark Landscape The conversation covers several recent benchmark papers across performance and security. The host demonstrates subject expertise by synthesizing the role of PsyCode versus expensive agentic benchmarks and skeptically critiques single-path user simulators like Vending Bench.9:12–14:31 · Guest teaching 2/10 Evaluating Impossible Tasks and Model Refusals The host explicitly pushes back against the guest's excitement for 24-hour unassisted autonomous execution, arguing that real developer workflows require fast, interactive human-in-the-loop iteration. The guest concedes and reframes long autonomy as suitable for repetitive background tasks.14:32–17:31 · Guest teaching 2/10 Telemetry Needs and Codebase Understanding Frontiers The guest asks for user interaction data and outlines multi-agent evaluation setups. The host details Cognition's upcoming work on codebase understanding and automatic context engineering, also correcting the guest on Google's internal team structure.1:07–3:08 · Guest disagreement 1/10 SWE-bench Variants and Benchmark Curation Methods The host inquires about SWE-bench spinoffs like SWE-bench Pro and asks about language diversification beyond Django. The guest clarifies that SWE-bench Pro was completely independent and explains their multilingual and multimodal expansion efforts.3:08–5:57 · Guest disagreement 1/10 Code Clash: Long-Horizon Tournaments and Game Arenas The guest introduces Code Clash and mentions Halite. The host demonstrates direct domain knowledge by correcting the guest on Halite being from Two Sigma rather than Jane Street, sharing his personal experience coding for it.5:57–9:11 · Guest disagreement 1/10 Surveying the SOTA Coding Benchmark Landscape The conversation covers several recent benchmark papers across performance and security. The host demonstrates subject expertise by synthesizing the role of PsyCode versus expensive agentic benchmarks and skeptically critiques single-path user simulators like Vending Bench.9:12–14:31 · Guest disagreement 2/10 Evaluating Impossible Tasks and Model Refusals The host explicitly pushes back against the guest's excitement for 24-hour unassisted autonomous execution, arguing that real developer workflows require fast, interactive human-in-the-loop iteration. The guest concedes and reframes long autonomy as suitable for repetitive background tasks.14:32–17:31 · Guest disagreement 1/10 Telemetry Needs and Codebase Understanding Frontiers The guest asks for user interaction data and outlines multi-agent evaluation setups. The host details Cognition's upcoming work on codebase understanding and automatic context engineering, also correcting the guest on Google's internal team structure.1:07–3:08 · The hosts pushing back 2/10 SWE-bench Variants and Benchmark Curation Methods The host inquires about SWE-bench spinoffs like SWE-bench Pro and asks about language diversification beyond Django. The guest clarifies that SWE-bench Pro was completely independent and explains their multilingual and multimodal expansion efforts.3:08–5:57 · The hosts pushing back 2/10 Code Clash: Long-Horizon Tournaments and Game Arenas The guest introduces Code Clash and mentions Halite. The host demonstrates direct domain knowledge by correcting the guest on Halite being from Two Sigma rather than Jane Street, sharing his personal experience coding for it.5:57–9:11 · The hosts pushing back 3/10 Surveying the SOTA Coding Benchmark Landscape The conversation covers several recent benchmark papers across performance and security. The host demonstrates subject expertise by synthesizing the role of PsyCode versus expensive agentic benchmarks and skeptically critiques single-path user simulators like Vending Bench.9:12–14:31 · The hosts pushing back 7/10 Evaluating Impossible Tasks and Model Refusals The host explicitly pushes back against the guest's excitement for 24-hour unassisted autonomous execution, arguing that real developer workflows require fast, interactive human-in-the-loop iteration. The guest concedes and reframes long autonomy as suitable for repetitive background tasks.14:32–17:31 · The hosts pushing back 3/10 Telemetry Needs and Codebase Understanding Frontiers The guest asks for user interaction data and outlines multi-agent evaluation setups. The host details Cognition's upcoming work on codebase understanding and automatic context engineering, also correcting the guest on Google's internal team structure.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 13:38 Guest defending multi-level autonomy against pure interactivity

The guest articulates that while human collaboration is key, developers still want full hands-off autonomy for tedious data tasks like JSON parsing.

Hardest push from the hosts ▶ 12:33 Host challenging the push for long autonomous runs

The host explicitly declares a pushback, asserting that long-horizon multi-hour autonomy acts more as a benchmark stunt than a practical real-world developer tool.

Biggest teaching moment ▶ 3:20 Guest detailing structural flaws in unit-test-based evals

The guest explains why single-instance unit tests fail to capture consequential, long-horizon multi-turn software development.

The host holds their own ▶ 4:46 Host fact-checking Halite's sponsor and origins

The host immediately catches and corrects the guest's attribution of Halite, referencing his direct past employment at Two Sigma.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
SWE-bench Variants and Benchmark Curation Methods 4312 The host inquires about SWE-bench spinoffs like SWE-bench Pro and asks about language diversification beyond Django. The guest clarifies that SWE-bench Pro was completely independent and explains their multilingual and multimodal expansion efforts.
Code Clash: Long-Horizon Tournaments and Game Arenas 7212 The guest introduces Code Clash and mentions Halite. The host demonstrates direct domain knowledge by correcting the guest on Halite being from Two Sigma rather than Jane Street, sharing his personal experience coding for it.
Surveying the SOTA Coding Benchmark Landscape 6313 The conversation covers several recent benchmark papers across performance and security. The host demonstrates subject expertise by synthesizing the role of PsyCode versus expensive agentic benchmarks and skeptically critiques single-path user simulators like Vending Bench.
Evaluating Impossible Tasks and Model Refusals 7227 The host explicitly pushes back against the guest's excitement for 24-hour unassisted autonomous execution, arguing that real developer workflows require fast, interactive human-in-the-loop iteration. The guest concedes and reframes long autonomy as suitable for repetitive background tasks.
Telemetry Needs and Codebase Understanding Frontiers 7213 The guest asks for user interaction data and outlines multi-agent evaluation setups. The host details Cognition's upcoming work on codebase understanding and automatic context engineering, also correcting the guest on Google's internal team structure.

Statements from this episode (9)

Assertion Not checkable as stated
Yang: SWE-bench adoption took off only after Cognition's Devin launch
“You know, we put it out October, 22, 23, and then people didn't really touch it too much. And then, of course, like, Cognition came on the scene, and Devon was an amazing release, and I think after that, it kind of kicked off the arms.”
John Yang Dec 31, 2025 ▶ 0:37
Disclosure
Yang: Walden Yan emailed benchmark results two weeks before Devin launched
“I got an email about, like, two weeks ago. I think it was from, I think it was from Walden. It was like, hey, you know, we have a good number on it.”
John Yang Dec 31, 2025 ▶ 0:51
Assertion Supported
Yang: SWE-bench Multilingual spans nine languages across roughly 40 repositories
“Yeah, multilingual, it's like nine languages across, like, 40 repos, but yeah, you got them, like, JavaScript, Rust, Java, C, you know, Ruby.”
John Yang Dec 31, 2025 ▶ 1:57
Opinion
Yang: Unit tests and independent task instances limit SWE-bench's effectiveness
“I don't like unit tests as a form of verification, and I also think there's an issue with Sweepbench where all of the task instances are independent of each other.”
John Yang Dec 31, 2025 ▶ 3:27
Disclosure
Yang: Code Clash is building economically valuable competitive coding arenas
“The current ongoing effort is, you know, to build economically valuable arenas.”
John Yang Dec 31, 2025 ▶ 5:27
Insight
Yang: Models should pass completion benchmarks before expensive multi-turn agentic evals
“Like, you know, you can do well on those first, and then sort of graduate to the multi-turn expensive stuff.”
John Yang Dec 31, 2025 ▶ 7:14
Assertion Supported
Yang: AI models falsely claim completion on impossible coding tasks
“I think they're all, the models are all kind of attempting and saying, like, oh, I did it, you know, so maybe not great.”
John Yang Dec 31, 2025 ▶ 10:54
Insight
Yang: TerminalBench enables more creative eval design than SWE-bench
“As we bench, you're confined in some sense to the domain of issues and PRs that already exist which I think has its benefits of being close to reality and natural, but I think with Terminal Bench, there's a lot of creativity that you can infuse into that”
John Yang Dec 31, 2025 ▶ 11:14
Insight
Yang: Evaluating Human-AI Code Interaction Requires Compelling Products or Simulators
“From an academic standpoint, it feels like there's two difficult approaches to resolving that. Either you build, like, a really compelling product, like Elmarina, that people have people use consistently, which is, I mean, really tricky in and of itself. Or yo…”
John Yang Dec 31, 2025 ▶ 14:48
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.