Dec 31, 2025 · 17m · latent-space
[State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this NeurIPS interview, SWE-bench co-creator John Yang reviews the current landscape of AI coding benchmarks, introduces the competitive long-horizon framework Code Clash, and discusses the future of autonomous agents versus interactive human-AI developer collaboration.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
The guest articulates that while human collaboration is key, developers still want full hands-off autonomy for tedious data tasks like JSON parsing.
Hardest push from the hosts ▶ 12:33 Host challenging the push for long autonomous runsThe host explicitly declares a pushback, asserting that long-horizon multi-hour autonomy acts more as a benchmark stunt than a practical real-world developer tool.
Biggest teaching moment ▶ 3:20 Guest detailing structural flaws in unit-test-based evalsThe guest explains why single-instance unit tests fail to capture consequential, long-horizon multi-turn software development.
The host holds their own ▶ 4:46 Host fact-checking Halite's sponsor and originsThe host immediately catches and corrects the guest's attribution of Halite, referencing his direct past employment at Two Sigma.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| SWE-bench Variants and Benchmark Curation Methods | 4 | 3 | 1 | 2 | The host inquires about SWE-bench spinoffs like SWE-bench Pro and asks about language diversification beyond Django. The guest clarifies that SWE-bench Pro was completely independent and explains their multilingual and multimodal expansion efforts. | |
| Code Clash: Long-Horizon Tournaments and Game Arenas | 7 | 2 | 1 | 2 | The guest introduces Code Clash and mentions Halite. The host demonstrates direct domain knowledge by correcting the guest on Halite being from Two Sigma rather than Jane Street, sharing his personal experience coding for it. | |
| Surveying the SOTA Coding Benchmark Landscape | 6 | 3 | 1 | 3 | The conversation covers several recent benchmark papers across performance and security. The host demonstrates subject expertise by synthesizing the role of PsyCode versus expensive agentic benchmarks and skeptically critiques single-path user simulators like Vending Bench. | |
| Evaluating Impossible Tasks and Model Refusals | 7 | 2 | 2 | 7 | The host explicitly pushes back against the guest's excitement for 24-hour unassisted autonomous execution, arguing that real developer workflows require fast, interactive human-in-the-loop iteration. The guest concedes and reframes long autonomy as suitable for repetitive background tasks. | |
| Telemetry Needs and Codebase Understanding Frontiers | 7 | 2 | 1 | 3 | The guest asks for user interaction data and outlines multi-agent evaluation setups. The host details Cognition's upcoming work on codebase understanding and automatic context engineering, also correcting the guest on Google's internal team structure. |