SWE-bench Verified
part of SWE Bench
11 statements across 6 episodes · 4 bullish · 4 bearish · 7 people on the record · first statement Aug 22, 2024 by Alistair Pullen · said 53 times in 13 episodes since 2024 · across every show →
Mentions by year
brought up most by Shawn Wang (8), Alessio Fanelli (6), Mia Glaese (4), Olivia Watkins (3), Jesse Hu (3), John Yang (2), Graham Neubig (1)
2026 21 mentions in 2 episodes 11 per episode
2025 12 mentions in 6 episodes 2 per episode
- [State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang
- Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
- The #1 SWE-Bench Verified Agent
- Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
- [State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
- 2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
- every mention in 2025, scene by scene →
2024 20 mentions in 5 episodes 4 per episode
- The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic
- [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
- Is finetuning GPT4o worth it?
- Building AGI in Real Time (OpenAI Dev Day 2024)
- Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
- every mention in 2024, scene by scene →