SWE-bench
includes SWE Bench Verified, SWE Bench Pro, SWE Bench Multimodal, SWE Bench Multilingual, SWE Bench Lite, SWE Bench Full, SWE Bench Light, SWE Bench Live
51 statements across 20 episodes · 16 bullish · 18 bearish · 18 people on the record · first statement Aug 22, 2024 by Alistair Pullen · said 299 times in 44 episodes since 2024 · across every show →
Mentions by year, the whole family
brought up most by Shawn Wang (60), John Yang (14), Jesse Hu (14), Graham Neubig (14), Erik Schluntz (14), Shawn Lewis (13), Alistair Pullen (12), Alessio Fanelli (12)
2026 38 mentions in 7 episodes 5 per episode
- The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
- 🔬 "The Most Innovative Diffusion Research Is Happening in Drug Discovery, Not Image Generation"
- Measuring Exponential Trends Rising (in AI) — Joel Becker, METR
- Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen
- When AI Agents Run Businesses — Lukas Petersson and Axel Backlund of Andon Labs
- Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
- Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
- every mention in 2026, scene by scene →
2025 128 mentions in 23 episodes 6 per episode
- [State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang
- Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
- The #1 SWE-Bench Verified Agent
- Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
- 2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
- 🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
- ChatGPT Codex: The Missing Manual
- GPT 4.1: The New OpenAI Workhorse
- 15 more episodes that year, every mention in 2025 →
2024 133 mentions in 14 episodes 10 per episode
- The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic
- [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
- Is finetuning GPT4o worth it?
- Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
- Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
- Building AGI in Real Time (OpenAI Dev Day 2024)
- Windsurf: The Enterprise AI IDE
- Building AGI with OpenAI's Structured Outputs API
- 6 more episodes that year, every mention in 2024 →