SWE Bench Verified
product on 2 shows · 12 statements across 7 episodes · said 53 times in 13 episodes since 2024
Latent Space 53
the MAD Podcast
Mentions by year, every show
tap a year for its mentions
Latent Space 53
2026 21 mentions in 2 episodes 11 per episode
2025 12 mentions in 6 episodes 2 per episode
-
[State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang -
Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits -
The #1 SWE-Bench Verified Agent -
Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis -
[State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor -
2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents - every mention in 2025, scene by scene →
2024 20 mentions in 5 episodes 4 per episode
-
The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic -
[Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu -
Is finetuning GPT4o worth it? -
Building AGI in Real Time (OpenAI Dev Day 2024) -
Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands) - every mention in 2024, scene by scene →
every mention on every show, scene by scene, with the transcript →
12 statements about SWE Bench Verified, every show
Watkins: SWE-bench Verified is saturated, contaminated, and should be retired
“SweetBenchVerified has been one of the Northstar coding benchmarks that the field has looked at to measure coding progress. But recently we've seen that progress has kind of stalled and this, we realized that this is because the eval is effectively saturated a…”
Watkins: OpenAI hired nearly 100 engineers to curate 500 SWE-bench tasks
“So folks at OpenAI did a pretty extensive human data campaign, hiring like almost a hundred real-world software engineers to go through the problems and figure out, like, are the tasks well-specified? Are the tests actually fair and kind of created a curated s…”
Watkins: About 90% of SWE-bench Verified tasks take under an hour
“For Sweep Edge Verified I think that's something like, 90% of the problems are things that were estimated to take, like, an expert software engineer like, less than an hour.”
Watkins: SWE-bench Verified is contaminated across OpenAI, Claude, and Gemini models
“And in SweetBenchVerified, we found many instances of contamination across like, across OpenEye models, across, like, Quad Opus, 4.5, Gemini Flash, and all of these, we saw things like regurgitating the ground truth solutions, things like in some cases giving,…”
Glaese: OpenAI no longer trusts further score improvements on SWE-bench Verified
“Issues with the benchmark that means that now that we're at like 80%, we don't really trust like further improvements on it, but like it does measure something that is like a real like capability of models.”
Kant: Poolside's Malibu agent matches initial Gemini 2.5 Pro SWE-bench scores
“But if you think about the Malibu agent as a coding agent, for instance, right now, it sits at a level of like sweet bench verified, for instance, where Gemini two and a half pro was when it came out.”
Merrill: AI evals will shift to observing real jobs over 2-3 years
“I think like the future of evals does look much more like this is like observing people who are doing their real jobs and then translating those real jobs into a format that allows you to evaluate language models and harnesses on them. And it's probably going …”
Merrill: 60% to 70% of SWE-bench Verified tasks come from Django
“If you go look at sweet bench verified, I think like 60, 70% of the tasks in there are from Django.”
Gur-Ari: Augment Code Achieved #1 on SWE-Bench Verified
“We just made number one on Sweepbench. So for us Sweepbench has been a useful tool for exploring how can we get the most out of agents. And so we were able to get the best result on Sweepbench verified right now.”
Schluntz: SWE-bench Verified was created in partnership with OpenAI
“SweetBench Verified was actually made in partnership with OpenAI, and they hired humans to go review all these tasks and pick out a subset to try to remove any obstacle like this that would make the tasks impossible.”