SWE-bench Verified

part of SWE Bench

11 statements across 6 episodes · 4 bullish · 4 bearish · 7 people on the record · first statement Aug 22, 2024 by Alistair Pullen · said 53 times in 13 episodes since 2024 · across every show →

Mentions by year

brought up most by Shawn Wang (8), Alessio Fanelli (6), Mia Glaese (4), Olivia Watkins (3), Jesse Hu (3), John Yang (2), Graham Neubig (1)

tap a year for its mentions
00133256202420252026episodesmentions
036202420252026episodes it came up in
007.53156202420252026episodesmentions per episode
2026 21 mentions in 2 episodes 11 per episode
2025 12 mentions in 6 episodes 2 per episode
2024 20 mentions in 5 episodes 4 per episode

every mention, scene by scene, with the transcript →

Everything said about SWE-bench Verified, oldest first

Aug 22, 2024 positive
Assertion Supported
Cosine's Genie achieved a state-of-the-art 43.8% on SWE-bench Verified
“We got 219 out of 500, which is 43.8%, which is To my knowledge, at least right now, state of the art also”
Alistair Pullen Aug 22, 2024 ▶ 52:38 Is finetuning GPT4o worth it?
Oct 19, 2024 bullish
Prediction Open · timeframe Oct 2029
Hu: AI Models Should Eventually Reach 100% on SWE-bench Verified
“And in that way, I think we should be able to hit up a hundred percent eventually.”
Jesse Hu Oct 19, 2024 ▶ 20:49 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
Nov 28, 2024 neutral
Assertion Supported
Schluntz: SWE-bench Verified was created in partnership with OpenAI
“SweetBench Verified was actually made in partnership with OpenAI, and they hired humans to go review all these tasks and pick out a subset to try to remove any obstacle like this that would make the tasks impossible.”
Erik Schluntz Nov 28, 2024 ▶ 10:03 The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic
Apr 2, 2025 positive
Assertion Supported
Gur-Ari: Augment Code Achieved #1 on SWE-Bench Verified
“We just made number one on Sweepbench. So for us Sweepbench has been a useful tool for exploring how can we get the most out of agents. And so we were able to get the best result on Sweepbench verified right now.”
Guy Gur-Ari Apr 2, 2025 ▶ 1:24 The #1 SWE-Bench Verified Agent
Oct 18, 2025 positive
Prediction Not checkable as stated
Merrill: AI evals will shift to observing real jobs over 2-3 years
“I think like the future of evals does look much more like this is like observing people who are doing their real jobs and then translating those real jobs into a format that allows you to evaluate language models and harnesses on them. And it's probably going …”
Mike Merrill Oct 18, 2025 ▶ 18:41 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
Oct 18, 2025 negative
Assertion Contradicted
Merrill: 60% to 70% of SWE-bench Verified tasks come from Django
“If you go look at sweet bench verified, I think like 60, 70% of the tasks in there are from Django.”
Mike Merrill Oct 18, 2025 ▶ 19:06 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
Feb 23, 2026 negative
Opinion
Watkins: SWE-bench Verified is saturated, contaminated, and should be retired
“SweetBenchVerified has been one of the Northstar coding benchmarks that the field has looked at to measure coding progress. But recently we've seen that progress has kind of stalled and this, we realized that this is because the eval is effectively saturated a…”
Olivia Watkins Feb 23, 2026 ▶ 1:08 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
Feb 23, 2026 neutral
Assertion Supported
Watkins: About 90% of SWE-bench Verified tasks take under an hour
“For Sweep Edge Verified I think that's something like, 90% of the problems are things that were estimated to take, like, an expert software engineer like, less than an hour.”
Olivia Watkins Feb 23, 2026 ▶ 10:59 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
Feb 23, 2026 bearish
Assertion Not checkable as stated
Glaese: OpenAI no longer trusts further score improvements on SWE-bench Verified
“Issues with the benchmark that means that now that we're at like 80%, we don't really trust like further improvements on it, but like it does measure something that is like a real like capability of models.”
Mia Glaese Feb 23, 2026 ▶ 14:34 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
Feb 23, 2026
Assertion Supported
Watkins: OpenAI hired nearly 100 engineers to curate 500 SWE-bench tasks
“So folks at OpenAI did a pretty extensive human data campaign, hiring like almost a hundred real-world software engineers to go through the problems and figure out, like, are the tasks well-specified? Are the tests actually fair and kind of created a curated s…”
Olivia Watkins Feb 23, 2026 ▶ 2:56 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
Feb 23, 2026 negative
Assertion Not checkable as stated
Watkins: SWE-bench Verified is contaminated across OpenAI, Claude, and Gemini models
“And in SweetBenchVerified, we found many instances of contamination across like, across OpenEye models, across, like, Quad Opus, 4.5, Gemini Flash, and all of these, we saw things like regurgitating the ground truth solutions, things like in some cases giving,…”
Olivia Watkins Feb 23, 2026 ▶ 11:54 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.