SWE-bench, every mention

126 scenes, the whole family · ← back to SWE-bench

tap a year for its mentions
00751315025202420252026episodesmentions
01325202420252026episodes it came up in
005131025202420252026episodesmentions per episode

every year anyone Shawn Wang 60John Yang 14Jesse Hu 14Graham Neubig 14Erik Schluntz 14Shawn Lewis 13Alistair Pullen 12Alessio Fanelli 12Anshul Ramachandran 11Guy Gur-Ari 10

Verbatim, from the transcripts: the passages where SWE-bench comes up

loading…

🔬 "The Most Innovative Diffusion Research Is Happening in Drug Discovery, Not Image Generation" Jun 30, 2026 · 2 mentions

Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen Jun 25, 2026 · 1 mention

  • ▶ 19:29 unnamed speaker I guess one aspect or area that seems very interesting are evals, um, and more specifically, have there been instances where you've seen, like, through just vibe checks that it's really good, but on the actual benchmarks it, like, performs…

When AI Agents Run Businesses — Lukas Petersson and Axel Backlund of Andon Labs Jun 4, 2026 · 1 mention

  • ▶ 7:34 unnamed speaker Like, uh, C bench verified, um, even vending bench one saturated, right?

Measuring Exponential Trends Rising (in AI) — Joel Becker, METR Feb 27, 2026 · 2 mentions

  • ▶ 52:26 Joel Becker Maybe one that I'll call out there is this difference between whether models pass, uh, unit tests, whether they, they succeed by, you know, SWE bench-like scoring, um, kind of meter-like scoring, benchmark-style scoring, versus whether… 2 times in the scene

Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis Feb 24, 2026 · 1 mention

  • ▶ 1:17:21 Doug O'Laughlin I think the difference is codex wants to code because it's RL to be so good at coding to win on sweet bench that like you're trying to use it for general information.

The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals Feb 23, 2026 · 30 mentions

  • ▶ 0:37 unnamed speaker And, uh, as by my understanding, you were part of the original team that worked on C-Bench Verified as well. 4 times in the scene
  • ▶ 1:33 unnamed speaker Like SweetBenchPro. 2 times in the scene
  • ▶ 1:52 unnamed speaker I think the, uh, let's, let's sort of reset on, like, what was the original work that you guys did for Sequence Verified, which I think was pretty substantial. 4 times in the scene
  • ▶ 2:14 Olivia Watkins SweetBench Verified was kind of a cleanup of original bench, academic benchmark from a lab at Princeton called SweetBench, and the agent is basically given a code base and a task that was sourced from a real-world repository and GitHub… 2 times in the scene
  • ▶ 8:57 Mia Glaese I think, I think also like at the time when three bench verified was published, I think it was like a very strong benchmark. 2 times in the scene
  • ▶ 10:46 unnamed speaker We're going to stop reporting CBench Verified, right? 5 times in the scene
  • ▶ 10:48 unnamed speaker And then, uh, CBench Pro will, will be some of the next one, which is an effort from scale. 3 times in the scene
  • ▶ 14:13 Mia Glaese They think three bench, three bench verified, obviously measured like some, that measures like some important capability, which is like, given like a description of a GitHub issue, can you produce like a patch that solves that issue, you… 2 times in the scene
  • ▶ 23:27 Olivia Watkins And that's kind of what ties most into the Sweebench, where coding is not all of automating research, but it is one very important key component.
  • ▶ 23:35 Olivia Watkins And so, uh, we initially created Sweebench Verified as part of our, like, building out evals for that model autonomy workstream. 3 times in the scene
  • ▶ 24:24 Mia Glaese Like Sweebench Pro, we're like, yes, but that's a better eval now. 2 times in the scene

Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith Jan 9, 2026 · 1 mention

  • ▶ 33:42 George Cameron And when the people that, that created this, like Minhui and, and, and actually Ophia, who was kind of behind Sweebench.

[State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang Dec 31, 2025 · 29 mentions

  • ▶ 0:12 unnamed speaker We're here at NeurIPS with John Yang of SweetBench and many other things, but welcome. 4 times in the scene
  • ▶ 1:07 unnamed speaker And then SweetBench Verified was, like, maybe last year.
  • ▶ 1:15 unnamed speaker You've, there's, like, a whole bunch of varieties of SweetBench now. 4 times in the scene
  • ▶ 1:23 John Yang One is, like, more SweetBenches, SweetBench Pro, SweetBench Live.
  • ▶ 1:23 John Yang One is, like, more SweetBenches, SweetBench Pro, SweetBench Live. 3 times in the scene
  • ▶ 1:47 unnamed speaker but yeah, uh, multimodal. 3 times in the scene
  • ▶ 1:49 John Yang Yeah, we did multimodal and multilingual, um, and I think, like, those have, multilingual seems to be, uh, is it, like, JavaScript? 4 times in the scene
  • ▶ 3:27 John Yang I don't like unit tests as a form of verification, and I also think there's an issue with Sweepbench where all of the task instances are independent of each other. 2 times in the scene
  • ▶ 7:06 unnamed speaker Sweetbench is expensive to run. 3 times in the scene
  • ▶ 7:43 John Yang I think the projections are, are quite interesting, and I definitely appreciate them kind of using SweetBench Verified to, to sort of proxy a lot of these things, but
  • ▶ 10:40 John Yang I don't know, but they basically took Sweet Bench Verified, and they changed the issues to make them impossible.
  • ▶ 11:06 John Yang I mean, honestly, I think, I think it's, people will make more suite benches. 2 times in the scene

[State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor Dec 30, 2025 · 1 mention

  • ▶ 8:42 unnamed speaker So Subay's verified, for sure.

⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust — Brian Fioca + Bill Chen, OpenAI Dec 26, 2025 · 1 mention

⚡️ 10x AI Engineers with $1m Salaries — Alex Lieberman & Arman Hezarkhani, Tenex Nov 19, 2025 · 1 mention

  • ▶ 17:59 Shawn Wang Like we think the models are good, but like actually they have been really trained into a certain sort of local minima of like, well, here's all the Python because Sweebench is all Python, all Django.

Anthropic, Glean & OpenRouter: How AI Moats Are Built with Deedy Das of Menlo Ventures Nov 14, 2025 · 2 mentions

  • ▶ 29:59 Shawn Wang Obviously this is benchmarks and evals and everyone has like, okay, today it's your turn to be best at SweetBench. 2 times in the scene

Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders Nov 8, 2025 · 3 mentions

  • ▶ 13:54 Alex Shaw Uh, it has a data set registry with many popular benchmarks, like SweetBench Verified, pre-integrated, uh, integrations with popular prompt optimization, and RL frameworks like SkyRail, and then out of the box cloud deployments using…
  • ▶ 23:29 Andy Konwinski Was really into sweep inch and mem GPT and we, and then Devon happened and I thought, wow, that's an interesting demo. 2 times in the scene

Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits Oct 18, 2025 · 10 mentions

  • ▶ 1:32 Alex Shaw And yeah, so he, he invited me to come work on the K prize, which was a one million dollar prize around sweet bench. 3 times in the scene
  • ▶ 16:08 unnamed speaker And then for example, you have sweet bench verified, but you don't have a sweet Lancer. 3 times in the scene
  • ▶ 16:12 unnamed speaker You don't have sweet bench pro like are all of those things. 2 times in the scene
  • ▶ 17:30 Mike Merrill So what made Sweebench so powerful was that you could just go on GitHub and find all of these repositories and 2 times in the scene

⚡️Claude Sonnet 4.5 and Anthropic's roadmap for Agents and Developers — Mike Krieger, Anthropic Sep 30, 2025 · 3 mentions

  • ▶ 17:01 Mike Krieger Maybe two weeks before the final snapshot of the final snapshot where yes, it improved on, on sweet bench, but even more so it went from, ah, it's mostly reliable to, yeah, it's great. 3 times in the scene

Long Live Context Engineering - with Jeff Huber of Chroma Aug 19, 2025 · 1 mention

  • ▶ 16:24 Jeff Huber So we're specifically looking at a couple of different datasets, uh, suite bench inclusive, and.

🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R) Jul 28, 2025 · 8 mentions

⚡️Warp 2.0: the Agentic Development Environment - Zach Lloyd and Ben Holmes Jun 25, 2025 · 3 mentions

  • ▶ 1:37 Zach Lloyd We should have some, I don't have the sweet bench score yet. 3 times in the scene

The AI Coding Factory May 29, 2025 · 3 mentions

  • ▶ 25:24 Shawn Wang Because let's say we talked about SweetBench before recording. 3 times in the scene

⚡️Multi-Turn RL for Multi-Hour Agents — with Will Brown, Prime Intellect May 23, 2025 · 1 mention

  • ▶ 6:58 Will Brown Like, you could imagine, like, a Sweebench kind of thing where there's a minimal diff that is really what you want, but there's, like, you could do a ton of other stuff and put all these other things in place that as long as you don't, as…

ChatGPT Codex: The Missing Manual May 16, 2025 · 4 mentions

  • ▶ 10:31 Alexander Embiricos You know, we don't just want it to be good at code, and, like, we don't just want it to, like, solve, like, say, like, Sweebench tasks. 2 times in the scene
  • ▶ 38:08 unnamed speaker Is this part of the, ah, you had cutoffs for a few, like, 23, um, SweetBench verified examples that were not runnable, was that part of it, ah, in terms of length, or was there just something else? 2 times in the scene

Why Every Agent needs Open Source Cloud Sandboxes Apr 24, 2025 · 1 mention

  • ▶ 57:03 Shawn Wang Like, it should be very easy to run, like, SuiteBench and all these on, on YouTube.

GPT 4.1: The New OpenAI Workhorse Apr 15, 2025 · 4 mentions

  • ▶ 22:12 unnamed speaker How much, and I, and I think I read, uh, that improves like the SWE, the SWE bench, like, 20% just by having like the persistence. 2 times in the scene
  • ▶ 29:10 Shawn Wang It's better than O-one and Sweetbench. 2 times in the scene

The Creators of Model Context Protocol Apr 3, 2025 · 1 mention

  • ▶ 57:11 unnamed speaker Who built your sort of suite bench projects on the podcast as well.

The #1 SWE-Bench Verified Agent Apr 2, 2025 · 12 mentions

  • ▶ 1:14 Alessio Fanelli And well, maybe you don't want to say to be humble, but this is going to be the number one sweep bench verified. 2 times in the scene
  • ▶ 1:24 Guy Gur-Ari Yeah, we just, we just made number one on Sweepbench. 6 times in the scene
  • ▶ 9:46 Guy Gur-Ari I don't think it's 2 times in the scene
  • ▶ 31:13 Guy Gur-Ari Maybe another thing is with, since we talked about Sweebench, we actually open sourced our implementation of Sweebench. 2 times in the scene

Fullstack-Bench: The Eval for Coding Agents — with Sujay Jayakar, Chief Scientist, Convex Mar 19, 2025 · 2 mentions

  • ▶ 14:00 Shawn Wang Um, the, the, how do you compare it versus the other benchmarks that you see out there, like the suite bench. 2 times in the scene

Solve coding, solve AGI [Reflection.ai launch w/ CEO Misha Laskin] Mar 7, 2025 · 3 mentions

Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis Jan 28, 2025 · 24 mentions

  • ▶ 0:13 unnamed speaker We have a new Sweebench King, and he, part of his, part of his, uh, appeal is also that he has a great name. 2 times in the scene
  • ▶ 4:58 unnamed speaker So we did an episode with Anthropic about their SWE agent work and like the SWE bench verified results that they had.
  • ▶ 7:46 Shawn Lewis So SweetBench is, you know, the, the best eval that we have for AI programming today. 6 times in the scene
  • ▶ 13:46 Shawn Lewis Um, and I think it was the first, you know, published result on SweetBench that, and, and maybe the only one, you know, still as of a week later, 2 times in the scene
  • ▶ 17:31 Shawn Lewis So this is like a simple table where along the rows we have each of the Sweebench problems in the dataset that we're looking at, and we can load up other columns from the three evals that we've selected on the side. 4 times in the scene
  • ▶ 25:19 Shawn Lewis I think, you know, I, I built this in the course of like really trying to solve that sweet bench or do really well on sweet bench first and foremost.
  • ▶ 26:03 unnamed speaker It's like, you know, there's kind of this, we bench verified benchmark, and then there's like, you know, the more programmatic, how am I going to use this agent? 2 times in the scene
  • ▶ 27:37 unnamed speaker Uh, and then also there's like just the general submission process for SweeBedge Verify. 6 times in the scene

The Unreasonable Effectiveness of Reasoning Distillation: using DeepSeek R1 to beat OpenAI o1 Jan 24, 2025 · 1 mention

  • ▶ 19:35 Shawn Wang And then we also, uh, on this podcast, we interviewed Cosign that actually fine tune four O to, on three bench, sorry, uh, to, to, to, to achieve data on three bench.

2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents Jan 1, 2025 · 10 mentions

Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands) Dec 25, 2024 · 18 mentions

  • ▶ 2:23 Graham Neubig This is a data science task, which says, uh, I want to create scatter plots that show the increase of the SWE bench score over time. 6 times in the scene
  • ▶ 23:06 Graham Neubig Um, and for code, we use Sweebench, which, um, I think a lot of people may have heard of.
  • ▶ 25:21 Graham Neubig So, right now, we have WebArena and Sweebench. 2 times in the scene
  • ▶ 26:33 Graham Neubig Um, right now we have 53% or 55% on sweet bench verified, which is real world GitHub PRS.
  • ▶ 31:40 unnamed speaker Uh, the first thing is that you said that you're estimating that your, um, your agent is successfully resolving, like, something like 30 to 40% of your issues, but that's, like, below what you saw on Sweebench, so I guess I'm wondering… 4 times in the scene
  • ▶ 37:30 Graham Neubig Oh, how would I make a successor to Sweebench? 4 times in the scene

The State of AI Startups in 2024 [LS Live @ NeurIPS] Dec 21, 2024 · 1 mention

  • ▶ 8:43 Sarah Guo If you recall, like a year ago, the point of view on sweet bench was like, it was impossible to surpass.

Windsurf: The Enterprise AI IDE Dec 13, 2024 · 5 mentions

The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic Nov 28, 2024 · 45 mentions

  • ▶ 0:21 Erik Schluntz I'm a member of technical staff at Anthropic working on tool use, computer use, and SweetBench. 3 times in the scene
  • ▶ 3:39 Shawn Wang So let's get right into SweetBench. 8 times in the scene
  • ▶ 4:14 Shawn Wang So yeah, maybe just give us a context about like why you looked at SweetBench Verified 3 times in the scene
  • ▶ 9:14 Alessio Fanelli How do we get Sweepbench verified to 92%? 4 times in the scene
  • ▶ 9:24 Erik Schluntz And actually, uh, maybe I'll start with SweetBench versus SweetBenchVerified, which is, I think, something I missed earlier. 3 times in the scene
  • ▶ 21:16 Shawn Wang You didn't need them? 6 times in the scene
  • ▶ 24:04 Erik Schluntz I will say, um, Sweetbench just released Sweetbench multimodal, um, which I believe is either entirely JavaScript or largely JavaScript.
  • ▶ 28:43 Erik Schluntz And again, like the tools we released as part of Sweetbench were, I'd say they're very specific for like editing files and doing bash, but at the same time, that's actually very general. 2 times in the scene
  • ▶ 29:22 Alessio Fanelli and then you're running it against Sweebench anyway, so it doesn't really need to write the test or? 6 times in the scene
  • ▶ 40:20 Shawn Wang You said you worked on the function calling and tool use before you actually started this three bench work, right? 3 times in the scene
  • ▶ 48:22 Erik Schluntz You know, we're not going to go and do lots more submissions to sweet bench and try to try to prompt engineer this and build a bigger system. 5 times in the scene
  • ▶ 58:35 Shawn Wang In some way that's, you know, solving sweet bench, like you, you, you, you should be allowed to use the internet or you should be allowed to use a computer to, to solve it and use your vision and use whatever.

Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI Nov 25, 2024 · 2 mentions

  • ▶ 35:47 Shawn Wang I would say if you, for example, it's interesting that like, for example, Sweebench, if you want to be considered for ranking, you have to submit your reasoning traces, and that has actually disqualified some of our past guests, like… 2 times in the scene

How NotebookLM Was Made Oct 25, 2024 · 1 mention

  • ▶ 32:59 Usama Shafqat Um, I think another weird thing here was like, we need it to be entertaining, and that's much harder to quantify than some of the other benchmarks that you can make for like, you know, Sweebench or like, get better at this math prompt.

[Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu Oct 19, 2024 · 13 mentions

  • ▶ 0:09 unnamed speaker Uh, I'm only seeing your SWE bench screen.
  • ▶ 0:27 Jesse Hu So, um, I'm sure a lot of folks have, uh, either seen or tried to, uh, get through SweetBench or even compete on SweetBench. 2 times in the scene
  • ▶ 8:40 unnamed speaker Yeah, uh, they see for the, uh, SWE Bench Verified, like, like, um, as you also, SWE Bench is very different, um, compared to the human evolve. 2 times in the scene
  • ▶ 8:40 unnamed speaker Yeah, uh, they see for the, uh, SWE Bench Verified, like, like, um, as you also, SWE Bench is very different, um, compared to the human evolve. 3 times in the scene
  • ▶ 13:10 Jesse Hu Well, at least on light, they were scoring.
  • ▶ 15:29 Jesse Hu And as a result, um, the sweep bench team, you know, like, uh,
  • ▶ 17:06 Jesse Hu Well, so now I'll move on to verified. 3 times in the scene
page 1 of 2 · 100 scenes per page · newest episode first next →
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.