SWE-bench, every mention
126 scenes, the whole family · ← back to SWE-bench
tap a year for its mentions
every year anyone Shawn Wang 60John Yang 14Jesse Hu 14Graham Neubig 14Erik Schluntz 14Shawn Lewis 13Alistair Pullen 12Alessio Fanelli 12Anshul Ramachandran 11Guy Gur-Ari 10
Verbatim, from the transcripts: the passages where SWE-bench comes up
🔬 "The Most Innovative Diffusion Research Is Happening in Drug Discovery, Not Image Generation"
- ▶ 55:29 Evan Feinberg Have you ever heard of Sweebench? 2 times in the scene
Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen
- ▶ 19:29 unnamed speaker I guess one aspect or area that seems very interesting are evals, um, and more specifically, have there been instances where you've seen, like, through just vibe checks that it's really good, but on the actual benchmarks it, like, performs…
When AI Agents Run Businesses — Lukas Petersson and Axel Backlund of Andon Labs
- ▶ 7:34 unnamed speaker Like, uh, C bench verified, um, even vending bench one saturated, right?
Measuring Exponential Trends Rising (in AI) — Joel Becker, METR
- ▶ 52:26 Joel Becker Maybe one that I'll call out there is this difference between whether models pass, uh, unit tests, whether they, they succeed by, you know, SWE bench-like scoring, um, kind of meter-like scoring, benchmark-style scoring, versus whether… 2 times in the scene
Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
- ▶ 1:17:21 Doug O'Laughlin I think the difference is codex wants to code because it's RL to be so good at coding to win on sweet bench that like you're trying to use it for general information.
The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
- ▶ 0:37 unnamed speaker And, uh, as by my understanding, you were part of the original team that worked on C-Bench Verified as well. 4 times in the scene
- ▶ 1:33 unnamed speaker Like SweetBenchPro. 2 times in the scene
- ▶ 1:52 unnamed speaker I think the, uh, let's, let's sort of reset on, like, what was the original work that you guys did for Sequence Verified, which I think was pretty substantial. 4 times in the scene
- ▶ 2:14 Olivia Watkins SweetBench Verified was kind of a cleanup of original bench, academic benchmark from a lab at Princeton called SweetBench, and the agent is basically given a code base and a task that was sourced from a real-world repository and GitHub… 2 times in the scene
- ▶ 8:57 Mia Glaese I think, I think also like at the time when three bench verified was published, I think it was like a very strong benchmark. 2 times in the scene
- ▶ 10:46 unnamed speaker We're going to stop reporting CBench Verified, right? 5 times in the scene
- ▶ 10:48 unnamed speaker And then, uh, CBench Pro will, will be some of the next one, which is an effort from scale. 3 times in the scene
- ▶ 14:13 Mia Glaese They think three bench, three bench verified, obviously measured like some, that measures like some important capability, which is like, given like a description of a GitHub issue, can you produce like a patch that solves that issue, you… 2 times in the scene
- ▶ 23:27 Olivia Watkins And that's kind of what ties most into the Sweebench, where coding is not all of automating research, but it is one very important key component.
- ▶ 23:35 Olivia Watkins And so, uh, we initially created Sweebench Verified as part of our, like, building out evals for that model autonomy workstream. 3 times in the scene
- ▶ 24:24 Mia Glaese Like Sweebench Pro, we're like, yes, but that's a better eval now. 2 times in the scene
Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
- ▶ 33:42 George Cameron And when the people that, that created this, like Minhui and, and, and actually Ophia, who was kind of behind Sweebench.
[State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang
- ▶ 0:12 unnamed speaker We're here at NeurIPS with John Yang of SweetBench and many other things, but welcome. 4 times in the scene
- ▶ 1:07 unnamed speaker And then SweetBench Verified was, like, maybe last year.
- ▶ 1:15 unnamed speaker You've, there's, like, a whole bunch of varieties of SweetBench now. 4 times in the scene
- ▶ 1:23 John Yang One is, like, more SweetBenches, SweetBench Pro, SweetBench Live.
- ▶ 1:23 John Yang One is, like, more SweetBenches, SweetBench Pro, SweetBench Live. 3 times in the scene
- ▶ 1:47 unnamed speaker but yeah, uh, multimodal. 3 times in the scene
- ▶ 1:49 John Yang Yeah, we did multimodal and multilingual, um, and I think, like, those have, multilingual seems to be, uh, is it, like, JavaScript? 4 times in the scene
- ▶ 3:27 John Yang I don't like unit tests as a form of verification, and I also think there's an issue with Sweepbench where all of the task instances are independent of each other. 2 times in the scene
- ▶ 7:06 unnamed speaker Sweetbench is expensive to run. 3 times in the scene
- ▶ 7:43 John Yang I think the projections are, are quite interesting, and I definitely appreciate them kind of using SweetBench Verified to, to sort of proxy a lot of these things, but
- ▶ 10:40 John Yang I don't know, but they basically took Sweet Bench Verified, and they changed the issues to make them impossible.
- ▶ 11:06 John Yang I mean, honestly, I think, I think it's, people will make more suite benches. 2 times in the scene
[State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
- ▶ 8:42 unnamed speaker So Subay's verified, for sure.
⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust — Brian Fioca + Bill Chen, OpenAI
⚡️ 10x AI Engineers with $1m Salaries — Alex Lieberman & Arman Hezarkhani, Tenex
- ▶ 17:59 Shawn Wang Like we think the models are good, but like actually they have been really trained into a certain sort of local minima of like, well, here's all the Python because Sweebench is all Python, all Django.
Anthropic, Glean & OpenRouter: How AI Moats Are Built with Deedy Das of Menlo Ventures
- ▶ 29:59 Shawn Wang Obviously this is benchmarks and evals and everyone has like, okay, today it's your turn to be best at SweetBench. 2 times in the scene
Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders
- ▶ 13:54 Alex Shaw Uh, it has a data set registry with many popular benchmarks, like SweetBench Verified, pre-integrated, uh, integrations with popular prompt optimization, and RL frameworks like SkyRail, and then out of the box cloud deployments using…
- ▶ 23:29 Andy Konwinski Was really into sweep inch and mem GPT and we, and then Devon happened and I thought, wow, that's an interesting demo. 2 times in the scene
Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
- ▶ 1:32 Alex Shaw And yeah, so he, he invited me to come work on the K prize, which was a one million dollar prize around sweet bench. 3 times in the scene
- ▶ 16:08 unnamed speaker And then for example, you have sweet bench verified, but you don't have a sweet Lancer. 3 times in the scene
- ▶ 16:12 unnamed speaker You don't have sweet bench pro like are all of those things. 2 times in the scene
- ▶ 17:30 Mike Merrill So what made Sweebench so powerful was that you could just go on GitHub and find all of these repositories and 2 times in the scene
⚡️Claude Sonnet 4.5 and Anthropic's roadmap for Agents and Developers — Mike Krieger, Anthropic
- ▶ 17:01 Mike Krieger Maybe two weeks before the final snapshot of the final snapshot where yes, it improved on, on sweet bench, but even more so it went from, ah, it's mostly reliable to, yeah, it's great. 3 times in the scene
Long Live Context Engineering - with Jeff Huber of Chroma
- ▶ 16:24 Jeff Huber So we're specifically looking at a couple of different datasets, uh, suite bench inclusive, and.
🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
- ▶ 1:50:34 Anshul Ramachandran Like, okay, you have SweeBench, that's cool, no actual 6 times in the scene
- ▶ 3:29:03 Kevin Hou And while the industry focuses heavily on things like Sweebench, 2 times in the scene
⚡️Warp 2.0: the Agentic Development Environment - Zach Lloyd and Ben Holmes
- ▶ 1:37 Zach Lloyd We should have some, I don't have the sweet bench score yet. 3 times in the scene
The AI Coding Factory
- ▶ 25:24 Shawn Wang Because let's say we talked about SweetBench before recording. 3 times in the scene
⚡️Multi-Turn RL for Multi-Hour Agents — with Will Brown, Prime Intellect
- ▶ 6:58 Will Brown Like, you could imagine, like, a Sweebench kind of thing where there's a minimal diff that is really what you want, but there's, like, you could do a ton of other stuff and put all these other things in place that as long as you don't, as…
ChatGPT Codex: The Missing Manual
- ▶ 10:31 Alexander Embiricos You know, we don't just want it to be good at code, and, like, we don't just want it to, like, solve, like, say, like, Sweebench tasks. 2 times in the scene
- ▶ 38:08 unnamed speaker Is this part of the, ah, you had cutoffs for a few, like, 23, um, SweetBench verified examples that were not runnable, was that part of it, ah, in terms of length, or was there just something else? 2 times in the scene
Why Every Agent needs Open Source Cloud Sandboxes
- ▶ 57:03 Shawn Wang Like, it should be very easy to run, like, SuiteBench and all these on, on YouTube.
GPT 4.1: The New OpenAI Workhorse
- ▶ 22:12 unnamed speaker How much, and I, and I think I read, uh, that improves like the SWE, the SWE bench, like, 20% just by having like the persistence. 2 times in the scene
- ▶ 29:10 Shawn Wang It's better than O-one and Sweetbench. 2 times in the scene
The Creators of Model Context Protocol
- ▶ 57:11 unnamed speaker Who built your sort of suite bench projects on the podcast as well.
The #1 SWE-Bench Verified Agent
- ▶ 1:14 Alessio Fanelli And well, maybe you don't want to say to be humble, but this is going to be the number one sweep bench verified. 2 times in the scene
- ▶ 1:24 Guy Gur-Ari Yeah, we just, we just made number one on Sweepbench. 6 times in the scene
- ▶ 9:46 Guy Gur-Ari I don't think it's 2 times in the scene
- ▶ 31:13 Guy Gur-Ari Maybe another thing is with, since we talked about Sweebench, we actually open sourced our implementation of Sweebench. 2 times in the scene
Fullstack-Bench: The Eval for Coding Agents — with Sujay Jayakar, Chief Scientist, Convex
- ▶ 14:00 Shawn Wang Um, the, the, how do you compare it versus the other benchmarks that you see out there, like the suite bench. 2 times in the scene
Solve coding, solve AGI [Reflection.ai launch w/ CEO Misha Laskin]
- ▶ 23:00 Shawn Wang Obviously, right now, people focus on SweetBench. 3 times in the scene
Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
- ▶ 0:13 unnamed speaker We have a new Sweebench King, and he, part of his, part of his, uh, appeal is also that he has a great name. 2 times in the scene
- ▶ 4:58 unnamed speaker So we did an episode with Anthropic about their SWE agent work and like the SWE bench verified results that they had.
- ▶ 7:46 Shawn Lewis So SweetBench is, you know, the, the best eval that we have for AI programming today. 6 times in the scene
- ▶ 13:46 Shawn Lewis Um, and I think it was the first, you know, published result on SweetBench that, and, and maybe the only one, you know, still as of a week later, 2 times in the scene
- ▶ 17:31 Shawn Lewis So this is like a simple table where along the rows we have each of the Sweebench problems in the dataset that we're looking at, and we can load up other columns from the three evals that we've selected on the side. 4 times in the scene
- ▶ 25:19 Shawn Lewis I think, you know, I, I built this in the course of like really trying to solve that sweet bench or do really well on sweet bench first and foremost.
- ▶ 26:03 unnamed speaker It's like, you know, there's kind of this, we bench verified benchmark, and then there's like, you know, the more programmatic, how am I going to use this agent? 2 times in the scene
- ▶ 27:37 unnamed speaker Uh, and then also there's like just the general submission process for SweeBedge Verify. 6 times in the scene
The Unreasonable Effectiveness of Reasoning Distillation: using DeepSeek R1 to beat OpenAI o1
- ▶ 19:35 Shawn Wang And then we also, uh, on this podcast, we interviewed Cosign that actually fine tune four O to, on three bench, sorry, uh, to, to, to, to achieve data on three bench.
2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
- ▶ 19:43 Shawn Wang It's very hard to stay on top of SweetBench. 2 times in the scene
- ▶ 1:06:34 Shawn Wang But also we have also seen like the emergence of Sweebench, Livebench, MMORPRO and AIME, AIME specifically for one. 6 times in the scene
- ▶ 1:07:40 Shawn Wang We care about Sweebench verified.
- ▶ 1:07:42 Shawn Wang Uh, we, we care about the Sweebench multimodal.
Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
- ▶ 2:23 Graham Neubig This is a data science task, which says, uh, I want to create scatter plots that show the increase of the SWE bench score over time. 6 times in the scene
- ▶ 23:06 Graham Neubig Um, and for code, we use Sweebench, which, um, I think a lot of people may have heard of.
- ▶ 25:21 Graham Neubig So, right now, we have WebArena and Sweebench. 2 times in the scene
- ▶ 26:33 Graham Neubig Um, right now we have 53% or 55% on sweet bench verified, which is real world GitHub PRS.
- ▶ 31:40 unnamed speaker Uh, the first thing is that you said that you're estimating that your, um, your agent is successfully resolving, like, something like 30 to 40% of your issues, but that's, like, below what you saw on Sweebench, so I guess I'm wondering… 4 times in the scene
- ▶ 37:30 Graham Neubig Oh, how would I make a successor to Sweebench? 4 times in the scene
The State of AI Startups in 2024 [LS Live @ NeurIPS]
Windsurf: The Enterprise AI IDE
- ▶ 13:24 Anshul Ramachandran Like, okay, you have sweet bench. 5 times in the scene
The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic
- ▶ 0:21 Erik Schluntz I'm a member of technical staff at Anthropic working on tool use, computer use, and SweetBench. 3 times in the scene
- ▶ 3:39 Shawn Wang So let's get right into SweetBench. 8 times in the scene
- ▶ 4:14 Shawn Wang So yeah, maybe just give us a context about like why you looked at SweetBench Verified 3 times in the scene
- ▶ 9:14 Alessio Fanelli How do we get Sweepbench verified to 92%? 4 times in the scene
- ▶ 9:24 Erik Schluntz And actually, uh, maybe I'll start with SweetBench versus SweetBenchVerified, which is, I think, something I missed earlier. 3 times in the scene
- ▶ 21:16 Shawn Wang You didn't need them? 6 times in the scene
- ▶ 24:04 Erik Schluntz I will say, um, Sweetbench just released Sweetbench multimodal, um, which I believe is either entirely JavaScript or largely JavaScript.
- ▶ 28:43 Erik Schluntz And again, like the tools we released as part of Sweetbench were, I'd say they're very specific for like editing files and doing bash, but at the same time, that's actually very general. 2 times in the scene
- ▶ 29:22 Alessio Fanelli and then you're running it against Sweebench anyway, so it doesn't really need to write the test or? 6 times in the scene
- ▶ 40:20 Shawn Wang You said you worked on the function calling and tool use before you actually started this three bench work, right? 3 times in the scene
- ▶ 48:22 Erik Schluntz You know, we're not going to go and do lots more submissions to sweet bench and try to try to prompt engineer this and build a bigger system. 5 times in the scene
- ▶ 58:35 Shawn Wang In some way that's, you know, solving sweet bench, like you, you, you, you should be allowed to use the internet or you should be allowed to use a computer to, to solve it and use your vision and use whatever.
Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI
- ▶ 35:47 Shawn Wang I would say if you, for example, it's interesting that like, for example, Sweebench, if you want to be considered for ranking, you have to submit your reasoning traces, and that has actually disqualified some of our past guests, like… 2 times in the scene
How NotebookLM Was Made
- ▶ 32:59 Usama Shafqat Um, I think another weird thing here was like, we need it to be entertaining, and that's much harder to quantify than some of the other benchmarks that you can make for like, you know, Sweebench or like, get better at this math prompt.
[Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
- ▶ 0:09 unnamed speaker Uh, I'm only seeing your SWE bench screen.
- ▶ 0:27 Jesse Hu So, um, I'm sure a lot of folks have, uh, either seen or tried to, uh, get through SweetBench or even compete on SweetBench. 2 times in the scene
- ▶ 8:40 unnamed speaker Yeah, uh, they see for the, uh, SWE Bench Verified, like, like, um, as you also, SWE Bench is very different, um, compared to the human evolve. 2 times in the scene
- ▶ 8:40 unnamed speaker Yeah, uh, they see for the, uh, SWE Bench Verified, like, like, um, as you also, SWE Bench is very different, um, compared to the human evolve. 3 times in the scene
- ▶ 13:10 Jesse Hu Well, at least on light, they were scoring.
- ▶ 15:29 Jesse Hu And as a result, um, the sweep bench team, you know, like, uh,
- ▶ 17:06 Jesse Hu Well, so now I'll move on to verified. 3 times in the scene
page 1 of 2 · 100 scenes per page · newest episode first next →