SWE-agent, every mention

13 scenes · ← back to SWE-agent

tap a year for its mentions
008315620242025episodesmentions
03620242025episodes it came up in
002.535620242025episodesmentions per episode

every year anyone Shawn Wang 14John Yang 3Charles Packer 3Alessio Fanelli 2Will Brown 1

Verbatim, from the transcripts: the passages where SWE-agent comes up

loading…

[State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang Dec 31, 2025 · 3 mentions

  • ▶ 7:35 John Yang They like the x-axis being sort of the runtime, and their, or, yeah, y-axis being the completion, you know, like, we can do more long-running sweet agent tasks.
  • ▶ 11:59 John Yang And then, of course, for myself, I think just like this long-running sweet agent kind of thing just feels very compelling.
  • ▶ 15:30 John Yang Like, I think for, for Code Clash, what I'm excited about is the current framing is really long-running suite agents, but, you know, you could have multi-agents, like, two agents work together on the code base, and what happens?

🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R) Jul 28, 2025 · 1 mention

  • ▶ 1:52:59 Alessio Fanelli Yeah, we did an episode with Anthropic about their recent, like, sweet agent, sweet bench results, and we talked about the human eval versus sweet bench, and early human eval is kind of like a Greenfield benchmark, you know, you need to be…

⚡️Multi-Turn RL for Multi-Hour Agents — with Will Brown, Prime Intellect May 23, 2025 · 1 mention

  • ▶ 2:07 Will Brown They're really like showing off their suite agent and like tool, tool use and like function calling benchmarks, multi turn stuff.

Sleep-Time Compute — Letta AI (Charles Packer, Charlie Snell, Kevin Lin) Apr 21, 2025 · 3 mentions

How Claude Plays Pokémon was made Mar 4, 2025 · 1 mention

Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis Jan 28, 2025 · 1 mention

  • ▶ 4:58 unnamed speaker So we did an episode with Anthropic about their SWE agent work and like the SWE bench verified results that they had.

Windsurf: The Enterprise AI IDE Dec 13, 2024 · 1 mention

  • ▶ 15:49 unnamed speaker Yeah, we did an episode with Anthropic about their recent, like, sweet agent, sweet bench, uh, results, and we talked about the human eval versus sweet bench, and, or, like, human eval is kind of like a Greenfield benchmark, you know, you…

The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic Nov 28, 2024 · 6 mentions

  • ▶ 17:17 Shawn Wang I think we'll go into sweet agent in a, in a little bit, but I kind of reject the fact that you need to choose one prompt and like have your whole performance be predicated on one prompt.
  • ▶ 41:14 Shawn Wang I also wanted to spend a little bit of time with Sweet Agent. 4 times in the scene
  • ▶ 49:19 Shawn Wang There's a run with SuiAgent plus Opus, but that's the official SuiBench guys doing it.

Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph Sep 27, 2024 · 8 mentions

  • ▶ 36:06 Shawn Wang I would consider Devin to be the sort of biggest launch of the year as far as AI startups go, and, uh, you guys in the Princeton group worked on SWE agents alongside of SWE bench. 8 times in the scene
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.