SWE-agent, every mention
13 scenes · ← back to SWE-agent
tap a year for its mentions
every year anyone Shawn Wang 14John Yang 3Charles Packer 3Alessio Fanelli 2Will Brown 1
Verbatim, from the transcripts: the passages where SWE-agent comes up
[State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang
- ▶ 7:35 John Yang They like the x-axis being sort of the runtime, and their, or, yeah, y-axis being the completion, you know, like, we can do more long-running sweet agent tasks.
- ▶ 11:59 John Yang And then, of course, for myself, I think just like this long-running sweet agent kind of thing just feels very compelling.
- ▶ 15:30 John Yang Like, I think for, for Code Clash, what I'm excited about is the current framing is really long-running suite agents, but, you know, you could have multi-agents, like, two agents work together on the code base, and what happens?
🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
- ▶ 1:52:59 Alessio Fanelli Yeah, we did an episode with Anthropic about their recent, like, sweet agent, sweet bench results, and we talked about the human eval versus sweet bench, and early human eval is kind of like a Greenfield benchmark, you know, you need to be…
⚡️Multi-Turn RL for Multi-Hour Agents — with Will Brown, Prime Intellect
- ▶ 2:07 Will Brown They're really like showing off their suite agent and like tool, tool use and like function calling benchmarks, multi turn stuff.
Sleep-Time Compute — Letta AI (Charles Packer, Charlie Snell, Kevin Lin)
- ▶ 24:17 Charles Packer There's some stuff on SuiAgent, I think. 3 times in the scene
How Claude Plays Pokémon was made
- ▶ 0:48 Alessio Fanelli We are Eric from the sweet agent before.
Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
- ▶ 4:58 unnamed speaker So we did an episode with Anthropic about their SWE agent work and like the SWE bench verified results that they had.
Windsurf: The Enterprise AI IDE
- ▶ 15:49 unnamed speaker Yeah, we did an episode with Anthropic about their recent, like, sweet agent, sweet bench, uh, results, and we talked about the human eval versus sweet bench, and, or, like, human eval is kind of like a Greenfield benchmark, you know, you…
The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic
- ▶ 17:17 Shawn Wang I think we'll go into sweet agent in a, in a little bit, but I kind of reject the fact that you need to choose one prompt and like have your whole performance be predicated on one prompt.
- ▶ 41:14 Shawn Wang I also wanted to spend a little bit of time with Sweet Agent. 4 times in the scene
- ▶ 49:19 Shawn Wang There's a run with SuiAgent plus Opus, but that's the official SuiBench guys doing it.
Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
- ▶ 36:06 Shawn Wang I would consider Devin to be the sort of biggest launch of the year as far as AI startups go, and, uh, you guys in the Princeton group worked on SWE agents alongside of SWE bench. 8 times in the scene