Apr 2, 2025 · 31m · latent-space

The #1 SWE-Bench Verified Agent

Guy Gur-Ari · 20m spoken Shawn Wang · 4m spoken Alessio Fanelli · 2m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Augment Code leader Guy Gur-Ari joins the Latent Space podcast to discuss their top-ranked SWE-bench Verified agent, breaking down the architecture, evaluation pipelines, and enterprise philosophy behind production coding agents.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 26.6% of the talking time here. How this is scored →

The hosts as informed peer 5.4 Guest teaching 2.7 Guest disagreement 1.0 The hosts pushing back 1.7
05100:0010:0020:0030:000:36–5:28 · The hosts as informed peer 6/10 Augment Agent Launch and SWE-Bench Verified Success Hosts demonstrate strong familiarity with SWE-Bench and Anthropic's sequential thinking MCP implementation. Guy explains that while off-the-shelf models are used for generation, codebase understanding relies on custom retrieval models, and ensembling gives an extra boost.5:29–7:44 · The hosts as informed peer 5/10 Experimentation Frameworks and Production Evaluation Pipelines Swyx asks about experimentation harnesses and eval frameworks. Guy outlines Augment's progression from 10-sample notebook testing to massive evals, code execution ground-truth validation, and human contractor evaluations for agent workflows.7:47–12:07 · The hosts as informed peer 6/10 Technical Wins, Orientation Agents, and Persistent Memory Alessio queries the specifics of ensembling costs and orientation agents from Augment's announcements. Guy details how the UX of multi-trajectory supervision is the real bottleneck rather than just compute cost, and how orientation agents build persistent memory.12:09–15:44 · The hosts as informed peer 6/10 Enterprise Coding Philosophy and Competitive Tooling Landscape Swyx presses Guy on enterprise competitive positioning versus Devin, Factory, Magic, and Poolside. Guy articulates Augment's philosophy of meeting enterprise developers inside existing IDEs with complex monorepos rather than replacing IDEs.15:44–23:20 · The hosts as informed peer 4/10 Live Demo Setup and Monorepo Build Configuration A live demo segment where Guy demonstrates the agent implementing a VS Code extension feature directly from a Linear ticket using Bazel and MCP. The hosts follow along and observe the tool creation and PR generation flow.23:25–26:48 · The hosts as informed peer 5/10 Future of the IDE and Model Context Protocol Alessio and Swyx ask whether native integrations will eventually be abandoned for Anthropic's Model Context Protocol (MCP). Guy notes that auth flows and developer friction currently necessitate native integrations despite broad MCP support.26:50–30:45 · The hosts as informed peer 6/10 Frontier Reinforcement Learning Research for Coding Models Alessio and Swyx prompt Guy on frontier RL methods for coding and Google's recent model revival. Guy references the SWIRL paper and DeepSeek R1/GRPO, while pushing back on the notion that current models have reached AGI headroom.0:36–5:28 · Guest teaching 3/10 Augment Agent Launch and SWE-Bench Verified Success Hosts demonstrate strong familiarity with SWE-Bench and Anthropic's sequential thinking MCP implementation. Guy explains that while off-the-shelf models are used for generation, codebase understanding relies on custom retrieval models, and ensembling gives an extra boost.5:29–7:44 · Guest teaching 4/10 Experimentation Frameworks and Production Evaluation Pipelines Swyx asks about experimentation harnesses and eval frameworks. Guy outlines Augment's progression from 10-sample notebook testing to massive evals, code execution ground-truth validation, and human contractor evaluations for agent workflows.7:47–12:07 · Guest teaching 3/10 Technical Wins, Orientation Agents, and Persistent Memory Alessio queries the specifics of ensembling costs and orientation agents from Augment's announcements. Guy details how the UX of multi-trajectory supervision is the real bottleneck rather than just compute cost, and how orientation agents build persistent memory.12:09–15:44 · Guest teaching 2/10 Enterprise Coding Philosophy and Competitive Tooling Landscape Swyx presses Guy on enterprise competitive positioning versus Devin, Factory, Magic, and Poolside. Guy articulates Augment's philosophy of meeting enterprise developers inside existing IDEs with complex monorepos rather than replacing IDEs.15:44–23:20 · Guest teaching 1/10 Live Demo Setup and Monorepo Build Configuration A live demo segment where Guy demonstrates the agent implementing a VS Code extension feature directly from a Linear ticket using Bazel and MCP. The hosts follow along and observe the tool creation and PR generation flow.23:25–26:48 · Guest teaching 3/10 Future of the IDE and Model Context Protocol Alessio and Swyx ask whether native integrations will eventually be abandoned for Anthropic's Model Context Protocol (MCP). Guy notes that auth flows and developer friction currently necessitate native integrations despite broad MCP support.26:50–30:45 · Guest teaching 3/10 Frontier Reinforcement Learning Research for Coding Models Alessio and Swyx prompt Guy on frontier RL methods for coding and Google's recent model revival. Guy references the SWIRL paper and DeepSeek R1/GRPO, while pushing back on the notion that current models have reached AGI headroom.0:36–5:28 · Guest disagreement 1/10 Augment Agent Launch and SWE-Bench Verified Success Hosts demonstrate strong familiarity with SWE-Bench and Anthropic's sequential thinking MCP implementation. Guy explains that while off-the-shelf models are used for generation, codebase understanding relies on custom retrieval models, and ensembling gives an extra boost.5:29–7:44 · Guest disagreement 1/10 Experimentation Frameworks and Production Evaluation Pipelines Swyx asks about experimentation harnesses and eval frameworks. Guy outlines Augment's progression from 10-sample notebook testing to massive evals, code execution ground-truth validation, and human contractor evaluations for agent workflows.7:47–12:07 · Guest disagreement 1/10 Technical Wins, Orientation Agents, and Persistent Memory Alessio queries the specifics of ensembling costs and orientation agents from Augment's announcements. Guy details how the UX of multi-trajectory supervision is the real bottleneck rather than just compute cost, and how orientation agents build persistent memory.12:09–15:44 · Guest disagreement 1/10 Enterprise Coding Philosophy and Competitive Tooling Landscape Swyx presses Guy on enterprise competitive positioning versus Devin, Factory, Magic, and Poolside. Guy articulates Augment's philosophy of meeting enterprise developers inside existing IDEs with complex monorepos rather than replacing IDEs.15:44–23:20 · Guest disagreement 0/10 Live Demo Setup and Monorepo Build Configuration A live demo segment where Guy demonstrates the agent implementing a VS Code extension feature directly from a Linear ticket using Bazel and MCP. The hosts follow along and observe the tool creation and PR generation flow.23:25–26:48 · Guest disagreement 1/10 Future of the IDE and Model Context Protocol Alessio and Swyx ask whether native integrations will eventually be abandoned for Anthropic's Model Context Protocol (MCP). Guy notes that auth flows and developer friction currently necessitate native integrations despite broad MCP support.26:50–30:45 · Guest disagreement 2/10 Frontier Reinforcement Learning Research for Coding Models Alessio and Swyx prompt Guy on frontier RL methods for coding and Google's recent model revival. Guy references the SWIRL paper and DeepSeek R1/GRPO, while pushing back on the notion that current models have reached AGI headroom.0:36–5:28 · The hosts pushing back 2/10 Augment Agent Launch and SWE-Bench Verified Success Hosts demonstrate strong familiarity with SWE-Bench and Anthropic's sequential thinking MCP implementation. Guy explains that while off-the-shelf models are used for generation, codebase understanding relies on custom retrieval models, and ensembling gives an extra boost.5:29–7:44 · The hosts pushing back 1/10 Experimentation Frameworks and Production Evaluation Pipelines Swyx asks about experimentation harnesses and eval frameworks. Guy outlines Augment's progression from 10-sample notebook testing to massive evals, code execution ground-truth validation, and human contractor evaluations for agent workflows.7:47–12:07 · The hosts pushing back 2/10 Technical Wins, Orientation Agents, and Persistent Memory Alessio queries the specifics of ensembling costs and orientation agents from Augment's announcements. Guy details how the UX of multi-trajectory supervision is the real bottleneck rather than just compute cost, and how orientation agents build persistent memory.12:09–15:44 · The hosts pushing back 2/10 Enterprise Coding Philosophy and Competitive Tooling Landscape Swyx presses Guy on enterprise competitive positioning versus Devin, Factory, Magic, and Poolside. Guy articulates Augment's philosophy of meeting enterprise developers inside existing IDEs with complex monorepos rather than replacing IDEs.15:44–23:20 · The hosts pushing back 1/10 Live Demo Setup and Monorepo Build Configuration A live demo segment where Guy demonstrates the agent implementing a VS Code extension feature directly from a Linear ticket using Bazel and MCP. The hosts follow along and observe the tool creation and PR generation flow.23:25–26:48 · The hosts pushing back 2/10 Future of the IDE and Model Context Protocol Alessio and Swyx ask whether native integrations will eventually be abandoned for Anthropic's Model Context Protocol (MCP). Guy notes that auth flows and developer friction currently necessitate native integrations despite broad MCP support.26:50–30:45 · The hosts pushing back 2/10 Frontier Reinforcement Learning Research for Coding Models Alessio and Swyx prompt Guy on frontier RL methods for coding and Google's recent model revival. Guy references the SWIRL paper and DeepSeek R1/GRPO, while pushing back on the notion that current models have reached AGI headroom.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 41.9% · guest 58.1%0:00 · the hosts 41.9% · guest 58.1%3:00 · the hosts 43.9% · guest 56.1%3:00 · the hosts 43.9% · guest 56.1%6:00 · the hosts 34.2% · guest 65.8%6:00 · the hosts 34.2% · guest 65.8%9:00 · the hosts 11.2% · guest 88.8%9:00 · the hosts 11.2% · guest 88.8%12:00 · the hosts 28.8% · guest 71.2%12:00 · the hosts 28.8% · guest 71.2%15:00 · the hosts 26.9% · guest 73.1%15:00 · the hosts 26.9% · guest 73.1%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 9.4% · guest 90.6%21:00 · the hosts 9.4% · guest 90.6%24:00 · the hosts 27.6% · guest 72.4%24:00 · the hosts 27.6% · guest 72.4%27:00 · the hosts 40.5% · guest 59.5%27:00 · the hosts 40.5% · guest 59.5%30:00 · the hosts 31.8% · guest 68.2%30:00 · the hosts 31.8% · guest 68.2%
Sharpest disagreement ▶ 29:51 Pushback on AGI claims in coding

Guy rejects the premise that current coding models represent solved AGI, noting that developers still have to manually decompose problems into bite-sized tasks.

Hardest push from the hosts ▶ 14:35 Challenging proprietary model development

Swyx presses Guy on why competitors like Poolside and Magic invest heavily in proprietary models while Augment chose off-the-shelf LLMs.

Biggest teaching moment ▶ 8:52 Explaining ensembling UX limitations

Guy educates the hosts on why ensembling agents is fundamentally a user experience and supervision problem rather than just an API cost issue.

The host holds their own ▶ 4:26 Framing the hybrid model cloud analogy

Swyx draws an analogy between historical hybrid cloud enterprise tooling and the modern necessity of multi-model agent orchestration.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Augment Agent Launch and SWE-Bench Verified Success 6312 Hosts demonstrate strong familiarity with SWE-Bench and Anthropic's sequential thinking MCP implementation. Guy explains that while off-the-shelf models are used for generation, codebase understanding relies on custom retrieval models, and ensembling gives an extra boost.
Experimentation Frameworks and Production Evaluation Pipelines 5411 Swyx asks about experimentation harnesses and eval frameworks. Guy outlines Augment's progression from 10-sample notebook testing to massive evals, code execution ground-truth validation, and human contractor evaluations for agent workflows.
Technical Wins, Orientation Agents, and Persistent Memory 6312 Alessio queries the specifics of ensembling costs and orientation agents from Augment's announcements. Guy details how the UX of multi-trajectory supervision is the real bottleneck rather than just compute cost, and how orientation agents build persistent memory.
Enterprise Coding Philosophy and Competitive Tooling Landscape 6212 Swyx presses Guy on enterprise competitive positioning versus Devin, Factory, Magic, and Poolside. Guy articulates Augment's philosophy of meeting enterprise developers inside existing IDEs with complex monorepos rather than replacing IDEs.
Live Demo Setup and Monorepo Build Configuration 4101 A live demo segment where Guy demonstrates the agent implementing a VS Code extension feature directly from a Linear ticket using Bazel and MCP. The hosts follow along and observe the tool creation and PR generation flow.
Future of the IDE and Model Context Protocol 5312 Alessio and Swyx ask whether native integrations will eventually be abandoned for Anthropic's Model Context Protocol (MCP). Guy notes that auth flows and developer friction currently necessitate native integrations despite broad MCP support.
Frontier Reinforcement Learning Research for Coding Models 6322 Alessio and Swyx prompt Guy on frontier RL methods for coding and Google's recent model revival. Guy references the SWIRL paper and DeepSeek R1/GRPO, while pushing back on the notion that current models have reached AGI headroom.

Statements from this episode (17)

Assertion Supported
Gur-Ari: Augment Code Achieved #1 on SWE-Bench Verified
“We just made number one on Sweepbench. So for us Sweepbench has been a useful tool for exploring how can we get the most out of agents. And so we were able to get the best result on Sweepbench verified right now.”
Guy Gur-Ari Apr 2, 2025 ▶ 1:24
Disclosure
Augment's #1 SWE-Bench agent uses off-the-shelf models, unlike their custom product
“The generation models for Sweebench, it's all off the shelf models. For the product, it includes our own custom trained models that help the model with, that help the agent with code-based understanding.”
Guy Gur-Ari Apr 2, 2025 ▶ 1:42
Insight
Gur-Ari: SWE-bench fails to effectively test true codebase understanding
“What we found with SuiteBench is an interesting benchmark, very useful for doing some things like refining prompts, for example, and trying out different tools. Specifically for codebase understanding, in the beginning, we were hopeful That it will tell us mor…”
Guy Gur-Ari Apr 2, 2025 ▶ 2:07
Assertion Not checkable as stated
Sequential Thinking MCP outperformed Claude 3.7 native reasoning mode in evaluations
“We tried reasoning mode as well with the new three seven. And we didn't see that much of a bump in performance. We don't know if this is something that's code specific or not. I don't have an insight. We tried both and yeah, sequential thinking worked better.”
Guy Gur-Ari Apr 2, 2025 ▶ 3:57
Insight
Engineers should begin AI feature evaluation with 10-sample interactive notebooks
“So I would say the process that I like to follow is in the beginning when developing a feature, come up with a curated set of samples. Could be as small as 10, 10 samples that you run through, and then yeah, it's all notebooks basically. You run through the sa…”
Guy Gur-Ari Apr 2, 2025 ▶ 6:10
Insight
Ground-truth evaluations without code execution provide significant mileage for AI
“So code execution, like checking the correctness of solutions and running tests automatically can help, although not, although you can get a lot of mileage out of evals that don't have code execution in them that just compare against ground truth.”
Guy Gur-Ari Apr 2, 2025 ▶ 6:48
Insight
Gur-Ari uses human contractors because chat and agent evaluations resist automation
“The last thing I can mention is we use contractors for evaluation where we cannot Do automatic evaluation. So with chats and agents, it becomes way harder to do things automatically. And so we use contractors for that.”
Guy Gur-Ari Apr 2, 2025 ▶ 7:33
Prediction Not checkable as stated
Ensembling coding agents will transition from a cost to a UX challenge
“I think it will become less of over time as cost will come down. I expect it will become less of a cost question and more of a UX question because with ensembling, well, what we found is that users really want to see what the agent is doing and follow along. A…”
Guy Gur-Ari Apr 2, 2025 ▶ 8:54
Insight
Better ensembling typically adds only a few percentage points on SWE-bench
“Ensembling is usually not the kind of thing that just solves a benchmark, just also in other domains, you typically can get like, yeah, and this, and first we bench probably a few percentage points is my guess.”
Guy Gur-Ari Apr 2, 2025 ▶ 9:46
Disclosure
Augment Code to ship multi-minute codebase orientation agent for thorough analysis
“We also have a feature that I think has not shipped yet, but we are planning to ship, which is a more thorough orientation. So this is something that will run for several minutes probably and try to do a pretty thorough job of trying to figure out You know, wh…”
Guy Gur-Ari Apr 2, 2025 ▶ 11:38
Prediction Not checkable as stated
Running multiple parallel agents will unlock most value for software developers
“Being able to run multiple agents and not just one is going to be the way to unlock Honestly, most of the value out of these agents, and that's what we're working toward.”
Guy Gur-Ari Apr 2, 2025 ▶ 13:56
Assertion Supported
Augment Code runs inside the Cursor editor and already has active users
“Yes, you can use augment inside of cursor. We actually have quite a few users doing that.”
Guy Gur-Ari Apr 2, 2025 ▶ 15:26
Assertion Supported
Augment Code supports MCP alongside native GitHub, Linear, and Notion integrations
“Within augment, we have we have both MCP support for complete extensibility of tools, but we, we've also built in a few first party integrations and serve them as tools to the agent. So we have GitHub linear notion that, that we use internally, and then we hav…”
Guy Gur-Ari Apr 2, 2025 ▶ 18:17
Prediction Not checkable as stated
Developers will eventually spend 80% of their time controlling agents outside IDEs
“I think in the future, at some point, my guess is that the IDE is going to become less of the focal point and more like an app that you can launch when you need to dig in deeper, but you spend most of your time away from it. So maybe 80% of your time is in a w…”
Guy Gur-Ari Apr 2, 2025 ▶ 24:18
Opinion
Gur-Ari highlights the SWIRL research paper for reinforcement learning in coding
“There was an interesting paper called SWIRL, which I thought was a really nice paper on how to do RL for, specifically for coding. So that's by Wei and other authors.”
Guy Gur-Ari Apr 2, 2025 ▶ 27:06
Opinion
Google's internal culture became significantly more aggressive and startup-like post-ChatGPT
“The culture has really shifted internally to be a lot more startup, be fast moving. Aggressive, and at least folks I talk to are optimistic about Google's chances to win this.”
Guy Gur-Ari Apr 2, 2025 ▶ 28:51
Assertion Supported
Augment Code open-sourced the SWE-bench implementation that reached number one
“We actually open sourced our implementation of Sweebench. So if you're curious how we got to number one, you'll be able to go see all the details of how we did it.”
Guy Gur-Ari Apr 2, 2025 ▶ 31:13
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.