Apr 2, 2025 · 31m · latent-space
The #1 SWE-Bench Verified Agent
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
Augment Code leader Guy Gur-Ari joins the Latent Space podcast to discuss their top-ranked SWE-bench Verified agent, breaking down the architecture, evaluation pipelines, and enterprise philosophy behind production coding agents.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 26.6% of the talking time here. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Guy rejects the premise that current coding models represent solved AGI, noting that developers still have to manually decompose problems into bite-sized tasks.
Hardest push from the hosts ▶ 14:35 Challenging proprietary model developmentSwyx presses Guy on why competitors like Poolside and Magic invest heavily in proprietary models while Augment chose off-the-shelf LLMs.
Biggest teaching moment ▶ 8:52 Explaining ensembling UX limitationsGuy educates the hosts on why ensembling agents is fundamentally a user experience and supervision problem rather than just an API cost issue.
The host holds their own ▶ 4:26 Framing the hybrid model cloud analogySwyx draws an analogy between historical hybrid cloud enterprise tooling and the modern necessity of multi-model agent orchestration.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Augment Agent Launch and SWE-Bench Verified Success | 6 | 3 | 1 | 2 | Hosts demonstrate strong familiarity with SWE-Bench and Anthropic's sequential thinking MCP implementation. Guy explains that while off-the-shelf models are used for generation, codebase understanding relies on custom retrieval models, and ensembling gives an extra boost. | |
| Experimentation Frameworks and Production Evaluation Pipelines | 5 | 4 | 1 | 1 | Swyx asks about experimentation harnesses and eval frameworks. Guy outlines Augment's progression from 10-sample notebook testing to massive evals, code execution ground-truth validation, and human contractor evaluations for agent workflows. | |
| Technical Wins, Orientation Agents, and Persistent Memory | 6 | 3 | 1 | 2 | Alessio queries the specifics of ensembling costs and orientation agents from Augment's announcements. Guy details how the UX of multi-trajectory supervision is the real bottleneck rather than just compute cost, and how orientation agents build persistent memory. | |
| Enterprise Coding Philosophy and Competitive Tooling Landscape | 6 | 2 | 1 | 2 | Swyx presses Guy on enterprise competitive positioning versus Devin, Factory, Magic, and Poolside. Guy articulates Augment's philosophy of meeting enterprise developers inside existing IDEs with complex monorepos rather than replacing IDEs. | |
| Live Demo Setup and Monorepo Build Configuration | 4 | 1 | 0 | 1 | A live demo segment where Guy demonstrates the agent implementing a VS Code extension feature directly from a Linear ticket using Bazel and MCP. The hosts follow along and observe the tool creation and PR generation flow. | |
| Future of the IDE and Model Context Protocol | 5 | 3 | 1 | 2 | Alessio and Swyx ask whether native integrations will eventually be abandoned for Anthropic's Model Context Protocol (MCP). Guy notes that auth flows and developer friction currently necessitate native integrations despite broad MCP support. | |
| Frontier Reinforcement Learning Research for Coding Models | 6 | 3 | 2 | 2 | Alessio and Swyx prompt Guy on frontier RL methods for coding and Google's recent model revival. Guy references the SWIRL paper and DeepSeek R1/GRPO, while pushing back on the notion that current models have reached AGI headroom. |