Prediction certainty 3/5 debate potential 2/5

Ensembling coding agents will transition from a cost to a UX challenge

Guy Gur-Ari · The #1 SWE-Bench Verified Agent · Apr 2, 2025 · at 8:54

Guy Gur-Ari, founding team/engineering leader at Augment Code, discusses the trade-offs of using ensembling techniques in autonomous coding agents.

0:00 / 0:25exact quote · 25.4s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“I think it will become less of over time as cost will come down. I expect it will become less of a cost question and more of a UX question because with ensembling, well, what we found is that users really want to see what the agent is doing and follow along. And if you're doing ensembling, you don't have one trajectory anymore. You have several. And then the question of how do users supervise that? As it's happening becomes pretty tricky.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Guy Gur-Ari

Insight
Gur-Ari: SWE-bench fails to effectively test true codebase understanding
“What we found with SuiteBench is an interesting benchmark, very useful for doing some things like refining prompts, for example, and trying out different tools. Specifically for codebase understanding, in the beginning, we were hopeful That it will tell us mor…”
Guy Gur-Ari Apr 2, 2025 ▶ 2:07 The #1 SWE-Bench Verified Agent
Assertion Not checkable as stated
Sequential Thinking MCP outperformed Claude 3.7 native reasoning mode in evaluations
“We tried reasoning mode as well with the new three seven. And we didn't see that much of a bump in performance. We don't know if this is something that's code specific or not. I don't have an insight. We tried both and yeah, sequential thinking worked better.”
Guy Gur-Ari Apr 2, 2025 ▶ 3:57 The #1 SWE-Bench Verified Agent
Prediction Not checkable as stated
Developers will eventually spend 80% of their time controlling agents outside IDEs
“I think in the future, at some point, my guess is that the IDE is going to become less of the focal point and more like an app that you can launch when you need to dig in deeper, but you spend most of your time away from it. So maybe 80% of your time is in a w…”
Guy Gur-Ari Apr 2, 2025 ▶ 24:18 The #1 SWE-Bench Verified Agent
Assertion Supported
Gur-Ari: Augment Code Achieved #1 on SWE-Bench Verified
“We just made number one on Sweepbench. So for us Sweepbench has been a useful tool for exploring how can we get the most out of agents. And so we were able to get the best result on Sweepbench verified right now.”
Guy Gur-Ari Apr 2, 2025 ▶ 1:24 The #1 SWE-Bench Verified Agent
Disclosure
Augment's #1 SWE-Bench agent uses off-the-shelf models, unlike their custom product
“The generation models for Sweebench, it's all off the shelf models. For the product, it includes our own custom trained models that help the model with, that help the agent with code-based understanding.”
Guy Gur-Ari Apr 2, 2025 ▶ 1:42 The #1 SWE-Bench Verified Agent
Insight
Ground-truth evaluations without code execution provide significant mileage for AI
“So code execution, like checking the correctness of solutions and running tests automatically can help, although not, although you can get a lot of mileage out of evals that don't have code execution in them that just compare against ground truth.”
Guy Gur-Ari Apr 2, 2025 ▶ 6:48 The #1 SWE-Bench Verified Agent
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.