Insight certainty 3/5 debate potential 2/5

Neubig: Inadequate Information Gathering Is The Biggest Agent Failure Mode

Graham Neubig · Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands) · Dec 25, 2024 · at 43:42

Graham Neubig, CMU professor and co-founder of AllHands AI (OpenHands), explains the primary failure patterns observed when running autonomous coding agents on software engineering tasks.

0:00 / 0:29exact quote · 29.1s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“So I think actually probably the biggest thing that it fails at is. Or that our agent plus Claude fails at is insufficient information gathering before trying to solve the task, and so if you provide all, if you provide instructions that it should do information gathering beforehand, it tends to do well. If you don't provide sufficient instructions, it will try to solve the task without, like, fully understanding the task first, and then fail, and then you need to go back and give you know, additional feedback.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Graham Neubig

Opinion
Neubig: Anthropic's MCP Duplicates Existing APIs With Little Added Value
“We already have an API for GitHub. So why do we need an MCP for GitHub, right? You know, like GitHub has an API. The GitHub API is evolving. We can look up the GitHub API documentation. So it seems like kind of duplicated a little bit. And also they have a set…”
Graham Neubig Dec 25, 2024 ▶ 41:29 Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
Insight
Neubig: Arbitrary Python Execution Beats Individual Agent Tool Calls
“And the method that we adopt in open hands instead is we provide these tools, but we provide them by just giving a coding agent the ability to call arbitrary Python code. And in the arbitrary Python code, it can call these tools. We expose these tools as APIs …”
Graham Neubig Dec 25, 2024 ▶ 8:21 Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
Assertion Supported
Neubig: SWE-bench Scores Are Inflated By Training Data Contamination
“Sweebench is on popular open source repos and all of these popular open source repos were included in the training data for all of the language models. And so, the language models already know these repos. In some cases, the language models already know the in…”
Graham Neubig Dec 25, 2024 ▶ 32:49 Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
Opinion
Neubig: GPT Loops On Errors While Claude Tries New Approaches
“So, like, GPT doesn't have very good air recovery ability. And so, because of this, it will go into loops and do the same thing over and over and over again, whereas Claude does not do this.”
Graham Neubig Dec 25, 2024 ▶ 14:25 Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
Opinion
Neubig: Claude Is The Best Agent Model, Open Models Lag Behind
“I still am under the impression that Claude is the best. The other closed models are, you know, not quite as good, and then the open models are a little bit behind that.”
Graham Neubig Dec 25, 2024 ▶ 15:16 Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
Insight
Neubig: Single Agents Adapt Better Than Rigid Multi-Agent Systems
“If you have a really, really good instruction following agent it will follow the instructions as long as things are working according to your plan, but let's say you need to deviate from your plan, you still have the flexibility to do this, and if you do expli…”
Graham Neubig Dec 25, 2024 ▶ 17:34 Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.