Graham Neubig

9 statements across 2 episodes · 4 bullish · 4 bearish · 2 people on the record · first statement Dec 25, 2024 by Graham Neubig · said 10 times in 4 episodes since 2024 · across every show →

On the record as a speaker too: Graham Neubig's record, appearances and statements → this page counts the times other people say the name.

Mentions by year

brought up most by Shawn Wang (5), Swyx (Marcos Swix) (4)

tap a year for its mentions
00418220242025episodesmentions
01220242025episodes it came up in
00214220242025episodesmentions per episode
2025 8 mentions in 2 episodes 4 per episode
2024 2 mentions in 2 episodes 1 per episode

every mention, scene by scene, with the transcript →

Everything said about Graham Neubig, oldest first

Dec 25, 2024 negative
Opinion
Neubig: AI Agents Including Claude Are Ineffective At Asking For Help
“I think it, my impression is that agents are not very good at asking for help, even Claude. So like when they ask for help, they'll ask for help when they don't need it and then won't ask for help when they do need it.”
Graham Neubig Dec 25, 2024 ▶ 34:08 Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
Dec 25, 2024 bearish
Assertion Supported
Neubig: SWE-bench Scores Are Inflated By Training Data Contamination
“Sweebench is on popular open source repos and all of these popular open source repos were included in the training data for all of the language models. And so, the language models already know these repos. In some cases, the language models already know the in…”
Graham Neubig Dec 25, 2024 ▶ 32:49 Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
Dec 25, 2024 positive
Disclosure
Neubig: Agents Solve 80 To 90 Percent Of Tasks With Feedback
“Actual ability of models is maybe closer to 30 to 40%. So 30 to 40% of the things that I want an agent to solve on my own repos, it just solves without any human intervention. 80 to 90% it can solve without me opening an IDE, but I need to give it feedback.”
Graham Neubig Dec 25, 2024 ▶ 26:47 Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
Dec 25, 2024
Disclosure
Neubig: OpenHands Uses A Single Prompt With No Multi-Agent Systems
“So in open hands, we do very light planning. We have a single prompt. We don't have any multi agent systems.”
Graham Neubig Dec 25, 2024 ▶ 17:09 Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
Dec 25, 2024 positive
Disclosure
Neubig Uses AI Coding Agents Five To Ten Times Daily
“I use coding agents maybe five to 10 times a day. To help me solve my own problems.”
Graham Neubig Dec 25, 2024 ▶ 2:15 Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
Dec 25, 2024 bullish
Prediction Not checkable as stated
Neubig: Every Major LLM Trainer Will Focus On Agents By Mid-2025
“My prediction is every large LM trainer will be focusing on training models as agents. So every large language model will be a better agent model. By mid 20, 25.”
Graham Neubig Dec 25, 2024 ▶ 24:08 Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
Dec 25, 2024 bearish
Assertion Not checkable as stated
Neubig: Most APIs Lack The Fine-Grained Authentication Needed For Agents
“For other things, they're totally not prepared to give that sort of fine-grained control. Like, most APIs don't have something like a fine-grained authentication token, and that goes into my, like, comment that we're gonna need to prepare the world for agents,…”
Graham Neubig Dec 25, 2024 ▶ 49:23 Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
Dec 25, 2024 negative
Insight
Neubig: Inadequate Information Gathering Is The Biggest Agent Failure Mode
“So I think actually probably the biggest thing that it fails at is. Or that our agent plus Claude fails at is insufficient information gathering before trying to solve the task, and so if you provide all, if you provide instructions that it should do informati…”
Graham Neubig Dec 25, 2024 ▶ 43:42 Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
Jan 1, 2025 positive
Assertion Supported
Swyx: OpenHands is ranked number one on SWE-bench Full
“He started open hand is currently still number one on SweetBench full, which is the hardest one.”
Shawn Wang Jan 1, 2025 ▶ 19:45 2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.