The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 10 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Opinion
Neubig: Anthropic's MCP Duplicates Existing APIs With Little Added Value
“We already have an API for GitHub. So why do we need an MCP for GitHub, right? You know, like GitHub has an API. The GitHub API is evolving. We can look up the GitHub API documentation. So it seems like kind of duplicated a little bit. And also they have a set…”
Graham Neubig Dec 25, 2024 ▶ 41:29 Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
Insight
Neubig: Arbitrary Python Execution Beats Individual Agent Tool Calls
“And the method that we adopt in open hands instead is we provide these tools, but we provide them by just giving a coding agent the ability to call arbitrary Python code. And in the arbitrary Python code, it can call these tools. We expose these tools as APIs …”
Graham Neubig Dec 25, 2024 ▶ 8:21 Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
Opinion
Neubig: Nobody Has A Good Answer For Human-Agent Interface Design
“I don't think anybody has a good answer to this, and I don't think we have a good answer to this”
Graham Neubig Dec 25, 2024 ▶ 11:04 Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
Disclosure
Neubig: OpenHands Uses A Single Prompt With No Multi-Agent Systems
“So in open hands, we do very light planning. We have a single prompt. We don't have any multi agent systems.”
Graham Neubig Dec 25, 2024 ▶ 17:09 Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
Insight
Neubig: Browsers, Terminals, And Code Editors Constitute The Core Agent Toolset
“Let's say I gave you a web browser and a terminal or a file system and the ability to edit text or code. What could you do with that? Everything. Yeah, probably a lot of things. This is like 99% of my, you know, daily daily life, I guess when I'm working. So I…”
Graham Neubig Dec 25, 2024 ▶ 0:42 Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
Disclosure
Neubig: OpenHands Equips Its Agent With Only Five Or Six Tools
“We're kind of extreme. And we're only giving the agent five tools or maybe six tools.”
Graham Neubig Dec 25, 2024 ▶ 9:08 Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
Insight
Neubig: Inadequate Information Gathering Is The Biggest Agent Failure Mode
“So I think actually probably the biggest thing that it fails at is. Or that our agent plus Claude fails at is insufficient information gathering before trying to solve the task, and so if you provide all, if you provide instructions that it should do informati…”
Graham Neubig Dec 25, 2024 ▶ 43:42 Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
Assertion Supported
Swyx: OpenHands is ranked number one on SWE-bench Full
“He started open hand is currently still number one on SweetBench full, which is the hardest one.”
Shawn Wang Jan 1, 2025 ▶ 19:45 2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
Assertion Supported
Swyx: SWE-bench resolution rates surged from 13% to ~50% in 2024
“Keep in mind, we started the year at 13%. And so now we're about 50 open hands is around there.”
Shawn Wang Jan 1, 2025 ▶ 1:08:29 2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
Disclosure
Neubig: All Hands AI Is Releasing A Unified Agent Benchmark
“We don't have benchmarks that test whether agents can code and do web navigation. But we're working on that and hoping to release something in the next week or two.”
Graham Neubig Dec 25, 2024 ▶ 23:27 Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.