Nov 8, 2025 · 35m · latent-space

Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders

Alex Shaw · 8m spoken Mike Merrill · 8m spoken Andy Konwinski · 3m spoken Ludwig Schmidt · 1m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this launch event and panel hosted by Latent Space, the creators unveil Terminal-Bench 2.0—a rigorously audited real-world benchmark for AI coding agents—alongside Harbor, a unified open-source framework for containerized agent evaluation, reinforcement learning, and distributed cloud execution.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 1.7 Guest teaching 1.2 Guest disagreement 0.3 The hosts pushing back 1.2
05100:0010:0020:0030:000:00–4:06 · The hosts as informed peer 0/10 Retrospective on Terminal-Bench 1.0 and Motivation Presentation monologue by Mike Merrill reviewing the motivation and shortcomings of Terminal-Bench 1.0. The host is not actively participating in this segment.4:08–7:01 · The hosts as informed peer 0/10 Introducing Terminal-Bench 2.0 Features and Verification Standards Continuation of the keynote presentation by Mike Merrill detailing Terminal-Bench 2.0 tasks and audit standards with zero host involvement.7:02–13:10 · The hosts as informed peer 0/10 The Repetitive Workflow of AI Agent Development Alex Shaw delivers a technical walkthrough on the repetitive steps required in agent evaluation loops. Monologue format without host presence.13:13–20:12 · The hosts as informed peer 0/10 Launching Harbor for Agent Evaluation and Optimization Presentation of the Harbor framework followed by Atosh demonstrating live cloud parallel trace generation. No host interaction.20:16–29:20 · The hosts as informed peer 4/10 Summary of the Dual Launch and Community Appreciation The host joins the panel and moderates a friendly discussion on academic origins, Datacomp, LAWD grant support, and author incentives.29:22–34:58 · The hosts as informed peer 6/10 Q&A on Scientific Computing Benchmarks and Agent Abstractions The host challenges the team's assertion that any task fits the terminal format, pointing to evolving GUI computer use and long-running agent state. The speakers politely defend Harbor's container abstraction.0:00–4:06 · Guest teaching 0/10 Retrospective on Terminal-Bench 1.0 and Motivation Presentation monologue by Mike Merrill reviewing the motivation and shortcomings of Terminal-Bench 1.0. The host is not actively participating in this segment.4:08–7:01 · Guest teaching 0/10 Introducing Terminal-Bench 2.0 Features and Verification Standards Continuation of the keynote presentation by Mike Merrill detailing Terminal-Bench 2.0 tasks and audit standards with zero host involvement.7:02–13:10 · Guest teaching 0/10 The Repetitive Workflow of AI Agent Development Alex Shaw delivers a technical walkthrough on the repetitive steps required in agent evaluation loops. Monologue format without host presence.13:13–20:12 · Guest teaching 0/10 Launching Harbor for Agent Evaluation and Optimization Presentation of the Harbor framework followed by Atosh demonstrating live cloud parallel trace generation. No host interaction.20:16–29:20 · Guest teaching 3/10 Summary of the Dual Launch and Community Appreciation The host joins the panel and moderates a friendly discussion on academic origins, Datacomp, LAWD grant support, and author incentives.29:22–34:58 · Guest teaching 4/10 Q&A on Scientific Computing Benchmarks and Agent Abstractions The host challenges the team's assertion that any task fits the terminal format, pointing to evolving GUI computer use and long-running agent state. The speakers politely defend Harbor's container abstraction.0:00–4:06 · Guest disagreement 0/10 Retrospective on Terminal-Bench 1.0 and Motivation Presentation monologue by Mike Merrill reviewing the motivation and shortcomings of Terminal-Bench 1.0. The host is not actively participating in this segment.4:08–7:01 · Guest disagreement 0/10 Introducing Terminal-Bench 2.0 Features and Verification Standards Continuation of the keynote presentation by Mike Merrill detailing Terminal-Bench 2.0 tasks and audit standards with zero host involvement.7:02–13:10 · Guest disagreement 0/10 The Repetitive Workflow of AI Agent Development Alex Shaw delivers a technical walkthrough on the repetitive steps required in agent evaluation loops. Monologue format without host presence.13:13–20:12 · Guest disagreement 0/10 Launching Harbor for Agent Evaluation and Optimization Presentation of the Harbor framework followed by Atosh demonstrating live cloud parallel trace generation. No host interaction.20:16–29:20 · Guest disagreement 0/10 Summary of the Dual Launch and Community Appreciation The host joins the panel and moderates a friendly discussion on academic origins, Datacomp, LAWD grant support, and author incentives.29:22–34:58 · Guest disagreement 2/10 Q&A on Scientific Computing Benchmarks and Agent Abstractions The host challenges the team's assertion that any task fits the terminal format, pointing to evolving GUI computer use and long-running agent state. The speakers politely defend Harbor's container abstraction.0:00–4:06 · The hosts pushing back 0/10 Retrospective on Terminal-Bench 1.0 and Motivation Presentation monologue by Mike Merrill reviewing the motivation and shortcomings of Terminal-Bench 1.0. The host is not actively participating in this segment.4:08–7:01 · The hosts pushing back 0/10 Introducing Terminal-Bench 2.0 Features and Verification Standards Continuation of the keynote presentation by Mike Merrill detailing Terminal-Bench 2.0 tasks and audit standards with zero host involvement.7:02–13:10 · The hosts pushing back 0/10 The Repetitive Workflow of AI Agent Development Alex Shaw delivers a technical walkthrough on the repetitive steps required in agent evaluation loops. Monologue format without host presence.13:13–20:12 · The hosts pushing back 0/10 Launching Harbor for Agent Evaluation and Optimization Presentation of the Harbor framework followed by Atosh demonstrating live cloud parallel trace generation. No host interaction.20:16–29:20 · The hosts pushing back 1/10 Summary of the Dual Launch and Community Appreciation The host joins the panel and moderates a friendly discussion on academic origins, Datacomp, LAWD grant support, and author incentives.29:22–34:58 · The hosts pushing back 6/10 Q&A on Scientific Computing Benchmarks and Agent Abstractions The host challenges the team's assertion that any task fits the terminal format, pointing to evolving GUI computer use and long-running agent state. The speakers politely defend Harbor's container abstraction.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 33:08 Alex Shaw defends the containerized program abstraction

Alex Shaw counters the host's critique by asserting that fundamentally any agent is simply an executable program, regardless of whether it interacts via terminal or GUI.

Hardest push from the hosts ▶ 32:18 Host pushes back on locking in text-based agent abstractions

The host directly challenges the speakers, citing trends at Cognition and arguing that Harbor risks locking in a rigid, text-based paradigm just as computer-use GUIs return.

Biggest teaching moment ▶ 33:10 Explaining that GUI agents operate inside container sandboxes

Alex Shaw clarifies for the host that Harbor's container model does not preclude visual agents because the GUI environment itself can execute inside the container.

The host holds their own ▶ 32:18 Host cites stateful memory and multimodal computer use

The host draws on current industry observations from Cognition and life sciences imaging to press the team on the limitations of short-horizon CLI benchmarking.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Retrospective on Terminal-Bench 1.0 and Motivation 0000 Presentation monologue by Mike Merrill reviewing the motivation and shortcomings of Terminal-Bench 1.0. The host is not actively participating in this segment.
Introducing Terminal-Bench 2.0 Features and Verification Standards 0000 Continuation of the keynote presentation by Mike Merrill detailing Terminal-Bench 2.0 tasks and audit standards with zero host involvement.
The Repetitive Workflow of AI Agent Development 0000 Alex Shaw delivers a technical walkthrough on the repetitive steps required in agent evaluation loops. Monologue format without host presence.
Launching Harbor for Agent Evaluation and Optimization 0000 Presentation of the Harbor framework followed by Atosh demonstrating live cloud parallel trace generation. No host interaction.
Summary of the Dual Launch and Community Appreciation 4301 The host joins the panel and moderates a friendly discussion on academic origins, Datacomp, LAWD grant support, and author incentives.
Q&A on Scientific Computing Benchmarks and Agent Abstractions 6426 The host challenges the team's assertion that any task fits the terminal format, pointing to evolving GUI computer use and long-running agent state. The speakers politely defend Harbor's container abstraction.

Statements from this episode (11)

Insight
Merrill: Launching benchmarks is the best way to steer AI development
“And so, if you have an opinion about how AI should work, what's the best way to get other people to work on it? And that's by launching a benchmark.”
Mike Merrill Nov 8, 2025 ▶ 1:35
Assertion Not checkable as stated
Merrill: Terminal-Bench is used by every frontier AI lab
“It's used by all Frontier Labs in some capacity.”
Mike Merrill Nov 8, 2025 ▶ 2:18
Assertion Not checkable as stated
Terminal-Bench 2.0 tasks are based on real paid workflows
“Every task in the benchmark is based off of real work that people do out in the world and are paid for.”
Mike Merrill Nov 8, 2025 ▶ 4:20
Assertion Not checkable as stated
Merrill: Terminal-Bench team spent 300+ hours auditing 89 benchmark tasks
“So I estimate we've spent at least 300 hours digging through each one of these 89 tasks, trying to figure out whether or not they're broken, whether or not they're some way to cheat, whether or not they're reproducible, and really obsessing over quality while …”
Mike Merrill Nov 8, 2025 ▶ 4:32
Assertion Supported
Merrill: Terminus-2 harness yields highest agent performance on Terminal-Bench 2.0
“Terminus-II, which is our harness for measuring language models, is the highest performing agent for many of these models.”
Mike Merrill Nov 8, 2025 ▶ 6:10
Insight
Shaw: Current AI coding benchmarks rely on redundant bespoke test harnesses
“In fact, every single benchmark that gets released at least I would say maybe all coding benchmarks that get released at this point are some form of instruction container tests with some bespoke harness that was coded up That feels very analogous to every othe…”
Alex Shaw Nov 8, 2025 ▶ 12:43
Assertion Not checkable as stated
Shaw: Harbor's standard task format can express most existing AI evaluations
“Specifically Harbor has a standard task format, which is an iteration of the terminal bench task format which we found to be very flexible and can often express most of the existing evaluations.”
Alex Shaw Nov 8, 2025 ▶ 13:33
Disclosure
Konwinski: Terminal-Bench is the largest LAWD Slingshot grant investment to date
“Terminal Bench is the largest investment we've made in any one of the slingshots so far.”
Andy Konwinski Nov 8, 2025 ▶ 26:26
Assertion Supported
Shaw: Terminal-Bench accepted 89 of 250 crowdsourced tasks for co-authorship
“We told people if they created three tasks for Terminal Bench, they could be a co-author on the paper that we eventually published, and I think we got maybe 250 task contributions, and 89 of them made it into the benchmark”
Alex Shaw Nov 8, 2025 ▶ 27:01
Insight
Schmidt: Co-authorship functions like startup equity for crowdsourced research
“If you make it a paper and you give everyone a share in the project, this is a little bit like in the startup ecosystem, right? Everyone gets a little bit of equity. Everyone gets a little bit of co-authorship.”
Ludwig Schmidt Nov 8, 2025 ▶ 28:47
Opinion
Merrill: Text-based agents are currently the frontier of performance
“We like text-based agents. We think right now they're the frontier of agentic performance. This very well may change.”
Mike Merrill Nov 8, 2025 ▶ 33:58
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.