Nov 8, 2025 · 35m · latent-space
Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this launch event and panel hosted by Latent Space, the creators unveil Terminal-Bench 2.0—a rigorously audited real-world benchmark for AI coding agents—alongside Harbor, a unified open-source framework for containerized agent evaluation, reinforcement learning, and distributed cloud execution.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Alex Shaw counters the host's critique by asserting that fundamentally any agent is simply an executable program, regardless of whether it interacts via terminal or GUI.
Hardest push from the hosts ▶ 32:18 Host pushes back on locking in text-based agent abstractionsThe host directly challenges the speakers, citing trends at Cognition and arguing that Harbor risks locking in a rigid, text-based paradigm just as computer-use GUIs return.
Biggest teaching moment ▶ 33:10 Explaining that GUI agents operate inside container sandboxesAlex Shaw clarifies for the host that Harbor's container model does not preclude visual agents because the GUI environment itself can execute inside the container.
The host holds their own ▶ 32:18 Host cites stateful memory and multimodal computer useThe host draws on current industry observations from Cognition and life sciences imaging to press the team on the limitations of short-horizon CLI benchmarking.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Retrospective on Terminal-Bench 1.0 and Motivation | 0 | 0 | 0 | 0 | Presentation monologue by Mike Merrill reviewing the motivation and shortcomings of Terminal-Bench 1.0. The host is not actively participating in this segment. | |
| Introducing Terminal-Bench 2.0 Features and Verification Standards | 0 | 0 | 0 | 0 | Continuation of the keynote presentation by Mike Merrill detailing Terminal-Bench 2.0 tasks and audit standards with zero host involvement. | |
| The Repetitive Workflow of AI Agent Development | 0 | 0 | 0 | 0 | Alex Shaw delivers a technical walkthrough on the repetitive steps required in agent evaluation loops. Monologue format without host presence. | |
| Launching Harbor for Agent Evaluation and Optimization | 0 | 0 | 0 | 0 | Presentation of the Harbor framework followed by Atosh demonstrating live cloud parallel trace generation. No host interaction. | |
| Summary of the Dual Launch and Community Appreciation | 4 | 3 | 0 | 1 | The host joins the panel and moderates a friendly discussion on academic origins, Datacomp, LAWD grant support, and author incentives. | |
| Q&A on Scientific Computing Benchmarks and Agent Abstractions | 6 | 4 | 2 | 6 | The host challenges the team's assertion that any task fits the terminal format, pointing to evolving GUI computer use and long-running agent state. The speakers politely defend Harbor's container abstraction. |