Oct 18, 2025 · 35m · latent-space

Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits

Mike Merrill · 13m spoken Alex Shaw · 11m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Alex and Mike Merrill join hosts Alessio and Swix to discuss Terminal-Bench, an open-source framework evaluating autonomous AI agents on diverse command-line tasks. They explore the advantages of terminal abstractions over GUIs, the benchmark's adoption across frontier AI labs, and methods for measuring agent intelligence and economic value.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 4.7 Guest teaching 4.7 Guest disagreement 1.3 The hosts pushing back 3.0
05100:0010:0020:0030:000:03–5:55 · The hosts as informed peer 2/10 Origins of Terminal-Bench and Frontier Lab Adoption The hosts welcome Alex and Mike, noting industry adoption by frontier labs. The guests recount the origins of Terminal-Bench from SWE-bench and the surprise mention in the Claude 3.5 Sonnet model card.5:55–10:31 · The hosts as informed peer 4/10 Terminal Abstraction and Expanding Beyond Pure Coding Alessio asks why they bet on terminal CLI over GUI and whether tasks reflect broader computing beyond coding. Mike explains why text-native terminal interfaces outperform GUI navigation, while Alex explains viewing tasks as general container environments.10:31–13:55 · The hosts as informed peer 3/10 Task Design, Machine Learning Benchmarks, and Biology Tools Swix inspects sample tasks on the live landing page. The guests explain complex task formulations like fastText ML training constraints and biology DNA sequence assembly designed by domain scientists.13:55–21:02 · The hosts as informed peer 6/10 Adapting External Benchmarks and the Meta-Framework Vision Alessio questions if adapting dozens of academic benchmarks yields diminishing returns versus first-principles task design. Mike and Alex explain why web scraping for evals is exhausted and why Terminal-Bench serves as a general meta-framework and RL training harness.21:02–26:37 · The hosts as informed peer 6/10 Disentangling Model Intelligence from Agent Harnesses Alessio asks how to decouple model intelligence from proprietary agent harnesses. Mike details Terminus as an unopinionated headless baseline, while Alex notes model capability differences far outweigh agent framework tuning.26:38–33:36 · The hosts as informed peer 7/10 Roadmap, Cloud Infrastructure, and Economic Evaluation Metrics Swix and Alessio push for multi-dimensional evals incorporating cost, latency, and scaled compute budgets. Mike argues the ultimate eval metric is economic profit and value generated, while Alex describes timeout ablation experiments.0:03–5:55 · Guest teaching 3/10 Origins of Terminal-Bench and Frontier Lab Adoption The hosts welcome Alex and Mike, noting industry adoption by frontier labs. The guests recount the origins of Terminal-Bench from SWE-bench and the surprise mention in the Claude 3.5 Sonnet model card.5:55–10:31 · Guest teaching 4/10 Terminal Abstraction and Expanding Beyond Pure Coding Alessio asks why they bet on terminal CLI over GUI and whether tasks reflect broader computing beyond coding. Mike explains why text-native terminal interfaces outperform GUI navigation, while Alex explains viewing tasks as general container environments.10:31–13:55 · Guest teaching 5/10 Task Design, Machine Learning Benchmarks, and Biology Tools Swix inspects sample tasks on the live landing page. The guests explain complex task formulations like fastText ML training constraints and biology DNA sequence assembly designed by domain scientists.13:55–21:02 · Guest teaching 6/10 Adapting External Benchmarks and the Meta-Framework Vision Alessio questions if adapting dozens of academic benchmarks yields diminishing returns versus first-principles task design. Mike and Alex explain why web scraping for evals is exhausted and why Terminal-Bench serves as a general meta-framework and RL training harness.21:02–26:37 · Guest teaching 5/10 Disentangling Model Intelligence from Agent Harnesses Alessio asks how to decouple model intelligence from proprietary agent harnesses. Mike details Terminus as an unopinionated headless baseline, while Alex notes model capability differences far outweigh agent framework tuning.26:38–33:36 · Guest teaching 5/10 Roadmap, Cloud Infrastructure, and Economic Evaluation Metrics Swix and Alessio push for multi-dimensional evals incorporating cost, latency, and scaled compute budgets. Mike argues the ultimate eval metric is economic profit and value generated, while Alex describes timeout ablation experiments.0:03–5:55 · Guest disagreement 1/10 Origins of Terminal-Bench and Frontier Lab Adoption The hosts welcome Alex and Mike, noting industry adoption by frontier labs. The guests recount the origins of Terminal-Bench from SWE-bench and the surprise mention in the Claude 3.5 Sonnet model card.5:55–10:31 · Guest disagreement 1/10 Terminal Abstraction and Expanding Beyond Pure Coding Alessio asks why they bet on terminal CLI over GUI and whether tasks reflect broader computing beyond coding. Mike explains why text-native terminal interfaces outperform GUI navigation, while Alex explains viewing tasks as general container environments.10:31–13:55 · Guest disagreement 0/10 Task Design, Machine Learning Benchmarks, and Biology Tools Swix inspects sample tasks on the live landing page. The guests explain complex task formulations like fastText ML training constraints and biology DNA sequence assembly designed by domain scientists.13:55–21:02 · Guest disagreement 2/10 Adapting External Benchmarks and the Meta-Framework Vision Alessio questions if adapting dozens of academic benchmarks yields diminishing returns versus first-principles task design. Mike and Alex explain why web scraping for evals is exhausted and why Terminal-Bench serves as a general meta-framework and RL training harness.21:02–26:37 · Guest disagreement 2/10 Disentangling Model Intelligence from Agent Harnesses Alessio asks how to decouple model intelligence from proprietary agent harnesses. Mike details Terminus as an unopinionated headless baseline, while Alex notes model capability differences far outweigh agent framework tuning.26:38–33:36 · Guest disagreement 2/10 Roadmap, Cloud Infrastructure, and Economic Evaluation Metrics Swix and Alessio push for multi-dimensional evals incorporating cost, latency, and scaled compute budgets. Mike argues the ultimate eval metric is economic profit and value generated, while Alex describes timeout ablation experiments.0:03–5:55 · The hosts pushing back 1/10 Origins of Terminal-Bench and Frontier Lab Adoption The hosts welcome Alex and Mike, noting industry adoption by frontier labs. The guests recount the origins of Terminal-Bench from SWE-bench and the surprise mention in the Claude 3.5 Sonnet model card.5:55–10:31 · The hosts pushing back 2/10 Terminal Abstraction and Expanding Beyond Pure Coding Alessio asks why they bet on terminal CLI over GUI and whether tasks reflect broader computing beyond coding. Mike explains why text-native terminal interfaces outperform GUI navigation, while Alex explains viewing tasks as general container environments.10:31–13:55 · The hosts pushing back 1/10 Task Design, Machine Learning Benchmarks, and Biology Tools Swix inspects sample tasks on the live landing page. The guests explain complex task formulations like fastText ML training constraints and biology DNA sequence assembly designed by domain scientists.13:55–21:02 · The hosts pushing back 5/10 Adapting External Benchmarks and the Meta-Framework Vision Alessio questions if adapting dozens of academic benchmarks yields diminishing returns versus first-principles task design. Mike and Alex explain why web scraping for evals is exhausted and why Terminal-Bench serves as a general meta-framework and RL training harness.21:02–26:37 · The hosts pushing back 4/10 Disentangling Model Intelligence from Agent Harnesses Alessio asks how to decouple model intelligence from proprietary agent harnesses. Mike details Terminus as an unopinionated headless baseline, while Alex notes model capability differences far outweigh agent framework tuning.26:38–33:36 · The hosts pushing back 5/10 Roadmap, Cloud Infrastructure, and Economic Evaluation Metrics Swix and Alessio push for multi-dimensional evals incorporating cost, latency, and scaled compute budgets. Mike argues the ultimate eval metric is economic profit and value generated, while Alex describes timeout ablation experiments.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 17:00 Rejecting automated benchmark scraping

Mike strongly rejects the idea that future benchmarks can be scraped automatically from GitHub, arguing that synthetic generation and online scraping fail to produce valid long-horizon tasks.

Hardest push from the hosts ▶ 15:50 Challenging benchmark adapter redundancy

Alessio presses the guests on whether continually adapting external benchmarks like SWE-bench variants creates diminishing returns rather than building unique first-principles tasks.

Biggest teaching moment ▶ 17:30 Explaining the shift to observation-based task generation

Mike educates the hosts on how benchmark design must transition from clear input/output code tasks to deeply specialized container environments modeled on real-world professional workflows.

The host holds their own ▶ 31:55 Probing compute-scaling dynamics in agent tasks

Alessio synthesizes the economic evaluation discussion by distinguishing between merely chaining short tasks together and testing whether large token allocations unlock fundamentally new agent capabilities.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Origins of Terminal-Bench and Frontier Lab Adoption 2311 The hosts welcome Alex and Mike, noting industry adoption by frontier labs. The guests recount the origins of Terminal-Bench from SWE-bench and the surprise mention in the Claude 3.5 Sonnet model card.
Terminal Abstraction and Expanding Beyond Pure Coding 4412 Alessio asks why they bet on terminal CLI over GUI and whether tasks reflect broader computing beyond coding. Mike explains why text-native terminal interfaces outperform GUI navigation, while Alex explains viewing tasks as general container environments.
Task Design, Machine Learning Benchmarks, and Biology Tools 3501 Swix inspects sample tasks on the live landing page. The guests explain complex task formulations like fastText ML training constraints and biology DNA sequence assembly designed by domain scientists.
Adapting External Benchmarks and the Meta-Framework Vision 6625 Alessio questions if adapting dozens of academic benchmarks yields diminishing returns versus first-principles task design. Mike and Alex explain why web scraping for evals is exhausted and why Terminal-Bench serves as a general meta-framework and RL training harness.
Disentangling Model Intelligence from Agent Harnesses 6524 Alessio asks how to decouple model intelligence from proprietary agent harnesses. Mike details Terminus as an unopinionated headless baseline, while Alex notes model capability differences far outweigh agent framework tuning.
Roadmap, Cloud Infrastructure, and Economic Evaluation Metrics 7525 Swix and Alessio push for multi-dimensional evals incorporating cost, latency, and scaled compute budgets. Mike argues the ultimate eval metric is economic profit and value generated, while Alex describes timeout ablation experiments.

Statements from this episode (11)

Assertion Supported
Merrill: Dario Amodei highlighted Terminal-Bench on the Claude model card
“I think one of the really key moments for us was getting onto the Claude IV model card. Being one of two benchmarks that Dario actually mentioned while releasing the model.”
Mike Merrill Oct 18, 2025 ▶ 3:58
Insight
Merrill: AI terminal interfaces will outpace GUI computer use development
“The bet here that we made was that GUI based computer use was going to be much slower to come into its own. Than terminal based use. And there's a few reasons for this. I mean, as I mentioned earlier, text is just the modality that works best with these models…”
Mike Merrill Oct 18, 2025 ▶ 6:29
Assertion Open · timeframe Oct 2026
Merrill: AI models now reliably solve Terminal-Bench's ML training task
“Unfortunately we are getting to the point where models do reliably get this one.”
Mike Merrill Oct 18, 2025 ▶ 11:36
Prediction Not checkable as stated
Shaw: Future agents will run containerized CLI tools behind the scenes
“I don't think the future of all agents is people NPM installing them onto their computers and then typing commands into their terminal. But I do think that behind the scenes, these agents are running in containers and their tools are actually programs that the…”
Alex Shaw Oct 18, 2025 ▶ 14:36
Prediction Not checkable as stated
Merrill: AI evals will shift to observing real jobs over 2-3 years
“I think like the future of evals does look much more like this is like observing people who are doing their real jobs and then translating those real jobs into a format that allows you to evaluate language models and harnesses on them. And it's probably going …”
Mike Merrill Oct 18, 2025 ▶ 18:41
Assertion Contradicted
Merrill: 60% to 70% of SWE-bench Verified tasks come from Django
“If you go look at sweet bench verified, I think like 60, 70% of the tasks in there are from Django.”
Mike Merrill Oct 18, 2025 ▶ 19:06
Disclosure
Shaw: Terminal-Bench has 30 third-party benchmark adapters in development
“I think we have 30 adapters on their way, and we have a couple of users who are building their benchmarks directly in Terminal Bench, and, like, they will use that as a way to distribute it.”
Alex Shaw Oct 18, 2025 ▶ 20:05
Insight
Alex Shaw: Base model selection matters much more than agent frameworks
“And if you look at the difference between the same model and the biggest spread across agent frameworks versus the same agent framework and the biggest spread across models, It's much larger across models, which I think goes to show that like improving models …”
Alex Shaw Oct 18, 2025 ▶ 23:51
Assertion Supported
Alex Shaw: Agent frameworks can impact benchmark scores by up to 15%
“Agent framework seems to make a difference as well, like up to 15% or something pretty significant.”
Alex Shaw Oct 18, 2025 ▶ 24:16
Prediction Not checkable as stated
Merrill: Frontier AI labs will center operations around vertical products
“And with the Claude codes and the codec CLIs and the deep researchers, researchers, you starting to see some evidence that the products are going to be a much more central part of how these frontier labs operate.”
Mike Merrill Oct 18, 2025 ▶ 26:19
Prediction Not checkable as stated
Merrill: AI agents could ace Terminal-Bench with $1,000 compute per task
“Like, I think there's probably a way in which you spend a thousand dollars on tokens for each terminal bench task. And like get them all right just by scaling inference time compute.”
Mike Merrill Oct 18, 2025 ▶ 30:33
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.