Oct 18, 2025 · 35m · latent-space
Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
Alex and Mike Merrill join hosts Alessio and Swix to discuss Terminal-Bench, an open-source framework evaluating autonomous AI agents on diverse command-line tasks. They explore the advantages of terminal abstractions over GUIs, the benchmark's adoption across frontier AI labs, and methods for measuring agent intelligence and economic value.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Mike strongly rejects the idea that future benchmarks can be scraped automatically from GitHub, arguing that synthetic generation and online scraping fail to produce valid long-horizon tasks.
Hardest push from the hosts ▶ 15:50 Challenging benchmark adapter redundancyAlessio presses the guests on whether continually adapting external benchmarks like SWE-bench variants creates diminishing returns rather than building unique first-principles tasks.
Biggest teaching moment ▶ 17:30 Explaining the shift to observation-based task generationMike educates the hosts on how benchmark design must transition from clear input/output code tasks to deeply specialized container environments modeled on real-world professional workflows.
The host holds their own ▶ 31:55 Probing compute-scaling dynamics in agent tasksAlessio synthesizes the economic evaluation discussion by distinguishing between merely chaining short tasks together and testing whether large token allocations unlock fundamentally new agent capabilities.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Origins of Terminal-Bench and Frontier Lab Adoption | 2 | 3 | 1 | 1 | The hosts welcome Alex and Mike, noting industry adoption by frontier labs. The guests recount the origins of Terminal-Bench from SWE-bench and the surprise mention in the Claude 3.5 Sonnet model card. | |
| Terminal Abstraction and Expanding Beyond Pure Coding | 4 | 4 | 1 | 2 | Alessio asks why they bet on terminal CLI over GUI and whether tasks reflect broader computing beyond coding. Mike explains why text-native terminal interfaces outperform GUI navigation, while Alex explains viewing tasks as general container environments. | |
| Task Design, Machine Learning Benchmarks, and Biology Tools | 3 | 5 | 0 | 1 | Swix inspects sample tasks on the live landing page. The guests explain complex task formulations like fastText ML training constraints and biology DNA sequence assembly designed by domain scientists. | |
| Adapting External Benchmarks and the Meta-Framework Vision | 6 | 6 | 2 | 5 | Alessio questions if adapting dozens of academic benchmarks yields diminishing returns versus first-principles task design. Mike and Alex explain why web scraping for evals is exhausted and why Terminal-Bench serves as a general meta-framework and RL training harness. | |
| Disentangling Model Intelligence from Agent Harnesses | 6 | 5 | 2 | 4 | Alessio asks how to decouple model intelligence from proprietary agent harnesses. Mike details Terminus as an unopinionated headless baseline, while Alex notes model capability differences far outweigh agent framework tuning. | |
| Roadmap, Cloud Infrastructure, and Economic Evaluation Metrics | 7 | 5 | 2 | 5 | Swix and Alessio push for multi-dimensional evals incorporating cost, latency, and scaled compute budgets. Mike argues the ultimate eval metric is economic profit and value generated, while Alex describes timeout ablation experiments. |