Terminal-Bench
also referred to as: terminal bench
includes Terminal Bench 2.0, Terminal Bench 1.0, Terminal Bench 3
13 statements across 3 episodes · 9 bullish · 0 bearish · 4 people on the record · first statement Oct 18, 2025 by Mike Merrill · said 88 times in 9 episodes since 2025 · across every show →
Mentions by year, the whole family
brought up most by Mike Merrill (33), Andy Konwinski (7), Alex Shaw (6), Shawn Wang (5), John Yang (5), Ivan Burazin (1), Alex Krentsel (1), Alessio Fanelli (1)
2026 6 mentions in 4 episodes 2 per episode
- Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
- Exo: Harnesses should see their own code and logs — Alex Krentsel, UC Berekeley / Google Research
- AI Agents Need Computers: 74% MoM Growth, 850K/Day Runs, & New Agent Cloud — Ivan Burazin, Daytona
- Measuring Exponential Trends Rising (in AI) — Joel Becker, METR
- every mention in 2026, scene by scene →
2025 82 mentions in 5 episodes 16 per episode
- Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders
- Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
- [State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang
- [State of Research Funding] Beyond NSF, Slingshots, Open Frontiers — Andy Konwinski, Laude Institute
- [State of AI Papers 2025] Fixing Research with Social Signals, OCR & Implementation — Team AlphaXiv
- every mention in 2025, scene by scene →