Terminal-Bench

also referred to as: terminal bench

includes Terminal Bench 2.0, Terminal Bench 1.0, Terminal Bench 3

13 statements across 3 episodes · 9 bullish · 0 bearish · 4 people on the record · first statement Oct 18, 2025 by Mike Merrill · said 88 times in 9 episodes since 2025 · across every show →

Mentions by year, the whole family

brought up most by Mike Merrill (33), Andy Konwinski (7), Alex Shaw (6), Shawn Wang (5), John Yang (5), Ivan Burazin (1), Alex Krentsel (1), Alessio Fanelli (1)

tap a year for its mentions
00503100520252026episodesmentions
03520252026episodes it came up in
00102.520520252026episodesmentions per episode

every mention, scene by scene, with the transcript →

Everything said about Terminal-Bench, oldest first

Oct 18, 2025 neutral
Prediction Not checkable as stated
Merrill: AI agents could ace Terminal-Bench with $1,000 compute per task
“Like, I think there's probably a way in which you spend a thousand dollars on tokens for each terminal bench task. And like get them all right just by scaling inference time compute.”
Mike Merrill Oct 18, 2025 ▶ 30:33 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
Oct 18, 2025 positive
Insight
Alex Shaw: Base model selection matters much more than agent frameworks
“And if you look at the difference between the same model and the biggest spread across agent frameworks versus the same agent framework and the biggest spread across models, It's much larger across models, which I think goes to show that like improving models …”
Alex Shaw Oct 18, 2025 ▶ 23:51 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
Oct 18, 2025 neutral
Assertion Supported
Alex Shaw: Agent frameworks can impact benchmark scores by up to 15%
“Agent framework seems to make a difference as well, like up to 15% or something pretty significant.”
Alex Shaw Oct 18, 2025 ▶ 24:16 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
Oct 18, 2025 positive
Prediction Not checkable as stated
Merrill: AI evals will shift to observing real jobs over 2-3 years
“I think like the future of evals does look much more like this is like observing people who are doing their real jobs and then translating those real jobs into a format that allows you to evaluate language models and harnesses on them. And it's probably going …”
Mike Merrill Oct 18, 2025 ▶ 18:41 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
Oct 18, 2025 bullish
Disclosure
Shaw: Terminal-Bench has 30 third-party benchmark adapters in development
“I think we have 30 adapters on their way, and we have a couple of users who are building their benchmarks directly in Terminal Bench, and, like, they will use that as a way to distribute it.”
Alex Shaw Oct 18, 2025 ▶ 20:05 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
Oct 18, 2025 bullish
Insight
Merrill: AI terminal interfaces will outpace GUI computer use development
“The bet here that we made was that GUI based computer use was going to be much slower to come into its own. Than terminal based use. And there's a few reasons for this. I mean, as I mentioned earlier, text is just the modality that works best with these models…”
Mike Merrill Oct 18, 2025 ▶ 6:29 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
Oct 18, 2025 positive
Assertion Supported
Merrill: Dario Amodei highlighted Terminal-Bench on the Claude model card
“I think one of the really key moments for us was getting onto the Claude IV model card. Being one of two benchmarks that Dario actually mentioned while releasing the model.”
Mike Merrill Oct 18, 2025 ▶ 3:58 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
Oct 18, 2025 bullish
Assertion Open · timeframe Oct 2026
Merrill: AI models now reliably solve Terminal-Bench's ML training task
“Unfortunately we are getting to the point where models do reliably get this one.”
Mike Merrill Oct 18, 2025 ▶ 11:36 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
Nov 8, 2025 positive
Assertion Supported
Shaw: Terminal-Bench accepted 89 of 250 crowdsourced tasks for co-authorship
“We told people if they created three tasks for Terminal Bench, they could be a co-author on the paper that we eventually published, and I think we got maybe 250 task contributions, and 89 of them made it into the benchmark”
Alex Shaw Nov 8, 2025 ▶ 27:01 Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders
Nov 8, 2025 neutral
Disclosure
Konwinski: Terminal-Bench is the largest LAWD Slingshot grant investment to date
“Terminal Bench is the largest investment we've made in any one of the slingshots so far.”
Andy Konwinski Nov 8, 2025 ▶ 26:26 Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders
Nov 8, 2025
Assertion Not checkable as stated
Merrill: Terminal-Bench is used by every frontier AI lab
“It's used by all Frontier Labs in some capacity.”
Mike Merrill Nov 8, 2025 ▶ 2:18 Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders
Nov 8, 2025 positive
Assertion Not checkable as stated
Shaw: Harbor's standard task format can express most existing AI evaluations
“Specifically Harbor has a standard task format, which is an iteration of the terminal bench task format which we found to be very flexible and can often express most of the existing evaluations.”
Alex Shaw Nov 8, 2025 ▶ 13:33 Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders
Dec 31, 2025 positive
Insight
Yang: TerminalBench enables more creative eval design than SWE-bench
“As we bench, you're confined in some sense to the domain of issues and PRs that already exist which I think has its benefits of being close to reality and natural, but I think with Terminal Bench, there's a lot of creativity that you can infuse into that”
John Yang Dec 31, 2025 ▶ 11:14 [State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.