Mike Merrill

Co-Creator, Terminal-Bench · 2 appearances on the record.

computed by AI from the episodes · how this works → · full disclaimer →

Mike Merrill is the co-creator of Terminal-Bench, a coding agent benchmark. He developed Terminal-Bench 2.0 to evaluate frontier AI models on verified terminal tasks.

13statements → 10claims → 3claims resolved → 67%fully supported → 3.46/5average certainty → 2.15/5average debate potential →

2 supported 0 partly supported 1 contradicted 1 not yet assessed 6 not checkable as stated how the 10 claims stand · each chip opens the sources

3 predictions · 7 assertions · 1 opinion · 2 insights · every statement was checked. The predictions and assertions are the 10 claims: statements the public record can support or contradict. 3 are resolved, 1 is not yet assessed, and 6 name no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Mike argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
Merrill: Terminus-2 harness yields highest agent performance on Terminal-Bench 2.0
“Terminus-II, which is our harness for measuring language models, is the highest performing agent for many of these models.”
Mike Merrill Nov 8, 2025 ▶ 6:10 Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders

Their most notable contradicted claim

Assertion Contradicted
Merrill: 60% to 70% of SWE-bench Verified tasks come from Django
“If you go look at sweet bench verified, I think like 60, 70% of the tasks in there are from Django.”
Mike Merrill Oct 18, 2025 ▶ 19:06 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits

Expressed certainty vs assessment result

none yet certainty 1
none yet certainty 2
0% certainty 3
100% certainty 4
none yet certainty 5

weighted support: a fully supported claim counts one, a partly supported claim counts half. Each filled bar is clickable and opens exactly those claims; "none yet" means nothing said at that certainty level has resolved yet

Everything Mike Merrill said on Latent Space that made the record, most notable first. Filter by type, assessment or year in the ledger →

Insight
Merrill: AI terminal interfaces will outpace GUI computer use development
“The bet here that we made was that GUI based computer use was going to be much slower to come into its own. Than terminal based use. And there's a few reasons for this. I mean, as I mentioned earlier, text is just the modality that works best with these models…”
Mike Merrill Oct 18, 2025 ▶ 6:29 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
Prediction Not checkable as stated
Merrill: AI evals will shift to observing real jobs over 2-3 years
“I think like the future of evals does look much more like this is like observing people who are doing their real jobs and then translating those real jobs into a format that allows you to evaluate language models and harnesses on them. And it's probably going …”
Mike Merrill Oct 18, 2025 ▶ 18:41 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
Prediction Not checkable as stated
Merrill: Frontier AI labs will center operations around vertical products
“And with the Claude codes and the codec CLIs and the deep researchers, researchers, you starting to see some evidence that the products are going to be a much more central part of how these frontier labs operate.”
Mike Merrill Oct 18, 2025 ▶ 26:19 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
Assertion Supported
Merrill: Terminus-2 harness yields highest agent performance on Terminal-Bench 2.0
“Terminus-II, which is our harness for measuring language models, is the highest performing agent for many of these models.”
Mike Merrill Nov 8, 2025 ▶ 6:10 Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders
Opinion
Merrill: Text-based agents are currently the frontier of performance
“We like text-based agents. We think right now they're the frontier of agentic performance. This very well may change.”
Mike Merrill Nov 8, 2025 ▶ 33:58 Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders
Prediction Not checkable as stated
Merrill: AI agents could ace Terminal-Bench with $1,000 compute per task
“Like, I think there's probably a way in which you spend a thousand dollars on tokens for each terminal bench task. And like get them all right just by scaling inference time compute.”
Mike Merrill Oct 18, 2025 ▶ 30:33 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
Insight
Merrill: Launching benchmarks is the best way to steer AI development
“And so, if you have an opinion about how AI should work, what's the best way to get other people to work on it? And that's by launching a benchmark.”
Mike Merrill Nov 8, 2025 ▶ 1:35 Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders
Assertion Not checkable as stated
Merrill: Terminal-Bench is used by every frontier AI lab
“It's used by all Frontier Labs in some capacity.”
Mike Merrill Nov 8, 2025 ▶ 2:18 Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders
Assertion Supported
Merrill: Dario Amodei highlighted Terminal-Bench on the Claude model card
“I think one of the really key moments for us was getting onto the Claude IV model card. Being one of two benchmarks that Dario actually mentioned while releasing the model.”
Mike Merrill Oct 18, 2025 ▶ 3:58 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
Assertion Open · timeframe Oct 2026
Merrill: AI models now reliably solve Terminal-Bench's ML training task
“Unfortunately we are getting to the point where models do reliably get this one.”
Mike Merrill Oct 18, 2025 ▶ 11:36 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
Assertion Contradicted
Merrill: 60% to 70% of SWE-bench Verified tasks come from Django
“If you go look at sweet bench verified, I think like 60, 70% of the tasks in there are from Django.”
Mike Merrill Oct 18, 2025 ▶ 19:06 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
Assertion Not checkable as stated
Terminal-Bench 2.0 tasks are based on real paid workflows
“Every task in the benchmark is based off of real work that people do out in the world and are paid for.”
Mike Merrill Nov 8, 2025 ▶ 4:20 Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders
Assertion Not checkable as stated
Merrill: Terminal-Bench team spent 300+ hours auditing 89 benchmark tasks
“So I estimate we've spent at least 300 hours digging through each one of these 89 tasks, trying to figure out whether or not they're broken, whether or not they're some way to cheat, whether or not they're reproducible, and really obsessing over quality while …”
Mike Merrill Nov 8, 2025 ▶ 4:32 Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders

Appearances (2)

EpisodeDateSpeaking time
Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w Nov 8, 2025 8m
Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits Oct 18, 2025 13m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.