Mike Merrill
Co-Creator, Terminal-Bench · 2 appearances on the record.
computed by AI from the episodes · how this works → · full disclaimer →
Mike Merrill is the co-creator of Terminal-Bench, a coding agent benchmark. He developed Terminal-Bench 2.0 to evaluate frontier AI models on verified terminal tasks.
2 supported 0 partly supported 1 contradicted 1 not yet assessed 6 not checkable as stated how the 10 claims stand · each chip opens the sources
3 predictions · 7 assertions · 1 opinion · 2 insights · every statement was checked. The predictions and assertions are the 10 claims: statements the public record can support or contradict. 3 are resolved, 1 is not yet assessed, and 6 name no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.
The record, in short
What the tape says about how Mike argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.
Their most notable supported claim
Expressed certainty vs assessment result
weighted support: a fully supported claim counts one, a partly supported claim counts half. Each filled bar is clickable and opens exactly those claims; "none yet" means nothing said at that certainty level has resolved yet
Everything Mike Merrill said on Latent Space that made the record, most notable first. Filter by type, assessment or year in the ledger →
Appearances (2)
| Episode | Date | Speaking time |
|---|---|---|
| Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w | Nov 8, 2025 | 8m |
| Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits | Oct 18, 2025 | 13m |