Disclosure certainty 3/5 debate potential 1/5

Shaw: Terminal-Bench has 30 third-party benchmark adapters in development

Alex Shaw · Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits · Oct 18, 2025 · at 20:05

Terminal-Bench co-creator Alex Shaw describes community adoption and the expansion of Terminal-Bench as an evaluation meta-framework.

0:00 / 0:09exact quote · 9.0s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“I think we have 30 adapters on their way, and we have a couple of users who are building their benchmarks directly in Terminal Bench, and, like, they will use that as a way to distribute it.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Alex Shaw

Insight
Alex Shaw: Base model selection matters much more than agent frameworks
“And if you look at the difference between the same model and the biggest spread across agent frameworks versus the same agent framework and the biggest spread across models, It's much larger across models, which I think goes to show that like improving models …”
Alex Shaw Oct 18, 2025 ▶ 23:51 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
Insight
Shaw: Current AI coding benchmarks rely on redundant bespoke test harnesses
“In fact, every single benchmark that gets released at least I would say maybe all coding benchmarks that get released at this point are some form of instruction container tests with some bespoke harness that was coded up That feels very analogous to every othe…”
Alex Shaw Nov 8, 2025 ▶ 12:43 Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders
Assertion Not checkable as stated
Shaw: Harbor's standard task format can express most existing AI evaluations
“Specifically Harbor has a standard task format, which is an iteration of the terminal bench task format which we found to be very flexible and can often express most of the existing evaluations.”
Alex Shaw Nov 8, 2025 ▶ 13:33 Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders
Prediction Not checkable as stated
Shaw: Future agents will run containerized CLI tools behind the scenes
“I don't think the future of all agents is people NPM installing them onto their computers and then typing commands into their terminal. But I do think that behind the scenes, these agents are running in containers and their tools are actually programs that the…”
Alex Shaw Oct 18, 2025 ▶ 14:36 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
Assertion Supported
Alex Shaw: Agent frameworks can impact benchmark scores by up to 15%
“Agent framework seems to make a difference as well, like up to 15% or something pretty significant.”
Alex Shaw Oct 18, 2025 ▶ 24:16 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
Assertion Supported
Shaw: Terminal-Bench accepted 89 of 250 crowdsourced tasks for co-authorship
“We told people if they created three tasks for Terminal Bench, they could be a co-author on the paper that we eventually published, and I think we got maybe 250 task contributions, and 89 of them made it into the benchmark”
Alex Shaw Nov 8, 2025 ▶ 27:01 Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.