Alex Shaw

Founding Member of Technical Staff, Laude Institute · 2 appearances on the record.

computed by AI from the episodes · how this works → · full disclaimer →

scientistengineer@alexgshaw ↗LinkedIn ↗alexgshaw.com ↗

Alex Shaw is the co-creator of Terminal-Bench, an industry-standard benchmark used to evaluate autonomous coding agents in command-line environments. He also developed Harbor, an open-source containerized framework for running agent evaluations and reinforcement learning rollouts.

7statements → 4claims → 2claims resolved → 3.29/5average certainty → 1.86/5average debate potential →

2 supported 0 partly supported 0 contradicted 2 not checkable as stated how the 4 claims stand · each chip opens the sources

1 prediction · 3 assertions · 2 insights · 1 disclosure · every statement was checked. The prediction and assertions are the 4 claims: statements the public record can support or contradict. 2 are resolved, and 2 name no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Alex argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
Alex Shaw: Agent frameworks can impact benchmark scores by up to 15%
“Agent framework seems to make a difference as well, like up to 15% or something pretty significant.”
Alex Shaw Oct 18, 2025 ▶ 24:16 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits

Everything Alex Shaw said on Latent Space that made the record, most notable first. Filter by type, assessment or year in the ledger →

Insight
Alex Shaw: Base model selection matters much more than agent frameworks
“And if you look at the difference between the same model and the biggest spread across agent frameworks versus the same agent framework and the biggest spread across models, It's much larger across models, which I think goes to show that like improving models …”
Alex Shaw Oct 18, 2025 ▶ 23:51 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
Insight
Shaw: Current AI coding benchmarks rely on redundant bespoke test harnesses
“In fact, every single benchmark that gets released at least I would say maybe all coding benchmarks that get released at this point are some form of instruction container tests with some bespoke harness that was coded up That feels very analogous to every othe…”
Alex Shaw Nov 8, 2025 ▶ 12:43 Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders
Assertion Not checkable as stated
Shaw: Harbor's standard task format can express most existing AI evaluations
“Specifically Harbor has a standard task format, which is an iteration of the terminal bench task format which we found to be very flexible and can often express most of the existing evaluations.”
Alex Shaw Nov 8, 2025 ▶ 13:33 Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders
Prediction Not checkable as stated
Shaw: Future agents will run containerized CLI tools behind the scenes
“I don't think the future of all agents is people NPM installing them onto their computers and then typing commands into their terminal. But I do think that behind the scenes, these agents are running in containers and their tools are actually programs that the…”
Alex Shaw Oct 18, 2025 ▶ 14:36 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
Assertion Supported
Alex Shaw: Agent frameworks can impact benchmark scores by up to 15%
“Agent framework seems to make a difference as well, like up to 15% or something pretty significant.”
Alex Shaw Oct 18, 2025 ▶ 24:16 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
Assertion Supported
Shaw: Terminal-Bench accepted 89 of 250 crowdsourced tasks for co-authorship
“We told people if they created three tasks for Terminal Bench, they could be a co-author on the paper that we eventually published, and I think we got maybe 250 task contributions, and 89 of them made it into the benchmark”
Alex Shaw Nov 8, 2025 ▶ 27:01 Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders
Disclosure
Shaw: Terminal-Bench has 30 third-party benchmark adapters in development
“I think we have 30 adapters on their way, and we have a couple of users who are building their benchmarks directly in Terminal Bench, and, like, they will use that as a way to distribute it.”
Alex Shaw Oct 18, 2025 ▶ 20:05 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits

Appearances (2)

EpisodeDateSpeaking time
Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w Nov 8, 2025 8m
Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits Oct 18, 2025 11m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.