People, every show

Alex Shaw

Founding Member of Technical Staff, Laude Institute. On 1 show, 2 appearances. The Shows tab opens the full record on each.

scientistengineer@alexgshaw ↗LinkedIn ↗alexgshaw.com ↗

Alex Shaw is the co-creator of Terminal-Bench, an industry-standard benchmark used to evaluate autonomous coding agents in command-line environments. He also developed Harbor, an open-source containerized framework for running agent evaluations and reinforcement learning rollouts.

1shows
2appearances
7statements
2resolved
2supported
0contradicted
100%fully supported

Everything Alex Shaw said on any show that made the record, most notable first. Each card names its show and opens the statement there.

Alex Shaw: Base model selection matters much more than agent frameworks
“And if you look at the difference between the same model and the biggest spread across agent frameworks versus the same agent framework and the biggest spread across models, It's much larger across models, which I think goes to show that like improving models …”
Alex Shaw Oct 18, 2025 ▶ 23:51 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
Shaw: Current AI coding benchmarks rely on redundant bespoke test harnesses
“In fact, every single benchmark that gets released at least I would say maybe all coding benchmarks that get released at this point are some form of instruction container tests with some bespoke harness that was coded up That feels very analogous to every othe…”
Alex Shaw Nov 8, 2025 ▶ 12:43 Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders
LATENT SPACE Assertion Not checkable as stated
Shaw: Harbor's standard task format can express most existing AI evaluations
“Specifically Harbor has a standard task format, which is an iteration of the terminal bench task format which we found to be very flexible and can often express most of the existing evaluations.”
Alex Shaw Nov 8, 2025 ▶ 13:33 Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders
LATENT SPACE Prediction Not checkable as stated
Shaw: Future agents will run containerized CLI tools behind the scenes
“I don't think the future of all agents is people NPM installing them onto their computers and then typing commands into their terminal. But I do think that behind the scenes, these agents are running in containers and their tools are actually programs that the…”
Alex Shaw Oct 18, 2025 ▶ 14:36 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
LATENT SPACE Assertion Supported
Alex Shaw: Agent frameworks can impact benchmark scores by up to 15%
“Agent framework seems to make a difference as well, like up to 15% or something pretty significant.”
Alex Shaw Oct 18, 2025 ▶ 24:16 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
LATENT SPACE Assertion Supported
Shaw: Terminal-Bench accepted 89 of 250 crowdsourced tasks for co-authorship
“We told people if they created three tasks for Terminal Bench, they could be a co-author on the paper that we eventually published, and I think we got maybe 250 task contributions, and 89 of them made it into the benchmark”
Alex Shaw Nov 8, 2025 ▶ 27:01 Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders
LATENT SPACE Disclosure
Shaw: Terminal-Bench has 30 third-party benchmark adapters in development
“I think we have 30 adapters on their way, and we have a couple of users who are building their benchmarks directly in Terminal Bench, and, like, they will use that as a way to distribute it.”
Alex Shaw Oct 18, 2025 ▶ 20:05 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits

One line per show, most statements first. The link opens Alex's full record on that show: the calibration, argument clarity, speaking style and every statement made there.

ShowRole thereEpsStatementsRecord
LATENT SPACELEDGER Founding Member of Technical Staff, Laude Institute 2 7 100% 2/2 full record on Latent Space →
Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.