Alex Shaw
Founding Member of Technical Staff, Laude Institute · 2 appearances on the record.
computed by AI from the episodes · how this works → · full disclaimer →
scientistengineer@alexgshaw ↗LinkedIn ↗alexgshaw.com ↗
Alex Shaw is the co-creator of Terminal-Bench, an industry-standard benchmark used to evaluate autonomous coding agents in command-line environments. He also developed Harbor, an open-source containerized framework for running agent evaluations and reinforcement learning rollouts.
2 supported 0 partly supported 0 contradicted 2 not checkable as stated how the 4 claims stand · each chip opens the sources
1 prediction · 3 assertions · 2 insights · 1 disclosure · every statement was checked. The prediction and assertions are the 4 claims: statements the public record can support or contradict. 2 are resolved, and 2 name no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.
The record, in short
What the tape says about how Alex argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.
Everything Alex Shaw said on Latent Space that made the record, most notable first. Filter by type, assessment or year in the ledger →
Appearances (2)
| Episode | Date | Speaking time |
|---|---|---|
| Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w | Nov 8, 2025 | 8m |
| Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits | Oct 18, 2025 | 11m |