Harbor

2 statements across 1 episodes · 1 bullish · 1 bearish · 1 people on the record · first statement Nov 8, 2025 by Alex Shaw · said 29 times in 4 episodes since 2025 · across every show →

Mentions by year

brought up most by Alex Shaw (7), Shawn Wang (3), Mike Merrill (3), Alex Krentsel (3)

tap a year for its mentions
0013125220252026episodesmentions
01220252026episodes it came up in
007.5115220252026episodesmentions per episode

every mention, scene by scene, with the transcript →

Everything said about Harbor, oldest first

Nov 8, 2025 negative
Insight
Shaw: Current AI coding benchmarks rely on redundant bespoke test harnesses
“In fact, every single benchmark that gets released at least I would say maybe all coding benchmarks that get released at this point are some form of instruction container tests with some bespoke harness that was coded up That feels very analogous to every othe…”
Alex Shaw Nov 8, 2025 ▶ 12:43 Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders
Nov 8, 2025 positive
Assertion Not checkable as stated
Shaw: Harbor's standard task format can express most existing AI evaluations
“Specifically Harbor has a standard task format, which is an iteration of the terminal bench task format which we found to be very flexible and can often express most of the existing evaluations.”
Alex Shaw Nov 8, 2025 ▶ 13:33 Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.