Benchmark

2 statements across 2 episodes · 2 bullish · 0 bearish · 2 people on the record · first statement Aug 28, 2024 by Nicholas Carlini · said 2 times in 2 episodes since 2024 · across every show →

Mentions by year

brought up most by Shawn Wang (1), Andrew Feldman (1)

tap a year for its mentions
00111120242025episodesmentions
01120242025episodes it came up in
000.50.51120242025episodesmentions per episode

every mention, scene by scene, with the transcript →

Everything said about Benchmark, oldest first

Aug 28, 2024 positive
Insight
Carlini: Unpopular AI benchmarks protect against model contamination and overfitting
“And by having a benchmark that is not very popular, you can be relatively certain that no one has tried to optimize their model for your benchmark.”
Nicholas Carlini Aug 28, 2024 ▶ 40:31 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Nov 8, 2025 positive
Insight
Merrill: Launching benchmarks is the best way to steer AI development
“And so, if you have an opinion about how AI should work, what's the best way to get other people to work on it? And that's by launching a benchmark.”
Mike Merrill Nov 8, 2025 ▶ 1:35 Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.