benchmarks

6 statements across 6 episodes · 1 bullish · 2 bearish · 6 people on the record · first statement Mar 14, 2024 by Mikey Shulman · across every show →

Everything said about benchmarks, oldest first

Mar 14, 2024 negative
Opinion
Shulman: AI evaluation benchmarks are far worse in audio than text
“As flawed as these benchmarks are in text, they're way worse in audio.”
Mikey Shulman Mar 14, 2024 ▶ 55:42 Making Transformers Sing - with Mikey Shulman of Suno
Jun 11, 2024 negative
Assertion Not checkable as stated
Conover: AI model developers are absolutely overfitting to public evaluation benchmarks
“And I think the work around over, you know, overfitting on the test, I think is like that. 100% is happening.”
Mike Conover Jun 11, 2024 ▶ 58:21 How AI is Eating Finance - with Mike Conover of Brightwave
Jun 25, 2024
Disclosure
Albrecht: Imbue reproduced 500-1,000 examples per dataset to stop eval contamination
“Let's just reproduce, you know, 500 to a thousand examples for every single one of these data sets ourselves and just make sure that this data is definitely not in the, you know, the training set. So we did that and then we're able to like now be confident abo…”
Josh Albrecht Jun 25, 2024 ▶ 1:00:11 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Aug 28, 2024 positive
Insight
Carlini: Users should build personalized AI benchmarks instead of relying on public leaderboards
“The argument that I tried to lay out in this post is that more people should make benchmarks that are tailored to them.”
Nicholas Carlini Aug 28, 2024 ▶ 39:10 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Jan 1, 2025
Insight
Swyx: Frontier AI labs distinguish themselves by adopting new benchmarks
“The labs that are not that frontier will keep measuring themselves on last year's benchmarks. And then the labs that are actually frontier will tell you about benchmarks you've never heard of.”
Shawn Wang Jan 1, 2025 ▶ 1:08:52 2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
Jul 29, 2025
Insight
Fortuna: Healthcare AI evaluation must measure worst-case failures over best-of-N
“I think one of the problems is, like, the benchmarks, they always report, like, the best event. So they report, like, you know, what's the best metric they got out of 64 attempts? But in healthcare, we're more interested in, like, what's the worst event, right…”
Brendan Fortuna Jul 29, 2025 ▶ 19:09 ⚡️Using RFT to Build Clinical Superintelligence
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.