general benchmarks

1 statements across 1 episodes · 0 bullish · 0 bearish · 1 people on the record · first statement Sep 25, 2025 by Hamel Husain · across every show →

Everything said about general benchmarks, oldest first

Sep 25, 2025 neutral
Insight
Husain: General LLM benchmarks do not correlate with product-specific evals
“Up until now, a lot of the big labs understandably focused on general benchmarks, like MMLU score, human eval, things like that, which are very important for foundation models. And, you know, those not very related to product specific evals, like the ones we t…”
Hamel Husain Sep 25, 2025 ▶ 1:21:35 Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.