general benchmarks

1 statements across 1 episodes · 0 bullish · 1 bearish · 1 people on the record · first statement Mar 13, 2025 by Hamel Husain · across every show →

Everything said about general benchmarks, oldest first

Mar 13, 2025 negative
Insight
Husain: General benchmarks for LLM judges provide very little value
“I put very little value in benchmarks, like general benchmarks. It's has some value, but you know, what you really need to do is like measure it in your domain and see if that alum as a judge is more aligned than like an off the shelf LLM. And what I've found …”
Hamel Husain Mar 13, 2025 ▶ 16:05 [Lightning Pod] Evals: How to Improve AI Consistently — with Hamel Husain and Shreya Shankar
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.