AI benchmark evaluations
1 statements across 1 episodes · 0 bullish · 0 bearish · 1 people on the record · first statement Mar 3, 2025 by Mike Krieger · across every show →
Everything said about AI benchmark evaluations, oldest first
Mar 3, 2025 neutral
Krieger: AI benchmark evals do not indicate real-world model performance
“Evals are really useful for hill climbing and for internal research, but they don't tell the story of like, is the model going to be excellent at what it needs to be excellent or deployed for, or even if it is excellent at that thing, is it only excellent at t…”