AI Benchmark Evaluations
topic on 1 show · 1 statements across 1 episodes
1 statements about AI Benchmark Evaluations, every show
Krieger: AI benchmark evals do not indicate real-world model performance
“Evals are really useful for hill climbing and for internal research, but they don't tell the story of like, is the model going to be excellent at what it needs to be excellent or deployed for, or even if it is excellent at that thing, is it only excellent at t…”