Static Benchmarks
topic on 2 shows · 2 statements across 2 episodes
2 statements about Static Benchmarks, every show
Angelopoulos: Chatbot Arena is immune to model overfitting by design
“Static benchmarks overfit. Why? It's because as Jan said earlier, you're giving the student the same test over and over. You have a model, you test it, you know, you look at whether or not it's improved on a static data set. Then you find another model, you te…”
Angelopoulos: Static benchmarks are intrinsically unable to evaluate generative models
“Static benchmarks are intrinsically, to some extent, unable to measure generative model performance. And the reason is because you cannot Pre-annotate all the outputs of a generative model. You change the model. It's like the distribution of your data is chang…”