MLE-bench
5 statements across 1 episodes · 2 bullish · 1 bearish · 1 people on the record · first statement Oct 19, 2024 by Jesse Hu · said 3 times in 2 episodes since 2024 · across every show →
Mentions by year
brought up most by Jesse Hu (2), Shawn Wang (1)
tap a year for its mentions
2025 1 mention in 1 episode
2024 2 mentions in 1 episode
Everything said about MLE-bench, oldest first
Oct 19, 2024 positive
Oct 19, 2024 negative
Hu: Single MLE-bench evaluation run with OpenAI o1-preview costs $4,000
“Just for one seed, For one run of these things cost 4000 dollars all in with the GPU plus the tokens. And a bulk of the cost was actually the token, so even if you cut the GPU out, it'll still cost you three grand to run on one preview.”
Oct 19, 2024 neutral
Hu: MLE-bench authors found obfuscating competition details did not show overfitting
“They do a lot of checks against overfitting on the Kaggle tasks themselves, and so they do something where they obfuscate some of the details of the Of the competitions, and then they rerun it. And I guess if they were overfitting on the competitions themselve…”
Oct 19, 2024 neutral