MMLU

6 statements across 6 episodes · 1 bullish · 2 bearish · 6 people on the record · first statement Sep 20, 2024 by Sander Schulhoff · said 22 times in 13 episodes since 2024 · across every show →

Mentions by year

brought up most by Shawn Wang (6), Alessio Fanelli (3), Ethan He (2), Will Brown (1), Vibhu Sapra (1), Micah Hill-Smith (1), Erik Schluntz (1), Ari Morcos (1)

tap a year for its mentions
0084158202420252026episodesmentions
048202420252026episodes it came up in
001.5438202420252026episodesmentions per episode
2026 3 mentions in 1 episode
2025 6 mentions in 5 episodes 1 per episode
2024 13 mentions in 7 episodes 2 per episode

every mention, scene by scene, with the transcript →

Everything said about MMLU, oldest first

Sep 20, 2024 negative
Opinion
Schulhoff: Role Prompting Does Not Improve Accuracy on Modern LLMs
“For accuracy-based tasks, like MMLU, you're trying to solve a math problem, and maybe you tell the AI that it's a math professor, and you expect it to have improved performance. I really don't think that works. I'm quite certain that doesn't work on more moder…”
Sander Schulhoff Sep 20, 2024 ▶ 17:08 The Ultimate Guide to Prompting - with Sander Schulhoff from LearnPrompting.org
Oct 29, 2024 positive
Assertion Supported
He: Upcycling a 15B model on 1T tokens yielded 4% MMLU gain
“On other scaling experiments, we tried on 15 B models upcycling and applied on one trillion tokens and achieved roughly about five percent improvement in terms of the validation loss and four percent improvement on MMLU.”
Ethan He Oct 29, 2024 ▶ 21:12 [Paper Club] Upcycling Large Language Models into Mixture of Experts
Dec 24, 2024
Assertion Supported
Synthetic college text boosts MMLU; middle school text boosts OpenBookQA
“College textbooks are really good for MLU or middle school textbooks are good for benchmarks like open book, UA and Pico.”
Loubna Ben Allal Dec 24, 2024 ▶ 7:44 Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
Jan 26, 2025
Insight
Beauchamp: Real AI progress is performance per dollar, not raw benchmarks
“For us, it doesn't make sense to think of AI as just the absolute performance. So if you look at like the MMLU score or the, you know, any of these benchmarks that people like to look at, If you just get that score, it doesn't really tell, tell you anything. C…”
William Beauchamp Jan 26, 2025 ▶ 25:24 Outlasting Noam Shazeer, Crowdsourcing Chai AI w/ 1.4m DAU — with William Beauchamp, Chai Research
Oct 20, 2025 neutral
Assertion Supported
Bakouch: SmolLM 2 scored random on MMLU until 6.5 trillion tokens
“And this is basically until like 6.5 trillion of tokens, which is a lot, to be honest. Until this amount of token, the MMLU in the QA format, meaning that the model have to select which answer is, the model have to output, for example, the right answer is A, o…”
Elie Bakouch Oct 20, 2025 ▶ 43:42 ⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
Jan 9, 2026 negative
Assertion Contradicted
Hill-Smith: Google used unpublished 32-shot CoT to claim Gemini beat GPT-4
“Back when I'm Googled a Gemini one when I ultra and needed a number that would say it was better than GPT four. And Like, constructed I think never published, like, chain of thought examples, 32 of them in every topic in MLU to run it, to get the score.”
Micah Hill-Smith Jan 9, 2026 ▶ 8:36 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.