MMLU

product on 7 shows · 11 statements across 11 episodes · said 40 times in 24 episodes since 2024

Latent Space 22 No Priors 5 the a16z Podcast 3 Big Technology 2 TBPN 2 the Y Combinator Startup Podcast 1 Lenny's Podcast

Mentions by year, every show

tap a year for its mentions
001052010202420252026episodesmentions
0510202420252026episodes it came up in
001.55310202420252026episodesmentions per episode

Latent Space 22No Priors 5the a16z Podcast 3Big Technology 2TBPN 2the Y Combinator Startup Podcast 1

2026 3 mentions in 1 episode
2025 13 mentions in 10 episodes 1 per episode
2024 19 mentions in 10 episodes 2 per episode

every mention on every show, scene by scene, with the transcript →

11 statements about MMLU, every show

LATENT SPACE Assertion Contradicted
Hill-Smith: Google used unpublished 32-shot CoT to claim Gemini beat GPT-4
“Back when I'm Googled a Gemini one when I ultra and needed a number that would say it was better than GPT four. And Like, constructed I think never published, like, chain of thought examples, 32 of them in every topic in MLU to run it, to get the score.”
Micah Hill-Smith Jan 9, 2026 ▶ 8:36 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
LATENT SPACE Assertion Supported
Bakouch: SmolLM 2 scored random on MMLU until 6.5 trillion tokens
“And this is basically until like 6.5 trillion of tokens, which is a lot, to be honest. Until this amount of token, the MMLU in the QA format, meaning that the model have to select which answer is, the model have to output, for example, the right answer is A, o…”
Elie Bakouch Oct 20, 2025 ▶ 43:42 ⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
Husain: General LLM benchmarks do not correlate with product-specific evals
“Up until now, a lot of the big labs understandably focused on general benchmarks, like MMLU score, human eval, things like that, which are very important for foundation models. And, you know, those not very related to product specific evals, like the ones we t…”
Hamel Husain Sep 25, 2025 ▶ 1:21:35 Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar
a16z Prediction Not checkable as stated
Midha: Real-time testing will replace static AI benchmarks like MMLU
“While benchmarks like MMLU and the idea of these static exams were useful three years ago, the future is about real-time evaluation, real-time systems, real-time testing in the wild.”
Anjney Midha May 29, 2025 ▶ 0:43 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Beauchamp: Real AI progress is performance per dollar, not raw benchmarks
“For us, it doesn't make sense to think of AI as just the absolute performance. So if you look at like the MMLU score or the, you know, any of these benchmarks that people like to look at, If you just get that score, it doesn't really tell, tell you anything. C…”
William Beauchamp Jan 26, 2025 ▶ 25:24 Outlasting Noam Shazeer, Crowdsourcing Chai AI w/ 1.4m DAU — with William Beauchamp, Chai Research
LATENT SPACE Assertion Supported
Synthetic college text boosts MMLU; middle school text boosts OpenBookQA
“College textbooks are really good for MLU or middle school textbooks are good for benchmarks like open book, UA and Pico.”
Loubna Ben Allal Dec 24, 2024 ▶ 7:44 Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
LATENT SPACE Assertion Supported
He: Upcycling a 15B model on 1T tokens yielded 4% MMLU gain
“On other scaling experiments, we tried on 15 B models upcycling and applied on one trillion tokens and achieved roughly about five percent improvement in terms of the validation loss and four percent improvement on MMLU.”
Ethan He Oct 29, 2024 ▶ 21:12 [Paper Club] Upcycling Large Language Models into Mixture of Experts
Schulhoff: Role Prompting Does Not Improve Accuracy on Modern LLMs
“For accuracy-based tasks, like MMLU, you're trying to solve a math problem, and maybe you tell the AI that it's a math professor, and you expect it to have improved performance. I really don't think that works. I'm quite certain that doesn't work on more moder…”
Sander Schulhoff Sep 20, 2024 ▶ 17:08 The Ultimate Guide to Prompting - with Sander Schulhoff from LearnPrompting.org
NO PRIORS Assertion Supported
Davis: Compound AI approach yielded a 3% MMLU performance bump
“The MMLU performance bump was about three percent. And to put that in perspective, the gap between some of the previous best models is often less than one percent. Between, for example, Gemini, 1.5, and Lama 3.1, and things like that. So, actually, 2.8% or thr…”
Jared Quincy Davis Aug 22, 2024 ▶ 39:55 No Priors Ep. 77 | With Foundry CEO and Founder Jared Quincy Davis
Standard AI benchmarks like MMLU are becoming saturated and inadequate
“I mean, like, people will come out and say, here's what we got on MMLU and so on, but they're getting saturated, and they're not, often not that great to begin with, so I, I'm more eager to see what it feels like to talk to one of these things than learn what …”
Dwarkesh Patel May 15, 2024 ▶ 5:06 AI Scaling, Alignment, and the Path to Superintelligence — With Dwarkesh Patel
a16z Assertion Supported
Ghodsi: MMLU benchmark scores are inflated due to web data contamination
“MMLU is just a multi-choice question that's on the web. Ask a question. Here's, is the answer A, B, C, D, and then it says what the right answer is. And it's on the web. You can deliberately train on it and create an LLM that crushes it on that. Okay. Or you c…”
Ali Ghodsi Sep 25, 2023 ▶ 18:15 AI Food Fights in the Enterprise with Databricks' Ali Ghodsi

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.