MMLU, every mention

14 scenes · ← back to MMLU

tap a year for its mentions
0084158202420252026episodesmentions
048202420252026episodes it came up in
001.5438202420252026episodesmentions per episode

every year anyone Shawn Wang 6Alessio Fanelli 3Ethan He 2Will Brown 1Vibhu Sapra 1Micah Hill-Smith 1Erik Schluntz 1Ari Morcos 1

Verbatim, from the transcripts: the passages where MMLU comes up

loading…

Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith Jan 9, 2026 · 3 mentions

  • ▶ 8:44 Micah Hill-Smith Like, constructed, um, I think never published, like, chain of thought examples, 32 of them in every topic in MLU to run it, to get the score.
  • ▶ 19:01 Shawn Wang Like you start out from the general, like MMU and GPQA stuff. 2 times in the scene

Better Data is All You Need — Ari Morcos, Datology Aug 29, 2025 · 1 mention

⚡️Multi-Turn RL for Multi-Hour Agents — with Will Brown, Prime Intellect May 23, 2025 · 1 mention

  • ▶ 30:45 Will Brown So, like, the, some questions would be, like, okay, here's some, like, MMLU style question.

Why is everyone cloning Deep Research? Feb 18, 2025 · 1 mention

The Agent Reasoning Interface: Claude, ChatGPT Canvas, Tasks, Operator — with Karina Nguyen, OpenAI Feb 1, 2025 · 1 mention

2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents Jan 1, 2025 · 2 mentions

  • ▶ 1:06:07 Shawn Wang This time last year, we were, we were still talking about MMLU, a little bit of, there's still like GSM-A-K. 2 times in the scene

The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic Nov 28, 2024 · 1 mention

  • ▶ 4:41 Erik Schluntz So, you know, we're not focused on sort of these more abstract general benchmarks like math problems or MMLU, but we really care about like finding the things that are really valuable and making sure the models are great at those.

[Paper Club] Upcycling Large Language Models into Mixture of Experts Oct 29, 2024 · 2 mentions

  • ▶ 35:46 Ethan He So, validation loss, 1.6, and MMLU is nine. 2 times in the scene

[Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz Oct 13, 2024 · 1 mention

Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind Aug 28, 2024 · 1 mention

  • ▶ 46:40 unnamed speaker I always say that here's a divergence between how models are marketed these days versus how people use it, which is when they test MMLU, they'll do like five shots, 25 shots, 50 shots, and no one's providing 50 examples.

Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI Jul 23, 2024 · 3 mentions

  • ▶ 38:21 Alessio Fanelli So the April, 15 checkpoint, MMLU on Instruct is like, 86, GPUA, 48, Human Eval, 84, GSMEK, 94, MAT, 57.8. 3 times in the scene

The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka Jul 5, 2024 · 3 mentions

  • ▶ 1:22:44 unnamed speaker When you say these kinds of things, like, most legit, um, it, it, obviously there's some, there's vibes eval, or whatever, um, but, like, I feel like a lot, um, people, the, the very common feeling is MMLU is kind of saturated. 3 times in the scene

A Comprehensive Overview of Large Language Models - Latent Space Paper Club Mar 15, 2024 · 2 mentions

  • ▶ 40:27 unnamed speaker So these are some of, I would say, your, uh, single task, uh, evals, and then you've got your multitask evaluation, things like glue, things like MMLU, 2 times in the scene
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.