MMLU, every mention
14 scenes · ← back to MMLU
tap a year for its mentions
every year anyone Shawn Wang 6Alessio Fanelli 3Ethan He 2Will Brown 1Vibhu Sapra 1Micah Hill-Smith 1Erik Schluntz 1Ari Morcos 1
Verbatim, from the transcripts: the passages where MMLU comes up
Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
- ▶ 8:44 Micah Hill-Smith Like, constructed, um, I think never published, like, chain of thought examples, 32 of them in every topic in MLU to run it, to get the score.
- ▶ 19:01 Shawn Wang Like you start out from the general, like MMU and GPQA stuff. 2 times in the scene
Better Data is All You Need — Ari Morcos, Datology
- ▶ 30:04 Ari Morcos So your MML use your arcs, your races, you know, et cetera.
⚡️Multi-Turn RL for Multi-Hour Agents — with Will Brown, Prime Intellect
- ▶ 30:45 Will Brown So, like, the, some questions would be, like, okay, here's some, like, MMLU style question.
Why is everyone cloning Deep Research?
- ▶ 49:17 Shawn Wang Should you care about humanities last exam or not MMIU, but whatever.
The Agent Reasoning Interface: Claude, ChatGPT Canvas, Tasks, Operator — with Karina Nguyen, OpenAI
- ▶ 15:44 Shawn Wang They're just like, my MMLU is 99.
2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
- ▶ 1:06:07 Shawn Wang This time last year, we were, we were still talking about MMLU, a little bit of, there's still like GSM-A-K. 2 times in the scene
The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic
- ▶ 4:41 Erik Schluntz So, you know, we're not focused on sort of these more abstract general benchmarks like math problems or MMLU, but we really care about like finding the things that are really valuable and making sure the models are great at those.
[Paper Club] Upcycling Large Language Models into Mixture of Experts
[Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
- ▶ 21:25 Vibhu Sapra So I guess there is MMLU
Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
- ▶ 46:40 unnamed speaker I always say that here's a divergence between how models are marketed these days versus how people use it, which is when they test MMLU, they'll do like five shots, 25 shots, 50 shots, and no one's providing 50 examples.
Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
- ▶ 38:21 Alessio Fanelli So the April, 15 checkpoint, MMLU on Instruct is like, 86, GPUA, 48, Human Eval, 84, GSMEK, 94, MAT, 57.8. 3 times in the scene
The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
- ▶ 1:22:44 unnamed speaker When you say these kinds of things, like, most legit, um, it, it, obviously there's some, there's vibes eval, or whatever, um, but, like, I feel like a lot, um, people, the, the very common feeling is MMLU is kind of saturated. 3 times in the scene
A Comprehensive Overview of Large Language Models - Latent Space Paper Club
- ▶ 40:27 unnamed speaker So these are some of, I would say, your, uh, single task, uh, evals, and then you've got your multitask evaluation, things like glue, things like MMLU, 2 times in the scene