GPQA, every mention
9 scenes, the whole family · ← back to GPQA
tap a year for its mentions
every year anyone Shawn Wang 9Karina Nguyen 3Michelle Pokrass 1Matt Fredrikson 1Alessio Fanelli 1
Verbatim, from the transcripts: the passages where GPQA comes up
AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
- ▶ 30:16 Matt Fredrikson So on the X axis, how capable is the model at, you know, GPQA diamond on, on the Y axis.
Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
- ▶ 19:01 Shawn Wang Like you start out from the general, like MMU and GPQA stuff. 2 times in the scene
[State of MechInterp] SAEs in Production, Circuit Tracing, AI4Science, "Pragmatic" Interp — Goodfire
- ▶ 5:23 unnamed speaker I was thinking, like, GPQ, QA, zero, and, uh, you know, I have eval, 100.
Better Data is All You Need — Ari Morcos, Datology
- ▶ 29:53 Shawn Wang And when you, when you say performance, you mean like GPQA or you mean loss?
- ▶ 1:05:23 Shawn Wang Like, uh, zero on GPQA, a hundred on BrowseConf.
GPT 4.1: The New OpenAI Workhorse
- ▶ 27:45 Michelle Pokrass So Amy, GPQA, stuff like that, you'll see the reasoning models do much better.
The Agent Reasoning Interface: Claude, ChatGPT Canvas, Tasks, Operator — with Karina Nguyen, OpenAI
- ▶ 14:34 Karina Nguyen I would say GPQA was kind of interesting. 3 times in the scene
2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
- ▶ 1:07:09 Shawn Wang And then, um, there was another one this time last year, it was GPQA. 5 times in the scene
Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
- ▶ 38:21 Alessio Fanelli So the April, 15 checkpoint, MMLU on Instruct is like, 86, GPUA, 48, Human Eval, 84, GSMEK, 94, MAT, 57.8.