Humanity's Last Exam, every mention
12 scenes · ← back to Humanity's Last Exam
tap a year for its mentions
every year anyone Jordi Hays 2Mike Knoop 1Mark Chen 1John Coogan 1Aravind Srinivas 1
Verbatim, from the transcripts: the passages where Humanity's Last Exam comes up
The AI lab market map, Robinhood brings startups to retail, GLPs & hedge funds | Diet TBPN
- ▶ 16:50 Jordi Hays Elon Musk announced that XAI is moving away from traditional academic benchmarks like humanity's last exam to focus Grok on maximal utility for real world engineering and 2 times in the scene
Google Gemini 3 Reactions, Google Antigravity, Anthropic-Nvidia-Microsoft Deal | Diet TBPN
- ▶ 6:34 unnamed speaker On Humanity's last exam, it's getting 37.5%.
Weekly Recap: Apple Sitting on $100B, Disney Enters The AI Race, Marc Andreessen, All About GPT-5
- ▶ 1:46:49 John Coogan Humanity's last exam.
Weekly Recap: Grok 4 Launch, Texas Floods, Web Browser War, Top Signals, Meta Smart Glasses
- ▶ 1:15:50 unnamed speaker It's number one on humanity's last exam, which interestingly was effectively like postgraduate PhD level problems, but across a bunch of different domains. 3 times in the scene
- ▶ 1:41:53 unnamed speaker Near said Grok on, uh, humanity's last exam, Grok four, uh, I'm not sure I buy even in the general case that there's a given humanity's last exam number, which implies you discover useful new physics. 2 times in the scene
- ▶ 1:43:34 unnamed speaker Like, I think ArcGIS more interesting, but just like the humanities last exam, the kind of general math, physics knowledge, it doesn't seem, uh, to be that like, it doesn't seem to line up with, like you see GPT, 4.5 kind of does very… 2 times in the scene
ELON MUSK's New Grok 4 Is Here! - We Break it Down
- ▶ 1:01 unnamed speaker It's number one on humanity's last exam, which interestingly was effectively like postgraduate PhD level problems, but across a bunch of different domains. 2 times in the scene
- ▶ 3:45 unnamed speaker Grok got number one on humanity's last exam at 44.4%.
- ▶ 27:04 unnamed speaker Near said Grok on, uh, Humanities Last Exam, Grok four, uh, I'm not sure I buy even in the general case that there's a given Humanities Last Exam number which implies you discover useful new physics. 3 times in the scene
Weekly Recap - Elon Vs Trump, Ukraine's Drone Attack, Cluely Update & OpenAI CRO
Perplexity Founder Explains What Comes Next - Aravind Srinivas on TBPN April 23rd
- ▶ 10:28 Aravind Srinivas Everyone has the same set of benchmarks, LMSIS, LM Arena, GAIA, GPQA, and, and humanity last exam, uh, and, and they're just trying to, like, show their AI is the best, and so they do all the same sort of things that you do in RLHF, um,…
Mike Knoop (Arc Prize) on Why Scaling AI Won’t Get Us to AGI
- ▶ 3:11 Mike Knoop What you're trying to, like, challenge this, like, PhD++ frontier, you know, like frontier math or, you know, humanities last exam, uh, you know, these are sort of AI benchmarks that you really do need to be, like, PhD level plus to be…