FrontierMath, every mention
12 scenes across 4 shows · ← back to FrontierMath
tap a year for its mentions
Latent Space 11
TBPN 3
the a16z Podcast 1
the MAD Podcast 1
every year every show
Latent Space 11
TBPN 3
the MAD Podcast 1
the a16z Podcast 1
Verbatim, from the transcripts: passages where FrontierMath comes up on Latent Space, TBPN, the MAD Podcast, the a16z Podcast
Astra Reactions, Jobs Print, Cybercab Launch | Diet TBPN
- ▶ 5:18 unnamed speaker Uh, realistically frontier math tier four, I think they scored 99% on that.
Model Mayhem, GPT-6 Astra, Why Nvidia Bought Hugging Face | Diet TBPN
- ▶ 21:08 Jordi Hays Uh, so almost fully saturated frontier math tier four V two gets a 97.6 deep sweet 74.1 exploit bench also saturated at a hundred percent.
Scaling Past Informal AI - Carina Hong, Axiom Math
- ▶ 51:10 Carina Hong Yeah, so you can say that if you're on Frontier Math, and you have like, sorry, Frontier Lab, and you have like infinite resources, or there's this, by definition, no running out of gas, right? 2 times in the scene
- ▶ 1:29:24 Carina Hong Like you have someone who's a core contributor, Frontier Math tier four, really great benchmark setter.
AI That Can Prove It’s Right: Verification as the Missing Layer in AI — Carina Hong
- ▶ 18:17 Corinna Hong Like, you know, we have seen from, say, Frontier Math and other benchmark, which only compels a numerical answer that it doesn't actually necessarily reflect the model's capability in the logical reasoning.
Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
- ▶ 33:27 Shawn Wang Or like a frontier map type? 2 times in the scene
- ▶ 1:15:28 Micah Hill-Smith In intelligence index, adding critical point, the, um, physics eval George was talking about similar to frontier math, that gives us completely new view with a brand new data set of very, very hard research problems.
Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question
- ▶ 15:14 Nathan Labenz Now we've got the frontier math benchmark that is, I think now like up to 25%.
⚡️Math Olympiad gold medalist explains OpenAI and Google DeepMind IMO Gold Performances
- ▶ 22:39 unnamed speaker This is a different word from, from the frontier math people. 2 times in the scene
- ▶ 31:22 unnamed speaker What, what other, like, IMO, uh, what other math, like, uh, you know, you have, like, frontier math, you have, uh, your benchmark that you're putting out, uh, any other, like, milestones that you're looking for in, in, in this part of the…
The Future of Artificial Intelligence
- ▶ 15:28 unnamed speaker Um, but it's on Arc AGI, SWE Bench, and Frontier Math, Advanced Mathematics, like, things that most consumers might not even care about the benchmarks on, but are potentially good proxies.
2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
- ▶ 1:06:55 Shawn Wang And so now people care a lot about frontier math coding, right? 2 times in the scene