Jul 24, 2025 · 33m · latent-space

⚡️Math Olympiad gold medalist explains OpenAI and Google DeepMind IMO Gold Performances

Dr. Jasper Zhang · 20m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Former Math Olympiad gold medalist Dr. Jasper Zhang joins Latent Space to analyze how Google DeepMind and OpenAI achieved historic IMO gold performances using natural language reasoning. He contrasts competitive math with frontier research, examines technical breakthroughs in multi-step reinforcement learning, and outlines the pathway toward true Mathematical AGI.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 4.1 Guest teaching 6.1 Guest disagreement 0.7 The hosts pushing back 0.7
05100:0010:0020:0030:001:25–4:15 · The hosts as informed peer 3/10 Timeline and Controversy of IMO Gold Claims Swix introduces the context of both DeepMind and OpenAI claiming gold-level IMO performances. Jasper provides detailed insider background on the rollout timeline and why OpenAI faced backlash for unverified claims while Google sought official IMO verification.4:17–7:21 · The hosts as informed peer 5/10 Shifting from Formal Lean Systems to Natural Language Models Swix notes the major shift from complex, slow formal Lean verifiers taking 60 hours to natural language reasoning. Jasper explains the structure of the IMO and expresses surprise that pure LLMs achieved gold without Lean formalization.7:22–9:31 · The hosts as informed peer 3/10 Analyzing IMO Problem 1 and Inductive Decomposition Jasper breaks down the exact inductive proof mechanism needed for IMO Problem 1, explaining how both AI labs reduced cases down to base cases.9:31–12:36 · The hosts as informed peer 4/10 Competition Math vs Mathematical Research and Taste Alessio asks whether mathematical taste matters compared to brute-forcing competition problems. Jasper explains that competition math is primarily brain teasers and does not test research creativity or conjecturing.12:36–15:08 · The hosts as informed peer 3/10 Combinatorics Bottlenecks and AI Limits on Problem 6 Jasper details findings from the Frontier Math symposium, explaining why LLMs fail at combinatorics problems like Problem 6 because they require constructing novel minimal bounds and experimentation.15:09–21:59 · The hosts as informed peer 3/10 Designing a First-Principles Holistic Math Benchmark Jasper lays out his comprehensive three-tier taxonomy for evaluating true mathematical capability beyond existing competition benchmarks, covering knowledge, communication, and creative meta-skills.21:59–27:08 · The hosts as informed peer 6/10 Expanding Beyond Next-Token Prediction and Benchmark Timeline Swix synthesizes Jasper's framework with broader AGI concepts and probes whether natural language or formal verification in Lean is superior. Jasper provides insights into data scarcity in Lean and ongoing efforts like Kevin Buzzard's Fermat proof.27:08–31:20 · The hosts as informed peer 7/10 Speculating on Multi-Step RL, Parallel Search, and Compute Swix pushes back on the claim that simple scaling explains the time reduction from 60 hours to 4.5 hours, arguing that efficiency requires qualitative algorithmic breakthroughs. Jasper counters with hardware and inference speed gains before reaching mutual agreement.31:20–32:52 · The hosts as informed peer 3/10 The Ultimate Milestone: Fields Medal and Math AGI Jasper articulates his ultimate milestone for Math AGI: moving beyond synthetic benchmarks to solving decades-old open conjectures and winning a Fields Medal.1:25–4:15 · Guest teaching 6/10 Timeline and Controversy of IMO Gold Claims Swix introduces the context of both DeepMind and OpenAI claiming gold-level IMO performances. Jasper provides detailed insider background on the rollout timeline and why OpenAI faced backlash for unverified claims while Google sought official IMO verification.4:17–7:21 · Guest teaching 6/10 Shifting from Formal Lean Systems to Natural Language Models Swix notes the major shift from complex, slow formal Lean verifiers taking 60 hours to natural language reasoning. Jasper explains the structure of the IMO and expresses surprise that pure LLMs achieved gold without Lean formalization.7:22–9:31 · Guest teaching 7/10 Analyzing IMO Problem 1 and Inductive Decomposition Jasper breaks down the exact inductive proof mechanism needed for IMO Problem 1, explaining how both AI labs reduced cases down to base cases.9:31–12:36 · Guest teaching 6/10 Competition Math vs Mathematical Research and Taste Alessio asks whether mathematical taste matters compared to brute-forcing competition problems. Jasper explains that competition math is primarily brain teasers and does not test research creativity or conjecturing.12:36–15:08 · Guest teaching 8/10 Combinatorics Bottlenecks and AI Limits on Problem 6 Jasper details findings from the Frontier Math symposium, explaining why LLMs fail at combinatorics problems like Problem 6 because they require constructing novel minimal bounds and experimentation.15:09–21:59 · Guest teaching 8/10 Designing a First-Principles Holistic Math Benchmark Jasper lays out his comprehensive three-tier taxonomy for evaluating true mathematical capability beyond existing competition benchmarks, covering knowledge, communication, and creative meta-skills.21:59–27:08 · Guest teaching 5/10 Expanding Beyond Next-Token Prediction and Benchmark Timeline Swix synthesizes Jasper's framework with broader AGI concepts and probes whether natural language or formal verification in Lean is superior. Jasper provides insights into data scarcity in Lean and ongoing efforts like Kevin Buzzard's Fermat proof.27:08–31:20 · Guest teaching 5/10 Speculating on Multi-Step RL, Parallel Search, and Compute Swix pushes back on the claim that simple scaling explains the time reduction from 60 hours to 4.5 hours, arguing that efficiency requires qualitative algorithmic breakthroughs. Jasper counters with hardware and inference speed gains before reaching mutual agreement.31:20–32:52 · Guest teaching 4/10 The Ultimate Milestone: Fields Medal and Math AGI Jasper articulates his ultimate milestone for Math AGI: moving beyond synthetic benchmarks to solving decades-old open conjectures and winning a Fields Medal.1:25–4:15 · Guest disagreement 1/10 Timeline and Controversy of IMO Gold Claims Swix introduces the context of both DeepMind and OpenAI claiming gold-level IMO performances. Jasper provides detailed insider background on the rollout timeline and why OpenAI faced backlash for unverified claims while Google sought official IMO verification.4:17–7:21 · Guest disagreement 0/10 Shifting from Formal Lean Systems to Natural Language Models Swix notes the major shift from complex, slow formal Lean verifiers taking 60 hours to natural language reasoning. Jasper explains the structure of the IMO and expresses surprise that pure LLMs achieved gold without Lean formalization.7:22–9:31 · Guest disagreement 0/10 Analyzing IMO Problem 1 and Inductive Decomposition Jasper breaks down the exact inductive proof mechanism needed for IMO Problem 1, explaining how both AI labs reduced cases down to base cases.9:31–12:36 · Guest disagreement 1/10 Competition Math vs Mathematical Research and Taste Alessio asks whether mathematical taste matters compared to brute-forcing competition problems. Jasper explains that competition math is primarily brain teasers and does not test research creativity or conjecturing.12:36–15:08 · Guest disagreement 1/10 Combinatorics Bottlenecks and AI Limits on Problem 6 Jasper details findings from the Frontier Math symposium, explaining why LLMs fail at combinatorics problems like Problem 6 because they require constructing novel minimal bounds and experimentation.15:09–21:59 · Guest disagreement 0/10 Designing a First-Principles Holistic Math Benchmark Jasper lays out his comprehensive three-tier taxonomy for evaluating true mathematical capability beyond existing competition benchmarks, covering knowledge, communication, and creative meta-skills.21:59–27:08 · Guest disagreement 0/10 Expanding Beyond Next-Token Prediction and Benchmark Timeline Swix synthesizes Jasper's framework with broader AGI concepts and probes whether natural language or formal verification in Lean is superior. Jasper provides insights into data scarcity in Lean and ongoing efforts like Kevin Buzzard's Fermat proof.27:08–31:20 · Guest disagreement 3/10 Speculating on Multi-Step RL, Parallel Search, and Compute Swix pushes back on the claim that simple scaling explains the time reduction from 60 hours to 4.5 hours, arguing that efficiency requires qualitative algorithmic breakthroughs. Jasper counters with hardware and inference speed gains before reaching mutual agreement.31:20–32:52 · Guest disagreement 0/10 The Ultimate Milestone: Fields Medal and Math AGI Jasper articulates his ultimate milestone for Math AGI: moving beyond synthetic benchmarks to solving decades-old open conjectures and winning a Fields Medal.1:25–4:15 · The hosts pushing back 0/10 Timeline and Controversy of IMO Gold Claims Swix introduces the context of both DeepMind and OpenAI claiming gold-level IMO performances. Jasper provides detailed insider background on the rollout timeline and why OpenAI faced backlash for unverified claims while Google sought official IMO verification.4:17–7:21 · The hosts pushing back 0/10 Shifting from Formal Lean Systems to Natural Language Models Swix notes the major shift from complex, slow formal Lean verifiers taking 60 hours to natural language reasoning. Jasper explains the structure of the IMO and expresses surprise that pure LLMs achieved gold without Lean formalization.7:22–9:31 · The hosts pushing back 0/10 Analyzing IMO Problem 1 and Inductive Decomposition Jasper breaks down the exact inductive proof mechanism needed for IMO Problem 1, explaining how both AI labs reduced cases down to base cases.9:31–12:36 · The hosts pushing back 0/10 Competition Math vs Mathematical Research and Taste Alessio asks whether mathematical taste matters compared to brute-forcing competition problems. Jasper explains that competition math is primarily brain teasers and does not test research creativity or conjecturing.12:36–15:08 · The hosts pushing back 0/10 Combinatorics Bottlenecks and AI Limits on Problem 6 Jasper details findings from the Frontier Math symposium, explaining why LLMs fail at combinatorics problems like Problem 6 because they require constructing novel minimal bounds and experimentation.15:09–21:59 · The hosts pushing back 0/10 Designing a First-Principles Holistic Math Benchmark Jasper lays out his comprehensive three-tier taxonomy for evaluating true mathematical capability beyond existing competition benchmarks, covering knowledge, communication, and creative meta-skills.21:59–27:08 · The hosts pushing back 1/10 Expanding Beyond Next-Token Prediction and Benchmark Timeline Swix synthesizes Jasper's framework with broader AGI concepts and probes whether natural language or formal verification in Lean is superior. Jasper provides insights into data scarcity in Lean and ongoing efforts like Kevin Buzzard's Fermat proof.27:08–31:20 · The hosts pushing back 5/10 Speculating on Multi-Step RL, Parallel Search, and Compute Swix pushes back on the claim that simple scaling explains the time reduction from 60 hours to 4.5 hours, arguing that efficiency requires qualitative algorithmic breakthroughs. Jasper counters with hardware and inference speed gains before reaching mutual agreement.31:20–32:52 · The hosts pushing back 0/10 The Ultimate Milestone: Fields Medal and Math AGI Jasper articulates his ultimate milestone for Math AGI: moving beyond synthetic benchmarks to solving decades-old open conjectures and winning a Fields Medal.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 0%33:00 · the hosts 0% · guest 0%
Sharpest disagreement ▶ 30:42 Jasper challenges the assumption that scaling alone could not yield the speedup

Jasper directly challenges Swix's assertion that scaling existing ideas could not explain the wall-clock speedup, pointing to B200 hardware and inference optimizations.

Hardest push from the hosts ▶ 30:20 Swix pushes back on brute-force scaling explanations

Swix firmly rejects the narrative that test-time scaling alone accounts for reducing runtime from 60 hours to 4.5 hours, asserting it represents an order-of-magnitude algorithmic difference.

Biggest teaching moment ▶ 15:49 Jasper outlines the first-principles breakdown of mathematical intelligence

Jasper delivers a detailed, structured masterclass on how human mathematical reasoning operates across knowledge, problem-solving, and creative meta-skills.

The host holds their own ▶ 29:30 Swix cites Gemini evaluation details and parallel search architectures

Swix demonstrates deep industry knowledge by citing the parallel thinking mechanics of DeepThink and o3-Pro, as well as Gemini's unassisted secondary submission.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Timeline and Controversy of IMO Gold Claims 3610 Swix introduces the context of both DeepMind and OpenAI claiming gold-level IMO performances. Jasper provides detailed insider background on the rollout timeline and why OpenAI faced backlash for unverified claims while Google sought official IMO verification.
Shifting from Formal Lean Systems to Natural Language Models 5600 Swix notes the major shift from complex, slow formal Lean verifiers taking 60 hours to natural language reasoning. Jasper explains the structure of the IMO and expresses surprise that pure LLMs achieved gold without Lean formalization.
Analyzing IMO Problem 1 and Inductive Decomposition 3700 Jasper breaks down the exact inductive proof mechanism needed for IMO Problem 1, explaining how both AI labs reduced cases down to base cases.
Competition Math vs Mathematical Research and Taste 4610 Alessio asks whether mathematical taste matters compared to brute-forcing competition problems. Jasper explains that competition math is primarily brain teasers and does not test research creativity or conjecturing.
Combinatorics Bottlenecks and AI Limits on Problem 6 3810 Jasper details findings from the Frontier Math symposium, explaining why LLMs fail at combinatorics problems like Problem 6 because they require constructing novel minimal bounds and experimentation.
Designing a First-Principles Holistic Math Benchmark 3800 Jasper lays out his comprehensive three-tier taxonomy for evaluating true mathematical capability beyond existing competition benchmarks, covering knowledge, communication, and creative meta-skills.
Expanding Beyond Next-Token Prediction and Benchmark Timeline 6501 Swix synthesizes Jasper's framework with broader AGI concepts and probes whether natural language or formal verification in Lean is superior. Jasper provides insights into data scarcity in Lean and ongoing efforts like Kevin Buzzard's Fermat proof.
Speculating on Multi-Step RL, Parallel Search, and Compute 7535 Swix pushes back on the claim that simple scaling explains the time reduction from 60 hours to 4.5 hours, arguing that efficiency requires qualitative algorithmic breakthroughs. Jasper counters with hardware and inference speed gains before reaching mutual agreement.
The Ultimate Milestone: Fields Medal and Math AGI 3400 Jasper articulates his ultimate milestone for Math AGI: moving beyond synthetic benchmarks to solving decades-old open conjectures and winning a Fields Medal.

Statements from this episode (15)

Assertion Supported
OpenAI's IMO performance was not officially verified by the IMO
“It turns out, like, OpenAI actually didn't involve officially with IMO. They just, like, use the problems, but, and then just, like, use their model to test the results, and ask, like, three previous IMO analysts to review them.”
Dr. Jasper Zhang Jul 24, 2025 ▶ 3:28
Assertion Not checkable as stated
Zhang confirms OpenAI's unverified IMO proofs are mathematically correct
“I read the proof that it's still correct, but it's just, like, less official, and that's why people kind of, like, kind of OpenAI received a few backlash over the weekend, and on Monday, DeepMind officially confirmed they have the gold medal and also fully ver…”
Dr. Jasper Zhang Jul 24, 2025 ▶ 3:51
Assertion Supported
DeepMind and OpenAI eliminated formal Lean translation for 2025 IMO solutions
“What surprised me is this time they don't use formal language, but instead they just use LM. And so last year when they tried to do the IMO, they need like a like a human to kind of translate the natural language. Problems to Lean, and then they use Lean to ki…”
Dr. Jasper Zhang Jul 24, 2025 ▶ 6:10
Assertion Supported
DeepMind and OpenAI used different reduction methods to solve IMO Problem 1
“Google has one method and then OpenAI AI also have another method but both kind of works. Basically you just reduced any n to three, and then you just do case by case analysis.”
Dr. Jasper Zhang Jul 24, 2025 ▶ 8:48
Insight
The IMO tests structured logic rather than mathematical research creativity
“IMO kind of like, it's a good exam to test certain aspect of math skills, which is like problem solving, or you need to have like really complete chain of thoughts. Like, logic, et cetera, but it doesn't test, like, your creativity in terms of, like, research.…”
Dr. Jasper Zhang Jul 24, 2025 ▶ 10:14
Assertion Not checkable as stated
Qualifying for China's math olympiad team is harder than winning IMO gold
“In terms of difficulty CMO is kind of similar to IMO. I would say it's super competitive in China. So there is a saying, like, if you get, it's harder to get selected for the national, to become the national team instead of then like getting, getting the gold …”
Dr. Jasper Zhang Jul 24, 2025 ▶ 11:46
Insight
AI models fail at creative combinatorics problems requiring example construction
“If it's like a textbook stuff or like a step-by-step problem, then AI can solve it. But, however, if it requires, like, creativity especially, like, in combinatorics, you kind of need to create, ah, some example, and then try to prove them, ah, that it's the m…”
Dr. Jasper Zhang Jul 24, 2025 ▶ 13:47
Opinion
Competition math benchmarks test only a narrow slice of mathematical ability
“Most of the benchmark that people are using right now is, like, AME, IMO, USAMO, all these, like, competition math problems, but, like, they're, they just test, like, a certain aspect of other math skills that a mathematician have”
Dr. Jasper Zhang Jul 24, 2025 ▶ 16:00
Prediction Not checkable as stated
Powerful math reasoning models will arrive soon using Lean as a verifier
“And so if you can just use lean to kind of become the verifier, then it's easy. So I think we probably will see very powerful reasoning models in, in math very soon.”
Dr. Jasper Zhang Jul 24, 2025 ▶ 19:37
Prediction Not checkable as stated
Scaling math AI becomes purely compute and data once auto-evaluation works
“And then my guess is, I believe in IL, so if for each category, we can figure out the A way to auto-evaluate the results, then after that, it will just be compute and data.”
Dr. Jasper Zhang Jul 24, 2025 ▶ 21:19
Opinion
FrontierMath is problem-focused but AI evaluation needs a skills-based breakdown
“It's so frontier math still problem focused, but then I think we need a more skills breakdown.”
Dr. Jasper Zhang Jul 24, 2025 ▶ 22:44
Assertion Contradicted
The Lean theorem proving language has only about one million training tokens
“For Lean there are not enough data for Lean, right? There are, like, probably one million tokens in about Lean. Right now, and I know a lot of, like, professors and PhD researchers are trying to build the biggest lean data set in the future.”
Dr. Jasper Zhang Jul 24, 2025 ▶ 24:50
Prediction Open · timeframe Jul 2028
Formalizing Fermat's Last Theorem in Lean is doable in 2-3 years
“It's I think definitely possible. Yeah. Like he, so the professor is Kevin buzzard and he got like a grant and now he just like focused on writing the proof for Ling. Like he's hoping to finish that in like two or three years. And then basically like if Ling i…”
Dr. Jasper Zhang Jul 24, 2025 ▶ 26:35
Opinion
Industry AI progress stems from scaling existing research systematically
“A lot of AI like, a lot of AI models built in industry is just, like, a bigger scale of, like in like, research results. They probably they will use, like, existing research, but it's just more, more, like, larger scale, more systematic and more data.”
Dr. Jasper Zhang Jul 24, 2025 ▶ 27:52
Opinion
Math AGI requires solving open problems and winning a Fields Medal
“The next step is trying to solve some open problems. That's like open for 30 years, 50 years. And then I think the ultimate goal is to Solve like a very challenging problem. I create a new theory and then win a field medal. That's how, how I define a math AGI …”
Dr. Jasper Zhang Jul 24, 2025 ▶ 32:06
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.