Jul 24, 2025 · 33m · latent-space
⚡️Math Olympiad gold medalist explains OpenAI and Google DeepMind IMO Gold Performances
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
Former Math Olympiad gold medalist Dr. Jasper Zhang joins Latent Space to analyze how Google DeepMind and OpenAI achieved historic IMO gold performances using natural language reasoning. He contrasts competitive math with frontier research, examines technical breakthroughs in multi-step reinforcement learning, and outlines the pathway toward true Mathematical AGI.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Jasper directly challenges Swix's assertion that scaling existing ideas could not explain the wall-clock speedup, pointing to B200 hardware and inference optimizations.
Hardest push from the hosts ▶ 30:20 Swix pushes back on brute-force scaling explanationsSwix firmly rejects the narrative that test-time scaling alone accounts for reducing runtime from 60 hours to 4.5 hours, asserting it represents an order-of-magnitude algorithmic difference.
Biggest teaching moment ▶ 15:49 Jasper outlines the first-principles breakdown of mathematical intelligenceJasper delivers a detailed, structured masterclass on how human mathematical reasoning operates across knowledge, problem-solving, and creative meta-skills.
The host holds their own ▶ 29:30 Swix cites Gemini evaluation details and parallel search architecturesSwix demonstrates deep industry knowledge by citing the parallel thinking mechanics of DeepThink and o3-Pro, as well as Gemini's unassisted secondary submission.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Timeline and Controversy of IMO Gold Claims | 3 | 6 | 1 | 0 | Swix introduces the context of both DeepMind and OpenAI claiming gold-level IMO performances. Jasper provides detailed insider background on the rollout timeline and why OpenAI faced backlash for unverified claims while Google sought official IMO verification. | |
| Shifting from Formal Lean Systems to Natural Language Models | 5 | 6 | 0 | 0 | Swix notes the major shift from complex, slow formal Lean verifiers taking 60 hours to natural language reasoning. Jasper explains the structure of the IMO and expresses surprise that pure LLMs achieved gold without Lean formalization. | |
| Analyzing IMO Problem 1 and Inductive Decomposition | 3 | 7 | 0 | 0 | Jasper breaks down the exact inductive proof mechanism needed for IMO Problem 1, explaining how both AI labs reduced cases down to base cases. | |
| Competition Math vs Mathematical Research and Taste | 4 | 6 | 1 | 0 | Alessio asks whether mathematical taste matters compared to brute-forcing competition problems. Jasper explains that competition math is primarily brain teasers and does not test research creativity or conjecturing. | |
| Combinatorics Bottlenecks and AI Limits on Problem 6 | 3 | 8 | 1 | 0 | Jasper details findings from the Frontier Math symposium, explaining why LLMs fail at combinatorics problems like Problem 6 because they require constructing novel minimal bounds and experimentation. | |
| Designing a First-Principles Holistic Math Benchmark | 3 | 8 | 0 | 0 | Jasper lays out his comprehensive three-tier taxonomy for evaluating true mathematical capability beyond existing competition benchmarks, covering knowledge, communication, and creative meta-skills. | |
| Expanding Beyond Next-Token Prediction and Benchmark Timeline | 6 | 5 | 0 | 1 | Swix synthesizes Jasper's framework with broader AGI concepts and probes whether natural language or formal verification in Lean is superior. Jasper provides insights into data scarcity in Lean and ongoing efforts like Kevin Buzzard's Fermat proof. | |
| Speculating on Multi-Step RL, Parallel Search, and Compute | 7 | 5 | 3 | 5 | Swix pushes back on the claim that simple scaling explains the time reduction from 60 hours to 4.5 hours, arguing that efficiency requires qualitative algorithmic breakthroughs. Jasper counters with hardware and inference speed gains before reaching mutual agreement. | |
| The Ultimate Milestone: Fields Medal and Math AGI | 3 | 4 | 0 | 0 | Jasper articulates his ultimate milestone for Math AGI: moving beyond synthetic benchmarks to solving decades-old open conjectures and winning a Fields Medal. |