Sep 8, 2026 · 1h 5m · a16z
Inside OpenAI’s Breakthroughs in Mathematical Reasoning
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
OpenAI researchers Mehtaab Sawhney and Mark Selke explore how frontier reasoning models are reshaping pure mathematics by overcoming human cognitive fatigue, automating intricate proof discovery, and resolving longstanding open conjectures in geometry and group theory. Through technical whiteboard breakdowns, they analyze the mechanics of AI deliberation while forecasting a future where human mathematicians shift from mechanical proof verification to conceptual synthesis.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the host, purple is the guest (3 minute bins)
Sawhney dismisses abstract notions of aesthetic taste, bluntly taking a utilitarian stance that taste simply means solving hard problems faster by making better pruning choices.
Hardest push from the host ▶ 9:51 Li challenges whether models reason or just roll lucky samplesLi interrupts to push back against the assumption that models possess genuine human-like mathematical intuition, asking whether their success is just parallel lucky sampling.
Biggest teaching moment ▶ 23:11 Sawhney explains linear programming bounds and Fourier transformsSawhney educates the host on how Viazovska's Fields Medal work uses Fourier transform constraints on test functions to establish optimal sphere packing bounds.
The host holds their own ▶ 11:45 Li's critique of Rudin and textbook analysis pedagogyLi exhibits deep domain awareness by pointing out that polished textbooks like Baby Rudin strip out the messy heuristics and struggle that actually teach mathematicians how to reason.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The host as informed peer | Guest teaching | Guest disagreement | The host pushing back | Why |
|---|---|---|---|---|---|---|
| Introductions and the Journey from Pure Math to OpenAI | 5 | 3 | 1 | 2 | Li frames the conversation around her own background in mathematics and probes into how the models moved from basic literature search to deep mathematical reasoning. Selke and Sawhney collaboratively explain how GPT models moved from resolving open Erdős problem references to executing detailed technical arguments without getting lost in epsilon-delta subtleties. | |
| How AI Reasons: Tree Search, Backtracking, and Cognitive Bias | 5 | 4 | 2 | 3 | Sawhney explains how mathematical research is often a gamble against problem difficulty and how reasoning models prune search trees without human cognitive biases. Li actively probes whether the model is merely running parallel lucky samples or actively backtracking like a human mathematician, prompting Selke and Sawhney to clarify the nuances of model restarts versus human context pollution. | |
| Training Reasoning Models and Interpreting AI Chains of Thought | 6 | 3 | 1 | 2 | Li offers a sophisticated critique of training on standard mathematical corpora, noting that formal textbooks like Rudin obscure the messy intuition and struggle of discovery. The guests explain that OpenAI focuses on general-purpose reasoning rather than narrow math auto-formalization, producing chains of thought that read like a colleague's candid email notes. | |
| Whiteboard Deep Dive: High-Dimensional Sphere Packing Breakthrough | 4 | 7 | 1 | 2 | Sawhney and Selke conduct an extensive whiteboard lecture covering sphere packing bounds from dimension 2 up to Viazovska's work in dimensions 8 and 24, and Astra's breakthrough asymptotic LP bound. Li asks clarifying technical questions regarding geometry and linear programming while the guests lead the pedagogical exposition. | |
| Whiteboard Deep Dive: Spherical Codes and Information Theory | 5 | 6 | 1 | 2 | Selke illustrates the connection between spherical codes, error-correcting codes, and representation theory on curved surfaces. Li demonstrates quick domain comprehension on Hamming distances and geometry, while Selke reveals the surprising iterative prompt where asking Astra to push the representation theory led to solving the full-space packing bound. | |
| Mathematical Taste, Prompting Harnesses, and Model Collaboration | 6 | 4 | 2 | 4 | Li pushes deeply into the philosophical and engineering implications of mathematical 'taste', debating whether taste is emergent from task execution or requires a separate supervising harness model. Sawhney counters with a utilitarian definition of taste as solving problems faster, and Selke notes that having a supervisor agent check an underling agent mimics human collaborative checks. | |
| Whiteboard Deep Dive: Disproving Gromov's Conjecture on Sofic Groups | 5 | 6 | 1 | 2 | Selke explains group theory fundamentals and how Astra resolved Gromov's question on whether all groups are sofic by constructing a remarkably short counterexample. Li engages with the Cayley graph formulation and asks what combinatorial structure in the literature resisted finite approximation. | |
| The Changing Role of Mathematicians and the Future of the Field | 5 | 3 | 1 | 2 | Li and the guests discuss how AI-generated mathematics will shift the profession from manual proof construction toward synthesizing, communicating, and verifying concepts. Selke and Sawhney reflect optimistically on math becoming more accessible while noting the upper ceiling of problems like P vs NP will keep mathematicians relevant. |