Rishi Mehta from Google DeepMind explains how AlphaProof solved a construction step in 2024 IMO Problem 6 that Fields Medalist Tim Gowers had struggled with.
Prediction Not checkable as stated
Mehta: Specialist data seeds LLMs, but RL drives superhuman capability
“The specialist humans are gonna serve to, like, get the LLM from, like, just a bunch of weights that knows how to do nothing to, like, something that, like is, like, surprisingly strong. And then the RL is gonna take you from there to, like, something that's, …”
Opinion
Mehta: AlphaProof's RL scaling and test-time compute generalize across domains
“Some of the sort of tech we developed here of like, you know, like scaling RL and like figuring out how to spend a lot of inference time compute stuff like this feels like it's Quite generally applicable to many other problems.”
Insight
Mehta: AlphaProof Uses Test-Time RL on Problem Variations to Find Proofs
“One of the ways in which it navigates this massive search space is via an idea that we came up with, which we call test time RL. So this is an idea where, like, let's say you're confronted with a problem that you don't know how to solve. And you know, you can …”
Insight
Mehta: As AI solves problems, human roles shift toward question framing
“As machines get better at finding the answers, like, we're going to have to get better at finding the questions. And, you know, like these systems don't have a, you know, their own sort of notion of what questions are interesting. And given a large question, h…”
Assertion Supported
Mehta: AlphaProof Is Strongest in IMO Algebra and Number Theory
“So the IMO problems have come in four categories. So there's algebra number theory, combinatrix, and geometry. The two that it's strongest at are algebra and number theory. It's relatively weaker at combinatrix, although it can do quite good at some IMO combin…”