Rishi Mehta, researcher at Google DeepMind, breaks down AlphaProof's mathematical strengths across International Mathematical Olympiad (IMO) subject domains.
Prediction Not checkable as stated
Mehta: Specialist data seeds LLMs, but RL drives superhuman capability
“The specialist humans are gonna serve to, like, get the LLM from, like, just a bunch of weights that knows how to do nothing to, like, something that, like is, like, surprisingly strong. And then the RL is gonna take you from there to, like, something that's, …”
Opinion
Mehta: AlphaProof's RL scaling and test-time compute generalize across domains
“Some of the sort of tech we developed here of like, you know, like scaling RL and like figuring out how to spend a lot of inference time compute stuff like this feels like it's Quite generally applicable to many other problems.”
Insight
Mehta: AlphaProof Uses Test-Time RL on Problem Variations to Find Proofs
“One of the ways in which it navigates this massive search space is via an idea that we came up with, which we call test time RL. So this is an idea where, like, let's say you're confronted with a problem that you don't know how to solve. And you know, you can …”
Assertion Not checkable as stated
Mehta: Fields Medalist Tim Gowers failed to find AlphaProof's IMO construction
“Tim Gowers, who was one of our judges and is also a fields medalist. Tried this question for a couple hours and he couldn't find the construction for function that had this property”
Insight
Mehta: As AI solves problems, human roles shift toward question framing
“As machines get better at finding the answers, like, we're going to have to get better at finding the questions. And, you know, like these systems don't have a, you know, their own sort of notion of what questions are interesting. And given a large question, h…”