Rishi Mehta, researcher at Google DeepMind, explains the test-time reinforcement learning mechanism enabling AlphaProof to tackle difficult math problems.
“One of the ways in which it navigates this massive search space is via an idea that we came up with, which we call test time RL. So this is an idea where, like, let's say you're confronted with a problem that you don't know how to solve. And you know, you can do some search with your what you know right now, and you're not, not able to find a proof to it. What the agent then does is it constructs many variations of the problem in the vicinity of that problem and attempts to solve all of them. And if it manages to solve any of them, it learns from that experience and sort of comes closer and closer to solving the original problem.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Rishi Mehta
PredictionNot checkable as stated
Mehta: Specialist data seeds LLMs, but RL drives superhuman capability
“The specialist humans are gonna serve to, like, get the LLM from, like, just a bunch of weights that knows how to do nothing to, like, something that, like is, like, surprisingly strong. And then the RL is gonna take you from there to, like, something that's, …”
Rishi MehtaNov 14, 2024▶ 27:32No Priors Ep. 90 | With Google's DeepMind's AlphaProof Team
Opinion
Mehta: AlphaProof's RL scaling and test-time compute generalize across domains
“Some of the sort of tech we developed here of like, you know, like scaling RL and like figuring out how to spend a lot of inference time compute stuff like this feels like it's Quite generally applicable to many other problems.”
Rishi MehtaNov 14, 2024▶ 21:49No Priors Ep. 90 | With Google's DeepMind's AlphaProof Team
AssertionNot checkable as stated
Mehta: Fields Medalist Tim Gowers failed to find AlphaProof's IMO construction
“Tim Gowers, who was one of our judges and is also a fields medalist. Tried this question for a couple hours and he couldn't find the construction for function that had this property”
Rishi MehtaNov 14, 2024▶ 33:55No Priors Ep. 90 | With Google's DeepMind's AlphaProof Team
Insight
Mehta: As AI solves problems, human roles shift toward question framing
“As machines get better at finding the answers, like, we're going to have to get better at finding the questions. And, you know, like these systems don't have a, you know, their own sort of notion of what questions are interesting. And given a large question, h…”
Rishi MehtaNov 14, 2024▶ 38:19No Priors Ep. 90 | With Google's DeepMind's AlphaProof Team
AssertionSupported
Mehta: AlphaProof Is Strongest in IMO Algebra and Number Theory
“So the IMO problems have come in four categories. So there's algebra number theory, combinatrix, and geometry. The two that it's strongest at are algebra and number theory. It's relatively weaker at combinatrix, although it can do quite good at some IMO combin…”
Rishi MehtaNov 14, 2024▶ 7:51No Priors Ep. 90 | With Google's DeepMind's AlphaProof Team
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 100 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.