Rishi Mehta, a Google DeepMind researcher on AlphaProof, discusses how the core RL and inference scaling techniques used for math generalize to other AI problem domains.
“Some of the sort of tech we developed here of like, you know, like scaling RL and like figuring out how to spend a lot of inference time compute stuff like this feels like it's Quite generally applicable to many other problems.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Rishi Mehta
PredictionNot checkable as stated
Mehta: Specialist data seeds LLMs, but RL drives superhuman capability
“The specialist humans are gonna serve to, like, get the LLM from, like, just a bunch of weights that knows how to do nothing to, like, something that, like is, like, surprisingly strong. And then the RL is gonna take you from there to, like, something that's, …”
Rishi MehtaNov 14, 2024▶ 27:32No Priors Ep. 90 | With Google's DeepMind's AlphaProof Team
Insight
Mehta: AlphaProof Uses Test-Time RL on Problem Variations to Find Proofs
“One of the ways in which it navigates this massive search space is via an idea that we came up with, which we call test time RL. So this is an idea where, like, let's say you're confronted with a problem that you don't know how to solve. And you know, you can …”
Rishi MehtaNov 14, 2024▶ 8:13No Priors Ep. 90 | With Google's DeepMind's AlphaProof Team
AssertionNot checkable as stated
Mehta: Fields Medalist Tim Gowers failed to find AlphaProof's IMO construction
“Tim Gowers, who was one of our judges and is also a fields medalist. Tried this question for a couple hours and he couldn't find the construction for function that had this property”
Rishi MehtaNov 14, 2024▶ 33:55No Priors Ep. 90 | With Google's DeepMind's AlphaProof Team
Insight
Mehta: As AI solves problems, human roles shift toward question framing
“As machines get better at finding the answers, like, we're going to have to get better at finding the questions. And, you know, like these systems don't have a, you know, their own sort of notion of what questions are interesting. And given a large question, h…”
Rishi MehtaNov 14, 2024▶ 38:19No Priors Ep. 90 | With Google's DeepMind's AlphaProof Team
AssertionSupported
Mehta: AlphaProof Is Strongest in IMO Algebra and Number Theory
“So the IMO problems have come in four categories. So there's algebra number theory, combinatrix, and geometry. The two that it's strongest at are algebra and number theory. It's relatively weaker at combinatrix, although it can do quite good at some IMO combin…”
Rishi MehtaNov 14, 2024▶ 7:51No Priors Ep. 90 | With Google's DeepMind's AlphaProof Team
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 100 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.