Dr. Jasper Zhang, a former Math Olympiad gold medalist and CEO of Hyperbolic, explains the differing evaluation methods used by OpenAI and Google DeepMind for their IMO AI benchmarks.
Prediction Not checkable as stated
Scaling math AI becomes purely compute and data once auto-evaluation works
“And then my guess is, I believe in IL, so if for each category, we can figure out the A way to auto-evaluate the results, then after that, it will just be compute and data.”
Prediction Open · timeframe Jul 2028
Formalizing Fermat's Last Theorem in Lean is doable in 2-3 years
“It's I think definitely possible. Yeah. Like he, so the professor is Kevin buzzard and he got like a grant and now he just like focused on writing the proof for Ling. Like he's hoping to finish that in like two or three years. And then basically like if Ling i…”
Assertion Not checkable as stated
Zhang confirms OpenAI's unverified IMO proofs are mathematically correct
“I read the proof that it's still correct, but it's just, like, less official, and that's why people kind of, like, kind of OpenAI received a few backlash over the weekend, and on Monday, DeepMind officially confirmed they have the gold medal and also fully ver…”
Insight
AI models fail at creative combinatorics problems requiring example construction
“If it's like a textbook stuff or like a step-by-step problem, then AI can solve it. But, however, if it requires, like, creativity especially, like, in combinatorics, you kind of need to create, ah, some example, and then try to prove them, ah, that it's the m…”
Prediction Not checkable as stated
Powerful math reasoning models will arrive soon using Lean as a verifier
“And so if you can just use lean to kind of become the verifier, then it's easy. So I think we probably will see very powerful reasoning models in, in math very soon.”
Opinion
Industry AI progress stems from scaling existing research systematically
“A lot of AI like, a lot of AI models built in industry is just, like, a bigger scale of, like in like, research results. They probably they will use, like, existing research, but it's just more, more, like, larger scale, more systematic and more data.”