Everything Dr. Jasper Zhang said on any show that made the record, most notable first. Each card names its show and opens the statement there.
OpenAI's IMO performance was not officially verified by the IMO
“It turns out, like, OpenAI actually didn't involve officially with IMO. They just, like, use the problems, but, and then just, like, use their model to test the results, and ask, like, three previous IMO analysts to review them.”
Scaling math AI becomes purely compute and data once auto-evaluation works
“And then my guess is, I believe in IL, so if for each category, we can figure out the A way to auto-evaluate the results, then after that, it will just be compute and data.”
Formalizing Fermat's Last Theorem in Lean is doable in 2-3 years
“It's I think definitely possible. Yeah. Like he, so the professor is Kevin buzzard and he got like a grant and now he just like focused on writing the proof for Ling. Like he's hoping to finish that in like two or three years. And then basically like if Ling i…”
Zhang confirms OpenAI's unverified IMO proofs are mathematically correct
“I read the proof that it's still correct, but it's just, like, less official, and that's why people kind of, like, kind of OpenAI received a few backlash over the weekend, and on Monday, DeepMind officially confirmed they have the gold medal and also fully ver…”
AI models fail at creative combinatorics problems requiring example construction
“If it's like a textbook stuff or like a step-by-step problem, then AI can solve it. But, however, if it requires, like, creativity especially, like, in combinatorics, you kind of need to create, ah, some example, and then try to prove them, ah, that it's the m…”
Powerful math reasoning models will arrive soon using Lean as a verifier
“And so if you can just use lean to kind of become the verifier, then it's easy. So I think we probably will see very powerful reasoning models in, in math very soon.”
Industry AI progress stems from scaling existing research systematically
“A lot of AI like, a lot of AI models built in industry is just, like, a bigger scale of, like in like, research results. They probably they will use, like, existing research, but it's just more, more, like, larger scale, more systematic and more data.”
DeepMind and OpenAI eliminated formal Lean translation for 2025 IMO solutions
“What surprised me is this time they don't use formal language, but instead they just use LM. And so last year when they tried to do the IMO, they need like a like a human to kind of translate the natural language. Problems to Lean, and then they use Lean to ki…”
The IMO tests structured logic rather than mathematical research creativity
“IMO kind of like, it's a good exam to test certain aspect of math skills, which is like problem solving, or you need to have like really complete chain of thoughts. Like, logic, et cetera, but it doesn't test, like, your creativity in terms of, like, research.…”
Qualifying for China's math olympiad team is harder than winning IMO gold
“In terms of difficulty CMO is kind of similar to IMO. I would say it's super competitive in China. So there is a saying, like, if you get, it's harder to get selected for the national, to become the national team instead of then like getting, getting the gold …”
Competition math benchmarks test only a narrow slice of mathematical ability
“Most of the benchmark that people are using right now is, like, AME, IMO, USAMO, all these, like, competition math problems, but, like, they're, they just test, like, a certain aspect of other math skills that a mathematician have”
FrontierMath is problem-focused but AI evaluation needs a skills-based breakdown
“It's so frontier math still problem focused, but then I think we need a more skills breakdown.”
The Lean theorem proving language has only about one million training tokens
“For Lean there are not enough data for Lean, right? There are, like, probably one million tokens in about Lean. Right now, and I know a lot of, like, professors and PhD researchers are trying to build the biggest lean data set in the future.”
Math AGI requires solving open problems and winning a Fields Medal
“The next step is trying to solve some open problems. That's like open for 30 years, 50 years. And then I think the ultimate goal is to Solve like a very challenging problem. I create a new theory and then win a field medal. That's how, how I define a math AGI …”
DeepMind and OpenAI used different reduction methods to solve IMO Problem 1
“Google has one method and then OpenAI AI also have another method but both kind of works. Basically you just reduced any n to three, and then you just do case by case analysis.”