Łukasz Kaiser, Lead Research Scientist at OpenAI, addresses whether pre-training scaling laws have run out of road and how compute scaling continues to yield model improvements.
“Pre-training, as I said, I think it has reached this upper level of the S-curve in terms of science, but it can scale smoothly. Meaning if you put More compute. You will get better losses if you do things right, which is extremely hard, and that's valuable.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Łukasz Kaiser
AssertionNot checkable as stated
Kaiser: Pre-training between GPT-4 and GPT-5 focused on reducing costs
“The pre-training part in that timeframe was mostly about making things cheaper. Not making things better.”
Łukasz KaiserNov 26, 2025▶ 40:36What’s Next for AI? OpenAI’s Łukasz Kaiser (Transformer Co-Author)
Insight
Kaiser: Reasoning models are the second major milestone after Transformers
“One point was, of course, the Transformers when it started, but the other point was reasoning models.”
Łukasz KaiserNov 26, 2025▶ 3:53What’s Next for AI? OpenAI’s Łukasz Kaiser (Transformer Co-Author)
AssertionNot checkable as stated
Kaiser: Pre-training scaling laws still hold across OpenAI and Google
“What scaling clause says is that your loss will log linearly decrease with your compute. We totally see that and clearly Google sees that and all other labs.”
Łukasz KaiserNov 26, 2025▶ 4:35What’s Next for AI? OpenAI’s Łukasz Kaiser (Transformer Co-Author)
AssertionNot checkable as stated
Kaiser: Model hallucinations are dramatically lower than two years ago
“There was these things called hallucinations. It's still with us to some extent, but dramatically less than two years ago.”
Łukasz KaiserNov 26, 2025▶ 41:53What’s Next for AI? OpenAI’s Łukasz Kaiser (Transformer Co-Author)
Insight
Kaiser: Test-time compute increases AI capabilities faster than pre-training
“Using more tokens to think increases your capability, and it increases it, given the computation, way faster than pre-training, right?”
Łukasz KaiserNov 26, 2025▶ 46:49What’s Next for AI? OpenAI’s Łukasz Kaiser (Transformer Co-Author)
AssertionNot checkable as stated
Kaiser: No frontier AI model can solve a specific first-grade math exercise
“I took one exercise from this math book and none of the frontier models is able to solve it.”
Łukasz KaiserNov 26, 2025▶ 48:11What’s Next for AI? OpenAI’s Łukasz Kaiser (Transformer Co-Author)
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 400 conversations transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.