Insight certainty 4/5 debate potential 3/5

Łukasz Kaiser: AI pre-training expands stored knowledge rather than generalization

Łukasz Kaiser · What’s Next for AI? OpenAI’s Łukasz Kaiser (Transformer Co-Author) · Nov 26, 2025 · at 51:47

Łukasz Kaiser, Lead Research Scientist at OpenAI and co-author of the Transformer paper, explains why pre-training compute scaling expands model knowledge without inherently solving generalization.

0:00 / 0:09exact quote · 9.1s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“Pre-training is a little different, right? Because it increases the data together with your increase in model size. So it doesn't necessarily increase generalization. It just uses more knowledge.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Łukasz Kaiser

Assertion Not checkable as stated
Kaiser: Pre-training between GPT-4 and GPT-5 focused on reducing costs
“The pre-training part in that timeframe was mostly about making things cheaper. Not making things better.”
Łukasz Kaiser Nov 26, 2025 ▶ 40:36 What’s Next for AI? OpenAI’s Łukasz Kaiser (Transformer Co-Author)
Insight
Kaiser: Reasoning models are the second major milestone after Transformers
“One point was, of course, the Transformers when it started, but the other point was reasoning models.”
Łukasz Kaiser Nov 26, 2025 ▶ 3:53 What’s Next for AI? OpenAI’s Łukasz Kaiser (Transformer Co-Author)
Assertion Not checkable as stated
Kaiser: Pre-training scaling laws still hold across OpenAI and Google
“What scaling clause says is that your loss will log linearly decrease with your compute. We totally see that and clearly Google sees that and all other labs.”
Łukasz Kaiser Nov 26, 2025 ▶ 4:35 What’s Next for AI? OpenAI’s Łukasz Kaiser (Transformer Co-Author)
Insight
Kaiser: Pre-training science is plateauing, but compute scaling still improves loss
“Pre-training, as I said, I think it has reached this upper level of the S-curve in terms of science, but it can scale smoothly. Meaning if you put More compute. You will get better losses if you do things right, which is extremely hard, and that's valuable.”
Łukasz Kaiser Nov 26, 2025 ▶ 33:07 What’s Next for AI? OpenAI’s Łukasz Kaiser (Transformer Co-Author)
Assertion Not checkable as stated
Kaiser: Model hallucinations are dramatically lower than two years ago
“There was these things called hallucinations. It's still with us to some extent, but dramatically less than two years ago.”
Łukasz Kaiser Nov 26, 2025 ▶ 41:53 What’s Next for AI? OpenAI’s Łukasz Kaiser (Transformer Co-Author)
Insight
Kaiser: Test-time compute increases AI capabilities faster than pre-training
“Using more tokens to think increases your capability, and it increases it, given the computation, way faster than pre-training, right?”
Łukasz Kaiser Nov 26, 2025 ▶ 46:49 What’s Next for AI? OpenAI’s Łukasz Kaiser (Transformer Co-Author)
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.