Assertion Supported AI assessment confidence: 95% certainty 4/5 debate potential 2/5

Kaiser: Reinforcement learning causes AI models to self-correct mistakes

Łukasz Kaiser · What’s Next for AI? OpenAI’s Łukasz Kaiser (Transformer Co-Author) · Nov 26, 2025 · at 20:59

Łukasz Kaiser, Lead Research Scientist at OpenAI, describes how reinforcement learning training induces self-verification and error correction behavior in reasoning models.

0:00 / 0:24exact quote · 24.1s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“Even for math and coding, you start seeing that the models start correcting their own mistakes, right? Earlier, if the model made a mistake, it generally just tell you what it did and insist that the mistake was right or something like that. With the thinking, it's like, oh, I often make mistakes, but I need to verify and correct myself to give the correct answer. So, so this just emerges from this reinforcement learning.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Łukasz Kaiser

Assertion Not checkable as stated
Kaiser: Pre-training between GPT-4 and GPT-5 focused on reducing costs
“The pre-training part in that timeframe was mostly about making things cheaper. Not making things better.”
Łukasz Kaiser Nov 26, 2025 ▶ 40:36 What’s Next for AI? OpenAI’s Łukasz Kaiser (Transformer Co-Author)
Insight
Kaiser: Reasoning models are the second major milestone after Transformers
“One point was, of course, the Transformers when it started, but the other point was reasoning models.”
Łukasz Kaiser Nov 26, 2025 ▶ 3:53 What’s Next for AI? OpenAI’s Łukasz Kaiser (Transformer Co-Author)
Assertion Not checkable as stated
Kaiser: Pre-training scaling laws still hold across OpenAI and Google
“What scaling clause says is that your loss will log linearly decrease with your compute. We totally see that and clearly Google sees that and all other labs.”
Łukasz Kaiser Nov 26, 2025 ▶ 4:35 What’s Next for AI? OpenAI’s Łukasz Kaiser (Transformer Co-Author)
Insight
Kaiser: Pre-training science is plateauing, but compute scaling still improves loss
“Pre-training, as I said, I think it has reached this upper level of the S-curve in terms of science, but it can scale smoothly. Meaning if you put More compute. You will get better losses if you do things right, which is extremely hard, and that's valuable.”
Łukasz Kaiser Nov 26, 2025 ▶ 33:07 What’s Next for AI? OpenAI’s Łukasz Kaiser (Transformer Co-Author)
Assertion Not checkable as stated
Kaiser: Model hallucinations are dramatically lower than two years ago
“There was these things called hallucinations. It's still with us to some extent, but dramatically less than two years ago.”
Łukasz Kaiser Nov 26, 2025 ▶ 41:53 What’s Next for AI? OpenAI’s Łukasz Kaiser (Transformer Co-Author)
Insight
Kaiser: Test-time compute increases AI capabilities faster than pre-training
“Using more tokens to think increases your capability, and it increases it, given the computation, way faster than pre-training, right?”
Łukasz Kaiser Nov 26, 2025 ▶ 46:49 What’s Next for AI? OpenAI’s Łukasz Kaiser (Transformer Co-Author)
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.