pre-training

18 statements across 9 episodes · 6 bullish · 4 bearish · 9 people on the record · first statement Oct 16, 2025 by Jerry Tworek · across every show →

Everything said about pre-training, oldest first

Oct 16, 2025 positive
Insight
Tworek: Pre-Training on Unlabeled Data Yields Far More Intelligence Than Supervised Mapping
“There are many more bits usually in the targets than in the labels and studying the structure of targets itself. It yields much more learning and much more intelligence than learning the mapping itself. So like spending a whole compute on just learning the dat…”
Jerry Tworek Oct 16, 2025 ▶ 49:27 How GPT-5 Thinks — OpenAI VP of Research Jerry Tworek
Oct 16, 2025 negative
Insight
Tworek: Reinforcement learning and pre-training require each other to succeed
“And like, I don't like in terms of a pure RL, I don't think like really pure RL makes sense. RL needs Pre-training to be successful. And I think pre-training, as I said before, needs RL to be successful as well.”
Jerry Tworek Oct 16, 2025 ▶ 1:13:14 How GPT-5 Thinks — OpenAI VP of Research Jerry Tworek
Oct 16, 2025
Insight
Jerry Tworek: Pre-training AI models is mathematically simple compared to RL
“The first thing that is important to know and understand, RL is hard. Like, conceptually, if you think about it, and there's still a lot of depth to it, but very conceptually, mathematically speaking, pre-training is dead simple.”
Jerry Tworek Oct 16, 2025 ▶ 53:31 How GPT-5 Thinks — OpenAI VP of Research Jerry Tworek
Oct 23, 2025 bearish
Prediction Not checkable as stated
Future AI models will continue to rely on pre-training data
“Personally, I think that's unlikely. Not, not because pre-training is strictly necessary. I think we may well be able to train something completely from scratch, as we've been able to do in other domains, but more because pre-training on this vast data sets th…”
Julian Schrittwieser Oct 23, 2025 ▶ 21:27 Are We Misreading the AI Exponential? Julian Schrittwieser on Move 37 & Scaling RL (Anthropic)
Oct 23, 2025 positive
Assertion Partly supported
Reinforcement learning scaling yields returns on compute similar to pre-training
“If you look at all the RL literature over time, we see very similar returns on compute in pre-training and in RL, where we can invest exponentially more compute in RL and keep getting benefits.”
Julian Schrittwieser Oct 23, 2025 ▶ 42:01 Are We Misreading the AI Exponential? Julian Schrittwieser on Move 37 & Scaling RL (Anthropic)
Oct 23, 2025
Insight
Schrittwieser: AI pre-training risks over-restricting an agent's exploration search space
“I think the main, you know, the main challenge or the main thing you need to watch out for is that you don't over encode or you don't restrict your search space too much. If your pre-training, if your prior knowledge prevents you from exploring something that …”
Julian Schrittwieser Oct 23, 2025 ▶ 38:41 Are We Misreading the AI Exponential? Julian Schrittwieser on Move 37 & Scaling RL (Anthropic)
Oct 30, 2025
Insight
Smarter pre-training methods will reduce massive AI data spending requirements
“So I think we're just getting smarter about how to do pre-training rather than shoving everything we have into a bucket and like seeing what happens. And so as a result of that, you might not necessarily have to spend the exact same amount of money to get a ca…”
Nathan Benaich Oct 30, 2025 ▶ 46:48 State of AI 2025 with Nathan Benaich: Power Deals, Reasoning Breakthroughs, Real Revenue
Nov 6, 2025 bearish
Assertion Not checkable as stated
Kant: Next-token pre-training gains hit a sigmoidal curve and slowed down
“The first paradigm of kind of pre-training of predicting the next token on the web was becoming sigmoidal and was slowing down in terms of the gains that it had.”
Eiso Kant Nov 6, 2025 ▶ 2:39 Intelligence Isn’t Enough: Why Energy & Compute Decide the AGI Race – Eiso Kant
Nov 26, 2025 bullish
Insight
Kaiser: Reasoning yields far greater AI capability gains per dollar than pre-training
“With the new paradigm of reasoning, you can get much more gains for the same amount of money because it's on this like lower and like, there are just discoveries to be made and these discoveries unlock insane capabilities.”
Łukasz Kaiser Nov 26, 2025 ▶ 4:56 What’s Next for AI? OpenAI’s Łukasz Kaiser (Transformer Co-Author)
Nov 26, 2025
Assertion Not checkable as stated
Kaiser: Pre-training consumes the most GPUs of any AI development stage
“Currently, pre-training just uses the most GPUs of all the parts, so it needs the most GPUs, right?”
Łukasz Kaiser Nov 26, 2025 ▶ 32:09 What’s Next for AI? OpenAI’s Łukasz Kaiser (Transformer Co-Author)
Nov 26, 2025 bullish
Insight
Kaiser: Test-time compute increases AI capabilities faster than pre-training
“Using more tokens to think increases your capability, and it increases it, given the computation, way faster than pre-training, right?”
Łukasz Kaiser Nov 26, 2025 ▶ 46:49 What’s Next for AI? OpenAI’s Łukasz Kaiser (Transformer Co-Author)
Nov 26, 2025 neutral
Insight
Łukasz Kaiser: AI pre-training expands stored knowledge rather than generalization
“Pre-training is a little different, right? Because it increases the data together with your increase in model size. So it doesn't necessarily increase generalization. It just uses more knowledge.”
Łukasz Kaiser Nov 26, 2025 ▶ 51:47 What’s Next for AI? OpenAI’s Łukasz Kaiser (Transformer Co-Author)
Dec 18, 2025 neutral
Insight
Bourgeau: Architecture and data innovation currently matter more than scale
“The other parts are architecture and data innovation. These also play a really, really important part in the Performance of pre-training and probably even more so than pure scale these days, but scaling is still an important factor as well.”
Sebastien Bourgeau Dec 18, 2025 ▶ 31:22 ”We’re Ahead of Where I Thought We’d Be” — Gemini 3 & the Future of AI
Jan 15, 2026
Assertion Not checkable as stated
Izmailov: AI researchers cannot reliably trace model behaviors to pre-training sources
“We don't really know what's the source of this type of behaviors, but that's also true for a lot of other behaviors in the models with, like, even the good ones. We don't really, we cannot always pin down, like, where they come from in the pre-training.”
Pavel Izmailov Jan 15, 2026 ▶ 3:57 The Evaluators Are Being Evaluated — Pavel Izmailov (Anthropic/NYU)
Jan 29, 2026 negative
Insight
Pre-training is no longer where the low-hanging AI gains lie
“Pre-training is not dead, but pre-training is boring. So it's not where the low hanging fruit is anymore.”
Sebastian Raschka Jan 29, 2026 ▶ 17:32 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Apr 2, 2026 neutral
Prediction Not checkable as stated
AI model progress will alternate between pre-training and post-training breakthroughs
“We're going to be having a bit of a swing back and forth between pre-training and post-training.”
Mostafa Dehghani Apr 2, 2026 ▶ 26:55 AI is Already Building AI — Google DeepMind’s Mostafa Dehghani
Apr 2, 2026 bullish
Prediction Not checkable as stated
New pre-training techniques will drastically boost base AI model capabilities
“The way that we used to do pre-training, maybe, like, you know, like two, a year ago or two years ago maybe, like, you know, diminishing return is, like, obvious, but I can see how new ideas are bringing, like, you know, fresh, fresh energy into the pre-traini…”
Mostafa Dehghani Apr 2, 2026 ▶ 29:18 AI is Already Building AI — Google DeepMind’s Mostafa Dehghani
Apr 2, 2026 bullish
Insight
Post-training techniques cannot compensate for a weak base AI model
“Pre-training is still the foundation and like, you can never post-train your way out of a week-based model.”
Mostafa Dehghani Apr 2, 2026 ▶ 27:01 AI is Already Building AI — Google DeepMind’s Mostafa Dehghani
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.