Pre Training

topic on 13 shows · 59 statements across 46 episodes

the Y Combinator Startup Podcast BG2 Pod Latent Space Lenny's Podcast No Priors the Official SaaStr Podcast Invest Like the Best Catalyst the MAD Podcast the a16z Podcast Big Technology TBPN 20VC

59 statements about Pre Training, every show

BIG TECHNOLOGY Prediction Not checkable as stated
Kedrosky: Investors will pressure AI labs to slash massive pre-training spend
“Once investors look under the hood and see more and more of this, they'll be questioning, why are we spending so much on pre-training? Why are you doing billion dollar training runs anymore? If most of the gains and models are coming from post training and RLH…”
Paul Kedrosky Aug 12, 2026 ▶ 53:37 Why The AI Bubble Will Burst: The Most Logical Case — With Paul Kedrosky
LATENT SPACE Prediction Not checkable as stated
Kant: Reinforcement learning will move earlier into LLM pre-training
“I have I would say a not commonly held opinion that reinforcement learning will move earlier and earlier into pre-training.”
Eiso Kant Jul 22, 2026 ▶ 45:09 The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
Kant: RL compute cannot scale like pre-training due to task batch constraints
“And RL is batch size constraint, right? So like you are ultimately in your batch size constraint because you don't have infinite tasks, right? When you've got the entire web, you can be much more flexible in scaling up your batch size because you've got the en…”
Eiso Kant Jul 22, 2026 ▶ 1:39:07 The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
20VC Prediction Open · timeframe Jul 2031
Sierra will not pre-train foundation models, leaving capex to major AI labs
“And we're not, you know, we're not doing our own pre-training. We'll leave the capital expense there to, you know, the labs and the larger companies.”
Clay Bavor Jul 4, 2026 ▶ 5:40 Open Models vs Frontier Models: Who Actually Wins? | The $100K Token Budget Every Engineer Will Need · 20VC with Harry Stebbings
LATENT SPACE Disclosure
OpenAI's three research pillars are pre-training, RL, and alignment
“At the very highest level, right, we have an org that focuses on pre-training, right, which is, you know, giving models a lot of world knowledge. We focus on RL, like, teaching the models how to reason with that knowledge, how to chain the little insights toge…”
Mark Chen Jun 25, 2026 ▶ 14:03 Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen
Chen: Pre-training is not dead and remains underrated in AI research
“Well, I think if you still have a pre-training is dead view of the world I think pre-training is definitely yeah, yeah, not, not dead. It's underrated.”
Mark Chen Jun 25, 2026 ▶ 38:01 Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen
CATALYST Assertion Not checkable as stated
AI reinforcement learning energy demand probably already exceeds traditional pre-training
“And this is an area that is becoming huge in terms of energy demand. It'll, it will, The probably already is bigger than what we have historically considered training, you know, pre-training”
Garth Sheldon-Coulson May 28, 2026 ▶ 43:57 Building inference data centers on the high seas
Y COMBINATOR Prediction Not checkable as stated
Hassabis: Current AI paradigms will be part of final AGI architecture
“The components that you just mentioned, I'm pretty sure will be part of the final architecture for AGI. So I think they've come such a long way now and we've proven out so many things about what they can do. I can't see a world in which we will sort of realize…”
Demis Hassabis Apr 29, 2026 ▶ 2:22 Demis Hassabis: Agents, AGI & The Next Big Scientific Breakthrough · Y Combinator
MAD Prediction Not checkable as stated
AI model progress will alternate between pre-training and post-training breakthroughs
“We're going to be having a bit of a swing back and forth between pre-training and post-training.”
Mostafa Dehghani Apr 2, 2026 ▶ 26:55 AI is Already Building AI — Google DeepMind’s Mostafa Dehghani
MAD Insight
Post-training techniques cannot compensate for a weak base AI model
“Pre-training is still the foundation and like, you can never post-train your way out of a week-based model.”
Mostafa Dehghani Apr 2, 2026 ▶ 27:01 AI is Already Building AI — Google DeepMind’s Mostafa Dehghani
MAD Prediction Not checkable as stated
New pre-training techniques will drastically boost base AI model capabilities
“The way that we used to do pre-training, maybe, like, you know, like two, a year ago or two years ago maybe, like, you know, diminishing return is, like, obvious, but I can see how new ideas are bringing, like, you know, fresh, fresh energy into the pre-traini…”
Mostafa Dehghani Apr 2, 2026 ▶ 29:18 AI is Already Building AI — Google DeepMind’s Mostafa Dehghani
Brockman: Pre-Training Capability Multiplies Through the Entire AI Model Pipeline
“Every single step of the model production pipeline multiplies. And so you want to improve all of them. And the thing that we see is we prove the pre-training. It makes all the other steps much easier. And it makes sense because it's a model is able to learn fa…”
Greg Brockman Apr 1, 2026 ▶ 51:56 OpenAI President Greg Brockman: AI Self-Improvement, The Superapp Bet, Path To AGI, Scaling Compute
LATENT SPACE Assertion Not checkable as stated
Lample: Mistral is far from reaching pre-training saturation
“We are still working a lot on the pre-training side. We are very, very far from any sort of situation on the pre-training.”
Guillaume Lample Mar 30, 2026 ▶ 45:23 Mistral: Voxtral TTS, Forge, Leanstral, & Mistral 4 — w/ Pavan Kumar Reddy & Guillaume Lample
LLMs favor CLI tools over APIs due to massive pre-training data volumes
“I think that in pre-training, there's just an enormous amount of command line data. Like even let's ignore, let's like, let's ignore RL. Like you're doing no harness post training. Just the amount of like CLI versus API documentation for just like navigating t…”
Kyle Kranen Mar 8, 2026 ▶ 1:10:19 Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
Core AI capabilities must be built during pre-training, not just fine-tuned
“If there is a core capability that you actually care about, that capability should be part of the foundation and not a fine-tuned artifact.”
Pratyush Maini Feb 10, 2026 ▶ 18:53 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
MAD Insight
Pre-training is no longer where the low-hanging AI gains lie
“Pre-training is not dead, but pre-training is boring. So it's not where the low hanging fruit is anymore.”
Sebastian Raschka Jan 29, 2026 ▶ 17:32 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
MAD Assertion Not checkable as stated
Izmailov: AI researchers cannot reliably trace model behaviors to pre-training sources
“We don't really know what's the source of this type of behaviors, but that's also true for a lot of other behaviors in the models with, like, even the good ones. We don't really, we cannot always pin down, like, where they come from in the pre-training.”
Pavel Izmailov Jan 15, 2026 ▶ 3:57 The Evaluators Are Being Evaluated — Pavel Izmailov (Anthropic/NYU)
MAD Insight
Bourgeau: Architecture and data innovation currently matter more than scale
“The other parts are architecture and data innovation. These also play a really, really important part in the Performance of pre-training and probably even more so than pure scale these days, but scaling is still an important factor as well.”
Sebastien Bourgeau Dec 18, 2025 ▶ 31:22 ”We’re Ahead of Where I Thought We’d Be” — Gemini 3 & the Future of AI
MAD Insight
Kaiser: Reasoning yields far greater AI capability gains per dollar than pre-training
“With the new paradigm of reasoning, you can get much more gains for the same amount of money because it's on this like lower and like, there are just discoveries to be made and these discoveries unlock insane capabilities.”
Łukasz Kaiser Nov 26, 2025 ▶ 4:56 What’s Next for AI? OpenAI’s Łukasz Kaiser (Transformer Co-Author)
MAD Assertion Not checkable as stated
Kaiser: Pre-training consumes the most GPUs of any AI development stage
“Currently, pre-training just uses the most GPUs of all the parts, so it needs the most GPUs, right?”
Łukasz Kaiser Nov 26, 2025 ▶ 32:09 What’s Next for AI? OpenAI’s Łukasz Kaiser (Transformer Co-Author)
MAD Insight
Kaiser: Test-time compute increases AI capabilities faster than pre-training
“Using more tokens to think increases your capability, and it increases it, given the computation, way faster than pre-training, right?”
Łukasz Kaiser Nov 26, 2025 ▶ 46:49 What’s Next for AI? OpenAI’s Łukasz Kaiser (Transformer Co-Author)
MAD Insight
Łukasz Kaiser: AI pre-training expands stored knowledge rather than generalization
“Pre-training is a little different, right? Because it increases the data together with your increase in model size. So it doesn't necessarily increase generalization. It just uses more knowledge.”
Łukasz Kaiser Nov 26, 2025 ▶ 51:47 What’s Next for AI? OpenAI’s Łukasz Kaiser (Transformer Co-Author)
a16z Assertion Not checkable as stated
David Owen: AI pre-training receives less focus due to post-training progress
“It seems as if pre-training is comparatively less of a focus than it was before, partly because, like, you have this exciting new direction of, well, new, newish direction of post-training where they've done so much about reasoning”
Epoch AI Researcher Nov 24, 2025 ▶ 6:16 The 2045 Superintelligence Timeline: Epoch AI’s Data-Driven Forecast
a16z Insight
David Owen: Post-training usage data generates feedback loops for pre-training
“A lot of this stuff is quite synergistic. You develop a better model. You, like, use post-training stuff to make it better. You get a load of data of the model actually being used successfully or not. A lot of that can probably go into pre-training next time.”
Epoch AI Researcher Nov 24, 2025 ▶ 6:41 The 2045 Superintelligence Timeline: Epoch AI’s Data-Driven Forecast
MAD Assertion Not checkable as stated
Kant: Next-token pre-training gains hit a sigmoidal curve and slowed down
“The first paradigm of kind of pre-training of predicting the next token on the web was becoming sigmoidal and was slowing down in terms of the gains that it had.”
Eiso Kant Nov 6, 2025 ▶ 2:39 Intelligence Isn’t Enough: Why Energy & Compute Decide the AGI Race – Eiso Kant
MAD Insight
Smarter pre-training methods will reduce massive AI data spending requirements
“So I think we're just getting smarter about how to do pre-training rather than shoving everything we have into a bucket and like seeing what happens. And so as a result of that, you might not necessarily have to spend the exact same amount of money to get a ca…”
Nathan Benaich Oct 30, 2025 ▶ 46:48 State of AI 2025 with Nathan Benaich: Power Deals, Reasoning Breakthroughs, Real Revenue
TBPN Insight
Nadella: Pre-training Remains More Efficient Than RL Due to Amortization
“Pre-training is a more efficient form of training. Because you can advertise it.”
Satya Nadella Oct 28, 2025 ▶ 11:33 Microsoft CEO Satya Nadella Live on TBPN
MAD Prediction Not checkable as stated
Future AI models will continue to rely on pre-training data
“Personally, I think that's unlikely. Not, not because pre-training is strictly necessary. I think we may well be able to train something completely from scratch, as we've been able to do in other domains, but more because pre-training on this vast data sets th…”
Julian Schrittwieser Oct 23, 2025 ▶ 21:27 Are We Misreading the AI Exponential? Julian Schrittwieser on Move 37 & Scaling RL (Anthropic)
MAD Insight
Schrittwieser: AI pre-training risks over-restricting an agent's exploration search space
“I think the main, you know, the main challenge or the main thing you need to watch out for is that you don't over encode or you don't restrict your search space too much. If your pre-training, if your prior knowledge prevents you from exploring something that …”
Julian Schrittwieser Oct 23, 2025 ▶ 38:41 Are We Misreading the AI Exponential? Julian Schrittwieser on Move 37 & Scaling RL (Anthropic)
MAD Assertion Partly supported
Reinforcement learning scaling yields returns on compute similar to pre-training
“If you look at all the RL literature over time, we see very similar returns on compute in pre-training and in RL, where we can invest exponentially more compute in RL and keep getting benefits.”
Julian Schrittwieser Oct 23, 2025 ▶ 42:01 Are We Misreading the AI Exponential? Julian Schrittwieser on Move 37 & Scaling RL (Anthropic)
Huyen: Internet data is maxed out, making post-training the key AI differentiator
“At some point, we are actually, like, have kind of maxed out on, like, internet data, right? And then people, like, text data, people max out. I think a lot of people are doing, like, with other data, like audios and videos, and, like, everyone's trying to thi…”
Chip Huyen Oct 23, 2025 ▶ 14:58 Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix)
MAD Insight
Tworek: Pre-Training on Unlabeled Data Yields Far More Intelligence Than Supervised Mapping
“There are many more bits usually in the targets than in the labels and studying the structure of targets itself. It yields much more learning and much more intelligence than learning the mapping itself. So like spending a whole compute on just learning the dat…”
Jerry Tworek Oct 16, 2025 ▶ 49:27 How GPT-5 Thinks — OpenAI VP of Research Jerry Tworek
MAD Insight
Jerry Tworek: Pre-training AI models is mathematically simple compared to RL
“The first thing that is important to know and understand, RL is hard. Like, conceptually, if you think about it, and there's still a lot of depth to it, but very conceptually, mathematically speaking, pre-training is dead simple.”
Jerry Tworek Oct 16, 2025 ▶ 53:31 How GPT-5 Thinks — OpenAI VP of Research Jerry Tworek
MAD Insight
Tworek: Reinforcement learning and pre-training require each other to succeed
“And like, I don't like in terms of a pure RL, I don't think like really pure RL makes sense. RL needs Pre-training to be successful. And I think pre-training, as I said before, needs RL to be successful as well.”
Jerry Tworek Oct 16, 2025 ▶ 1:13:14 How GPT-5 Thinks — OpenAI VP of Research Jerry Tworek
Harris: Best AI innovations happen in post-training as data runs out
“It does seem like post training is where the best innovations are happening now and the pre-training and the amount of data, like they've, we've used up a lot of the data. They're trying to create synthetic data to try to improve model performance.”
Parker Harris Oct 14, 2025 ▶ 19:32 Will AI Kill Software? — With Salesforce Co-Founder Parker Harris
Patel: AI models cannot learn external memory usage from pre-training alone
“How do I train it to interact with these databases and these word documents that it writes to? Because it's never gonna learn that from pre-training. Has to learn that from an environment.”
Dylan Patel Sep 30, 2025 ▶ 43:03 Inside the Trillion-Dollar AI Buildout | Dylan Patel Interview · Invest Like The Best
Morcos: Post-training techniques are better applied in pre- and mid-training
“Most of what we do in post-training is better were done in pre and mid training and earlier on in training in general.”
Ari Morcos Aug 29, 2025 ▶ 42:16 Better Data is All You Need — Ari Morcos, Datology
Morcos: Post-training alignment is ineffective long-term compared to pre-training alignment
“Like fundamentally, I think alignment and post training doesn't really make sense as a long-term solution. If you can easily align a model through post training, you can easily misalign a model through post training. If it's easy to put it in, it's easy to tak…”
Ari Morcos Aug 29, 2025 ▶ 53:37 Better Data is All You Need — Ari Morcos, Datology
LENNY'S PODCAST Prediction Open · timeframe Aug 2030
Sharma: Industry spending on AI post-training will eventually surpass pre-training
“Like, I believe we will see, you know, just as much money spent on post-training as we will on pre-training, and in the future, more on post-training.”
Asha Sharma Aug 28, 2025 ▶ 45:03 How 80,000 companies build with AI: Products as organisms and the death of org charts | Asha Sharma
LENNY'S PODCAST Assertion Supported
Lord: AI pre-training gains asymptoted 18 to 24 months ago
“And about 18 months ago, 24 months ago, we started to really see, like, an asymptoting of gains coming from, because they had essentially, like, sucked up all of the knowledge on the internet. And so labs really shifted towards most of the gains now coming fro…”
Garrett Lord Aug 24, 2025 ▶ 6:35 Inside the expert network training every frontier AI model | Garrett Lord
SAASTR Insight
Pre-Training Proprietary Models Does Not Build Moats for AI Application Startups
“Training your own model, certainly pre-training it, not helpful. Yes, you need data to fine tune, but it's not actually a ton of data. And so a lot of people have access to that. And so in many ways the differentiation that I think will continue to compound is…”
Jacob Effron Aug 13, 2025 ▶ 29:16 Redpoint Ventures Playbook: How Top VCs Are Really Investing in AI with Jacob Effron
Kaplan: Compute scaling drives AI progress more than researcher cleverness
“Basically you can Scale up the compute in both pre-training and RL and get better and better performance. And I think that's sort of the fundamental thing that is driving AI progress. It's not that AI researchers are really smart or they suddenly got smart. It…”
Jared Kaplan Jul 29, 2025 ▶ 7:54 Scaling and the Road to Human-Level AI | Anthropic Co-founder Jared Kaplan · Y Combinator
NO PRIORS Insight
Laskin: RL requires far fewer FLOPs than pre-training for frontier models
“We're in this brief period in history right now where the RL flops are still manageable. Like you can really have a best in class product if you're focused. And yes, you'll need to put, you know, you still need a decent amount of GPUs, but from a flops perspec…”
Misha Laskin Jul 17, 2025 ▶ 30:41 No Priors Ep. 123 | With ReflectionAI Co-Founder and CEO Misha Laskin
TBPN Opinion
Patel: Pre-training returns are plateauing; GPT-4.5 was unimpressive and deprecated
“Pre-training seems to have been giving us these plateauing returns. We make these models bigger. GPT-Fort .5 didn't seem to be all that impressive. They had to deprecate it.”
Dwarkesh Patel Jul 8, 2025 ▶ 5:24 DWARKESH PATEL on the Biggest Problem of AI
Kantrowitz: Pre-training data walls justify $100M+ packages for top AI talent
“And this is a strength, a sound strategy because you have everybody talking about how pre-training is hitting diminishing returns. You have everybody talking about how data is hitting a wall. And so what do you need? You just need these algorithmic development…”
Alex Kantrowitz Jul 7, 2025 ▶ 15:15 $100 Million AI Engineers, Vending Machine Claude, Legend Of Soham
BIG TECHNOLOGY Assertion Not checkable as stated
Patel: Pre-training scaling is seeing diminishing returns
“Pre-training, which is this idea that you just make the model bigger that has had diminishing returns.”
Dwarkesh Patel Jun 18, 2025 ▶ 9:13 Dwarkesh Patel: AI Continuous Improvement, Intelligence Explosion, Memory, Frontier Lab Competition
LATENT SPACE Prediction Held up
Ameisen: Deceptive Backward Reasoning Exists in Base Pre-Trained Models
“I bet, I don't know how much I bet a hundred bucks. So somebody can like, they would get a hundred bucks from me if they prove that I'm wrong, that this behavior for a model that does a drink fine tuning, it also does it post pre-training.”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:33:39 The Utility of Interpretability — Emmanuel Amiesen
TBPN Opinion
Kilpatrick: Pre-Training Isn't Dead; Gains Multiply Through Post-Training and RL
“And this is why, like, I don't subscribe to the, like pre-training is, you know, dead and all that stuff, because the more work that you can do at the pre-training level, those capabilities, as you do post-training and as you give the models RL capability, it'…”
Logan Kilpatrick Apr 25, 2025 ▶ 12:50 Google's AI Comeback in Their Own Words - Logan Kilpatrick
Patel: AI models ingest dangerous data during pre-training for world knowledge
“So you don't want to just filter out everything so that the model doesn't know anything about it but at the same time, you don't want it to output, you know, how to build a bomb so there's like a fine balance here, and that's why pre-training is defined as pre…”
Dylan Patel Apr 23, 2025 ▶ 12:57 Generative AI 101: Tokens, Pre-training, Fine-tuning, Reasoning — With SemiAnalysis CEO Dylan Patel
Chen: AI models cannot learn reasoning from scratch without pre-trained knowledge
“You need knowledge in order to build reasoning on top of it. Right. a model can't kind of go in blind and just learn reasoning from scratch. So we find these two paradigms to be fairly complementary and we think, you know, they have feedback loops on each oth…”
Mark Chen Feb 27, 2025 ▶ 4:45 OpenAI's Chief Research Officer on GPT 4.5's Debut, Scaling Laws, And Teaching EQ to Models
BG2 Insight
Patel: AI pre-training gains are becoming logarithmically more expensive
“So, the whole paradigm of training, you know, pre-training is, is, is not slowing down. It's just, it's logarithmically more expensive each, for each generation, for each incremental improvement.”
Dylan Patel Dec 23, 2024 ▶ 40:22 AI Semiconductor Landscape feat. Dylan Patel | BG2 w/ Bill Gurley & Brad Gerstner · Bg2 Pod
NO PRIORS Insight
Fine-Tuning Cannot Effectively Add New Languages to Large Language Models
“There's just no way you can do that without intervening on pre-training. You can't like fine tune or post train Japanese into a model effectively. And so you have to start from scratch.”
Aidan Gomez Nov 21, 2024 ▶ 18:50 No Priors Ep. 91 | With Cohere Co-Founder and CEO Aidan Gomez
BG2 Insight
Huang: AI Inference and Post-Training Are Now Just as Hard as Pre-Training
“People used to think that pre-training was hard, and inference was easy. Now everything is hard.”
Jensen Huang Oct 13, 2024 ▶ 5:38 Ep17. Welcome Jensen Huang | BG2 w/ Bill Gurley & Brad Gerstner · Bg2 Pod
20VC Opinion
Bret Taylor: AI pre-training outside AGI labs wastes capital
“Unless you are an AGI research lab, doing pre-training on a model I believe is just burning capital.”
Bret Taylor Oct 2, 2024 ▶ 0:07 Bret Taylor: Why Pre-Training is for Morons & Companies Will Build Their Own Software | E1209 · 20VC with Harry Stebbings
NO PRIORS Prediction Not checkable as stated
Taylor: AI pre-training will consolidate into a small number of frontier model builders
“I think that will probably play out with the frontier models. We'll end up with a relatively small number of companies doing pre-training you know, which is the really capital intensive part of model building not because, you know, there's not, they're the onl…”
Bret Taylor Sep 19, 2024 ▶ 15:41 No Priors Ep. 82 | With CEO of Sierra Bret Taylor
Howard: Training stages form a continuum allowing deep modification of pre-trained models
“Sorry, it wasn't the end of fine-tuning, but more that we should treat it as a continuum, and we should have much higher expectations of how much you can do with an already trained model. You can really add a lot of behavior to it. You can change its behavior.…”
Jeremy Howard Aug 17, 2024 ▶ 2:04 Answer.ai & AI Magic with Jeremy Howard
NO PRIORS Prediction Not checkable as stated
Vinyals predicts pre-training compute will drop to ~50% as RL expands
“So to me, that balance feels correct, like some on pre-training, and here we, we're trying to learn every task. So certainly that's going to be, you know, let's say it can be as high as 50%, not as high as over 90 like today. And then the rest mostly on reinfo…”
Oriol Vinyals Aug 1, 2024 ▶ 25:29 No Priors Ep. 74 | With Google DeepMind VP of Research Oriol Vinyals
Lambert: RLHF provides a richer signal per unit of compute than pre-training
“As reinforcement learning is so much less compute, like it, Is a richer signal in terms of its impact, because if they could do what RLHF is doing at pre-training, they would, but they don't know how to have that effect in like a stable manner. Otherwise, ever…”
Nathan Lambert Jan 11, 2024 ▶ 22:14 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
NO PRIORS Insight
Liang: Controlling AI hallucination is easier once models understand the concept
“So I think there's pre-training, which is predicting the next word and developing a world model, so to speak. And with those capabilities, then you can, you still have to say don't hallucinate, but it will be much easier to control that model if it has a notio…”
Dr. Percy Liang Apr 25, 2023 ▶ 26:56 No Priors Ep. 7 | With Stanford Professor Dr. Percy Liang

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.