Insight certainty 4/5 debate potential 3/5

Bourgeau: AI models must be trained on harmful data to avoid it

Sebastien Bourgeau · ”We’re Ahead of Where I Thought We’d Be” — Gemini 3 & the Future of AI · Dec 18, 2025 · at 43:14

Sebastien Bourgeau, Google DeepMind pre-training lead, explains why frontier AI models cannot completely omit harmful internet content from pre-training datasets.

0:00 / 0:10exact quote · 10.5s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“So at a fundamental level, you did, you do need the model to know about those things. So you have to train a bit at least on those so that it knows what those things are and knows to stay away from those, right?”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Sebastien Bourgeau

Assertion Not checkable as stated
Bourgeau: AI progress from pre-training improvements is not slowing down
“It's still remarkable how much progress we're able to achieve in this way, and it's not really slowing down.”
Sebastien Bourgeau Dec 18, 2025 ▶ 2:26 ”We’re Ahead of Where I Thought We’d Be” — Gemini 3 & the Future of AI
Assertion Not checkable as stated
Bourgeau: Google and DeepMind are actively researching post-Transformer architectures
“I believe so. There's groups doing research on the model architecture side, for sure, within Google and within DeepMind”
Sebastien Bourgeau Dec 18, 2025 ▶ 10:38 ”We’re Ahead of Where I Thought We’d Be” — Gemini 3 & the Future of AI
Insight
Bourgeau: Architecture and data innovation currently matter more than scale
“The other parts are architecture and data innovation. These also play a really, really important part in the Performance of pre-training and probably even more so than pure scale these days, but scaling is still an important factor as well.”
Sebastien Bourgeau Dec 18, 2025 ▶ 31:22 ”We’re Ahead of Where I Thought We’d Be” — Gemini 3 & the Future of AI
Assertion Not checkable as stated
Bourgeau: AI development is not running out of training data
“The other part of your question are we running out of data? I don't think so, so there's more.”
Sebastien Bourgeau Dec 18, 2025 ▶ 34:15 ”We’re Ahead of Where I Thought We’d Be” — Gemini 3 & the Future of AI
Insight
Bourgeau: AI research is shifting to a data-limited paradigm
“I think what might be happening instead is kind of a shift in paradigm where before we were kind of scaling in the data unlimited regime where, where data would scale as much as you would like. And we're kind of shifting more to a data limited regime, which ac…”
Sebastien Bourgeau Dec 18, 2025 ▶ 34:26 ”We’re Ahead of Where I Thought We’d Be” — Gemini 3 & the Future of AI
Prediction Not checkable as stated
Bourgeau: End-to-end differentiable retrieval and search in training will take years
“I think deep down, I do believe that the long-term answer is to learn this differentiable end-to-end way, which means probably doing pre-training or whatever that looks like in the future, Learn to retrieve as part of the training and learn how to do search as…”
Sebastien Bourgeau Dec 18, 2025 ▶ 40:09 ”We’re Ahead of Where I Thought We’d Be” — Gemini 3 & the Future of AI
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.