May 24, 2026 · 29m · latent-space

⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind

Omar Sanseviero · 18m spoken Shawn Wang · 5m spoken Alessio Fanelli · 3m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this interview, Google DeepMind's Omar Sanseviero joins hosts Shawn 'swyx' Wang and Alessio Fanelli to explore the architecture, multimodality, and open ecosystem strategy behind Gemma 4 and Gemini research. The discussion details breakthroughs in on-device parameter efficiency, mechanistic interpretability, text diffusion, and the emerging frontier of autonomous agent-driven research.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 33.2% of the talking time here. How this is scored →

The hosts as informed peer 5.8 Guest teaching 2.9 Guest disagreement 1.1 The hosts pushing back 2.6
05100:0010:0020:000:04–3:13 · The hosts as informed peer 6/10 Gemma 4 Architecture and Effective Parameters Breakdown Alessio and Swyx demonstrate familiarity with parameter offloading and the 3N architecture paper, but Omar educates them on how Gemma 4 implements per-layer lookup tables to decouple disk/CPU storage from GPU active memory.3:15–6:26 · The hosts as informed peer 5/10 Shipping Gemma 4 Ecosystem Integrations and Cloud Comparison Alessio presses on product positioning, questioning whether local models cannibalize larger Gemini API endpoints, prompting Omar to delineate capability versus world knowledge limits.6:27–9:36 · The hosts as informed peer 6/10 Gemma 4 Multimodality and Multilingual Tokenizer Enhancements The hosts probe into multimodal nuances like video-audio interleaving and tokenizers, with Omar explaining how the Gemini-inherited multilingual tokenizer outperforms competing baselines on lower-resource languages.9:36–13:36 · The hosts as informed peer 6/10 DeepMind London Research and Text Diffusion Models Swyx proposes a dual system 1 and system 2 architecture for diffusion and autoregressive text models, while Omar tempers expectations regarding text diffusion quality and fine-tunability.13:37–16:28 · The hosts as informed peer 6/10 Evolution of Fine-Tuning and On-Device LoRA Constraints Alessio brings up Apple's multi-LoRA on-device strategy, prompting Omar to explain the developer maintenance burden and battery life trade-offs of managing numerous LoRA adapters across base model updates.16:28–20:08 · The hosts as informed peer 7/10 Dense Versus MoE Architectures and Parameter Efficiency Limits The hosts challenge conventional wisdom around why MoE models are supposedly harder to fine-tune than dense models, pushing Omar on routing mechanics and information superposition.20:08–24:01 · The hosts as informed peer 6/10 Mechanistic Interpretability and Open Research Collaboration Swyx discusses why applied engineers care about mechanistic interpretability and GemmaScope, while Omar agrees and highlights how open-source community experiments often predate formal lab papers.24:01–26:05 · The hosts as informed peer 6/10 Auto-Research Frontiers and Agent-Driven Model Tuning Alessio pushes back against the premise that auto-research is merely automated hyperparameter search, arguing that true breakthrough potential lies in AI generating novel Move 37 research trajectories.26:05–29:49 · The hosts as informed peer 4/10 Global DevRel Expansion and Kaggle Benchmark Integration A relaxed closing segment where Swyx and Alessio ask about DeepMind's DevRel expansion in Singapore and the integration of Kaggle benchmark ecosystems for model evaluations.0:04–3:13 · Guest teaching 5/10 Gemma 4 Architecture and Effective Parameters Breakdown Alessio and Swyx demonstrate familiarity with parameter offloading and the 3N architecture paper, but Omar educates them on how Gemma 4 implements per-layer lookup tables to decouple disk/CPU storage from GPU active memory.3:15–6:26 · Guest teaching 3/10 Shipping Gemma 4 Ecosystem Integrations and Cloud Comparison Alessio presses on product positioning, questioning whether local models cannibalize larger Gemini API endpoints, prompting Omar to delineate capability versus world knowledge limits.6:27–9:36 · Guest teaching 4/10 Gemma 4 Multimodality and Multilingual Tokenizer Enhancements The hosts probe into multimodal nuances like video-audio interleaving and tokenizers, with Omar explaining how the Gemini-inherited multilingual tokenizer outperforms competing baselines on lower-resource languages.9:36–13:36 · Guest teaching 2/10 DeepMind London Research and Text Diffusion Models Swyx proposes a dual system 1 and system 2 architecture for diffusion and autoregressive text models, while Omar tempers expectations regarding text diffusion quality and fine-tunability.13:37–16:28 · Guest teaching 3/10 Evolution of Fine-Tuning and On-Device LoRA Constraints Alessio brings up Apple's multi-LoRA on-device strategy, prompting Omar to explain the developer maintenance burden and battery life trade-offs of managing numerous LoRA adapters across base model updates.16:28–20:08 · Guest teaching 3/10 Dense Versus MoE Architectures and Parameter Efficiency Limits The hosts challenge conventional wisdom around why MoE models are supposedly harder to fine-tune than dense models, pushing Omar on routing mechanics and information superposition.20:08–24:01 · Guest teaching 2/10 Mechanistic Interpretability and Open Research Collaboration Swyx discusses why applied engineers care about mechanistic interpretability and GemmaScope, while Omar agrees and highlights how open-source community experiments often predate formal lab papers.24:01–26:05 · Guest teaching 2/10 Auto-Research Frontiers and Agent-Driven Model Tuning Alessio pushes back against the premise that auto-research is merely automated hyperparameter search, arguing that true breakthrough potential lies in AI generating novel Move 37 research trajectories.26:05–29:49 · Guest teaching 2/10 Global DevRel Expansion and Kaggle Benchmark Integration A relaxed closing segment where Swyx and Alessio ask about DeepMind's DevRel expansion in Singapore and the integration of Kaggle benchmark ecosystems for model evaluations.0:04–3:13 · Guest disagreement 1/10 Gemma 4 Architecture and Effective Parameters Breakdown Alessio and Swyx demonstrate familiarity with parameter offloading and the 3N architecture paper, but Omar educates them on how Gemma 4 implements per-layer lookup tables to decouple disk/CPU storage from GPU active memory.3:15–6:26 · Guest disagreement 1/10 Shipping Gemma 4 Ecosystem Integrations and Cloud Comparison Alessio presses on product positioning, questioning whether local models cannibalize larger Gemini API endpoints, prompting Omar to delineate capability versus world knowledge limits.6:27–9:36 · Guest disagreement 1/10 Gemma 4 Multimodality and Multilingual Tokenizer Enhancements The hosts probe into multimodal nuances like video-audio interleaving and tokenizers, with Omar explaining how the Gemini-inherited multilingual tokenizer outperforms competing baselines on lower-resource languages.9:36–13:36 · Guest disagreement 2/10 DeepMind London Research and Text Diffusion Models Swyx proposes a dual system 1 and system 2 architecture for diffusion and autoregressive text models, while Omar tempers expectations regarding text diffusion quality and fine-tunability.13:37–16:28 · Guest disagreement 1/10 Evolution of Fine-Tuning and On-Device LoRA Constraints Alessio brings up Apple's multi-LoRA on-device strategy, prompting Omar to explain the developer maintenance burden and battery life trade-offs of managing numerous LoRA adapters across base model updates.16:28–20:08 · Guest disagreement 2/10 Dense Versus MoE Architectures and Parameter Efficiency Limits The hosts challenge conventional wisdom around why MoE models are supposedly harder to fine-tune than dense models, pushing Omar on routing mechanics and information superposition.20:08–24:01 · Guest disagreement 1/10 Mechanistic Interpretability and Open Research Collaboration Swyx discusses why applied engineers care about mechanistic interpretability and GemmaScope, while Omar agrees and highlights how open-source community experiments often predate formal lab papers.24:01–26:05 · Guest disagreement 1/10 Auto-Research Frontiers and Agent-Driven Model Tuning Alessio pushes back against the premise that auto-research is merely automated hyperparameter search, arguing that true breakthrough potential lies in AI generating novel Move 37 research trajectories.26:05–29:49 · Guest disagreement 0/10 Global DevRel Expansion and Kaggle Benchmark Integration A relaxed closing segment where Swyx and Alessio ask about DeepMind's DevRel expansion in Singapore and the integration of Kaggle benchmark ecosystems for model evaluations.0:04–3:13 · The hosts pushing back 2/10 Gemma 4 Architecture and Effective Parameters Breakdown Alessio and Swyx demonstrate familiarity with parameter offloading and the 3N architecture paper, but Omar educates them on how Gemma 4 implements per-layer lookup tables to decouple disk/CPU storage from GPU active memory.3:15–6:26 · The hosts pushing back 3/10 Shipping Gemma 4 Ecosystem Integrations and Cloud Comparison Alessio presses on product positioning, questioning whether local models cannibalize larger Gemini API endpoints, prompting Omar to delineate capability versus world knowledge limits.6:27–9:36 · The hosts pushing back 2/10 Gemma 4 Multimodality and Multilingual Tokenizer Enhancements The hosts probe into multimodal nuances like video-audio interleaving and tokenizers, with Omar explaining how the Gemini-inherited multilingual tokenizer outperforms competing baselines on lower-resource languages.9:36–13:36 · The hosts pushing back 3/10 DeepMind London Research and Text Diffusion Models Swyx proposes a dual system 1 and system 2 architecture for diffusion and autoregressive text models, while Omar tempers expectations regarding text diffusion quality and fine-tunability.13:37–16:28 · The hosts pushing back 2/10 Evolution of Fine-Tuning and On-Device LoRA Constraints Alessio brings up Apple's multi-LoRA on-device strategy, prompting Omar to explain the developer maintenance burden and battery life trade-offs of managing numerous LoRA adapters across base model updates.16:28–20:08 · The hosts pushing back 4/10 Dense Versus MoE Architectures and Parameter Efficiency Limits The hosts challenge conventional wisdom around why MoE models are supposedly harder to fine-tune than dense models, pushing Omar on routing mechanics and information superposition.20:08–24:01 · The hosts pushing back 2/10 Mechanistic Interpretability and Open Research Collaboration Swyx discusses why applied engineers care about mechanistic interpretability and GemmaScope, while Omar agrees and highlights how open-source community experiments often predate formal lab papers.24:01–26:05 · The hosts pushing back 4/10 Auto-Research Frontiers and Agent-Driven Model Tuning Alessio pushes back against the premise that auto-research is merely automated hyperparameter search, arguing that true breakthrough potential lies in AI generating novel Move 37 research trajectories.26:05–29:49 · The hosts pushing back 1/10 Global DevRel Expansion and Kaggle Benchmark Integration A relaxed closing segment where Swyx and Alessio ask about DeepMind's DevRel expansion in Singapore and the integration of Kaggle benchmark ecosystems for model evaluations.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 38.6% · guest 61.4%0:00 · the hosts 38.6% · guest 61.4%3:00 · the hosts 27.2% · guest 72.8%3:00 · the hosts 27.2% · guest 72.8%6:00 · the hosts 20.1% · guest 79.9%6:00 · the hosts 20.1% · guest 79.9%9:00 · the hosts 39.6% · guest 60.4%9:00 · the hosts 39.6% · guest 60.4%12:00 · the hosts 30% · guest 70%12:00 · the hosts 30% · guest 70%15:00 · the hosts 34.4% · guest 65.6%15:00 · the hosts 34.4% · guest 65.6%18:00 · the hosts 31.5% · guest 68.5%18:00 · the hosts 31.5% · guest 68.5%21:00 · the hosts 49.4% · guest 50.6%21:00 · the hosts 49.4% · guest 50.6%24:00 · the hosts 45.5% · guest 54.5%24:00 · the hosts 45.5% · guest 54.5%27:00 · the hosts 14.3% · guest 85.7%27:00 · the hosts 14.3% · guest 85.7%
Sharpest disagreement ▶ 17:47 Omar defends MoE fine-tuning nuances against host skepticism

When hosts dismiss MoE fine-tuning difficulties as 'just training,' Omar pushes back by explaining how router distribution shifts and activation dynamics create real empirical barriers.

Hardest push from the hosts ▶ 24:38 Alessio rejects the framing of auto-research as mere automation

Alessio explicitly rejects Swyx and Omar's modest view of auto-research, asserting that the true frontier is finding unpredictable Move 37 discoveries rather than routine script running.

Biggest teaching moment ▶ 0:31 Omar breaks down per-layer lookup tables versus active parameters

Omar provides a clear technical breakdown of how Gemma 4 utilizes per-layer embedding lookup tables to keep effective parameters on CPU/disk without full GPU matrix multiplication.

The host holds their own ▶ 17:47 Hosts challenge MoE fine-tuning orthodoxy

Swyx and Alessio assert their technical grasp of backpropagation, pressing on why MoE routing would inherently prevent fine-tuning if initial pretraining succeeded.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Gemma 4 Architecture and Effective Parameters Breakdown 6512 Alessio and Swyx demonstrate familiarity with parameter offloading and the 3N architecture paper, but Omar educates them on how Gemma 4 implements per-layer lookup tables to decouple disk/CPU storage from GPU active memory.
Shipping Gemma 4 Ecosystem Integrations and Cloud Comparison 5313 Alessio presses on product positioning, questioning whether local models cannibalize larger Gemini API endpoints, prompting Omar to delineate capability versus world knowledge limits.
Gemma 4 Multimodality and Multilingual Tokenizer Enhancements 6412 The hosts probe into multimodal nuances like video-audio interleaving and tokenizers, with Omar explaining how the Gemini-inherited multilingual tokenizer outperforms competing baselines on lower-resource languages.
DeepMind London Research and Text Diffusion Models 6223 Swyx proposes a dual system 1 and system 2 architecture for diffusion and autoregressive text models, while Omar tempers expectations regarding text diffusion quality and fine-tunability.
Evolution of Fine-Tuning and On-Device LoRA Constraints 6312 Alessio brings up Apple's multi-LoRA on-device strategy, prompting Omar to explain the developer maintenance burden and battery life trade-offs of managing numerous LoRA adapters across base model updates.
Dense Versus MoE Architectures and Parameter Efficiency Limits 7324 The hosts challenge conventional wisdom around why MoE models are supposedly harder to fine-tune than dense models, pushing Omar on routing mechanics and information superposition.
Mechanistic Interpretability and Open Research Collaboration 6212 Swyx discusses why applied engineers care about mechanistic interpretability and GemmaScope, while Omar agrees and highlights how open-source community experiments often predate formal lab papers.
Auto-Research Frontiers and Agent-Driven Model Tuning 6214 Alessio pushes back against the premise that auto-research is merely automated hyperparameter search, arguing that true breakthrough potential lies in AI generating novel Move 37 research trajectories.
Global DevRel Expansion and Kaggle Benchmark Integration 4201 A relaxed closing segment where Swyx and Alessio ask about DeepMind's DevRel expansion in Singapore and the integration of Kaggle benchmark ecosystems for model evaluations.

Statements from this episode (24)

Assertion Supported
Sanseviero: Gemma 4 is Google's most capable open model yet
“Gemma four is just out. It's the most capable open model we've released so far. We already tried to compact as much intelligence per parameter as we could, bring all of these multimodal capabilities.”
Omar Sanseviero May 24, 2026 ▶ 0:10
Assertion Supported
Gemma 4 E2B loads only 2B of 5B parameters into GPU
“So the GEMA for model is a E to B. That means that it effectively has two billion parameters loaded into the GPU. It actually has almost five billion parameters, but those three billion parameters can be in the CPU, they can be in the disk, which means that yo…”
Omar Sanseviero May 24, 2026 ▶ 0:52
Insight
Sanseviero: Embedding offloading suits edge devices; larger models require MoEs or dense architectures
“This is really optimized and designed for, like, on-device. And when I say on-device, I mean, like, running in a phone, Android, Raspberry Pi, and so on, right? When you go larger, you usually want to come back more You want to have more, like, dense architect…”
Omar Sanseviero May 24, 2026 ▶ 1:19
Assertion Supported
Sanseviero: Gemini Nano on Pixel and Samsung phones is built on Gemma
“If you buy a Pixel phone or a high-end Samsung, they come with a Gemini Nano, and Gemini Nano is packed into the operating system, and Gemini Nano is really built on top of Gemma.”
Omar Sanseviero May 24, 2026 ▶ 2:13
Disclosure
DeepMind's Gemma team runs with two to three PMs and one marketer
“The Gemma team is actually relatively small. We have, like, two or three PMs. We have one marketing person, and then there are, like, engineers and researchers working on shipping this.”
Omar Sanseviero May 24, 2026 ▶ 3:21
Opinion
Gemma 4 matches frontier AI state of the art from 18 months ago
“I mean, if you look at Gemma, you compare to how we were one year ago, I would say Gemma four is matching state of the art from one and a half years ago for most. Things.”
Omar Sanseviero May 24, 2026 ▶ 5:16
Insight
Sanseviero: Local models handle agentic capabilities well, but world knowledge requires scale
“With local models or models that you can run in your own hardware, you can get capabilities, so you can get agent capabilities, function calling, system instructions, like conversational, and that kind of stuff. Knowledge is much trickier, so for knowledge, yo…”
Omar Sanseviero May 24, 2026 ▶ 5:26
Prediction Open · timeframe May 2028
Smartphones will run Gemini 3 Pro-level models locally within two years
“I do think We are heading towards a future in one, two years where imagine like you can run a Gemini three pro powerful model directly in your phone, right?”
Omar Sanseviero May 24, 2026 ▶ 5:51
Assertion Supported
Sanseviero: Smaller Gemma 4 models process audio and 30-60 second videos
“Multimodal wise, the smaller models can understand audio Images and short videos, so, 30 to 62nd videos and audios.”
Omar Sanseviero May 24, 2026 ▶ 6:43
Assertion Supported
Sanseviero: Gemma 4 cannot yet process video and audio simultaneously
“The other thing we do not support yet is video with audio, so we can understand, like, video input or audio input separately, but if you want to pass, like, in the same, from both the visual part and the audio part, we still need to do some improvements around…”
Omar Sanseviero May 24, 2026 ▶ 7:28
Assertion Not checkable as stated
Sanseviero: Gemma 3 outperforms stronger general models when fine-tuned on non-English languages
“If you compare Gemma III to other models from back then, maybe the other models were better than Gemma III like as general model, But if you train all of these models for, I don't know, a specific Southeast Asian language, I don't know, Vietnamese, let's say, …”
Omar Sanseviero May 24, 2026 ▶ 8:50
Assertion Not checkable as stated
Sanseviero: Text diffusion model quality is still worse than autoregressive models
“I think especially like the model quality is still a bit worse from what you would get from the normal autoregressive model.”
Omar Sanseviero May 24, 2026 ▶ 12:40
Disclosure
Sanseviero: Many Gemma 4 Launch Partners Skipped Fine-Tuning Due to Base Performance
“For Gemma four, we had 50 To 60 partners. And some of them were like, oh yeah, we're going to try and fine tune the 27 B model for this vision task. And they were like, oh, actually the model works too well out of the box. We don't need to fine tune it. Yeah. …”
Omar Sanseviero May 24, 2026 ▶ 14:01
Insight
Sanseviero: Most Conversational Model Behavior Changes Can Be Done via Prompting
“Just changing how the model behaves, you can do most, most of that via prompting nowadays, and in terms of capabilities, the models are very good out of the box.”
Omar Sanseviero May 24, 2026 ▶ 14:32
Disclosure
Sanseviero: MedGemma 1.5 Is Gemma 3 Fine-Tuned on Google Medical Datasets
“MedGemma, the last MedGemma, which we released three months ago, MedGemma 1.5, it's based on Gemma three. Gemma three. Yeah, Gemma three. So it's pretty much Gemma three and then additional training with some of our medical data sets.”
Omar Sanseviero May 24, 2026 ▶ 15:08
Insight
Updating mobile base models breaks per-app LoRAs, creating severe maintenance hurdles
“From a developer point of view, I think it will be very tricky because one, you don't want to have 20 different base models in the phone of the users. The battery will just die. You also don't want to have to update 20 LoRa every time you update the base model…”
Omar Sanseviero May 24, 2026 ▶ 15:56
Assertion Contradicted
Sanseviero: 31B is the largest quantized model fitting consumer GPUs
“The 31 is really like the largest model size that quantize would fit in a consumer GPU.”
Omar Sanseviero May 24, 2026 ▶ 17:17
Insight
Sanseviero: MoE models are great for inference but hard to fine-tune
“MOEs are challenging to fine tune. I don't know if we've talked about that in the past, but MOEs in general are like an extremely good architecture. They work great for inference. But when people fine tune them, they struggle a bit. Like they are not as easy t…”
Omar Sanseviero May 24, 2026 ▶ 17:28
Opinion
Wang: Mechanistic interpretability is easiest path to research for engineers
“And Macinterp is probably the easiest, single easiest way that engineers can get into research if they want to.”
Shawn Wang May 24, 2026 ▶ 21:50
Disclosure
Sanseviero: DeepMind is building agentic tools for research ablations and evaluations
“So for example, within the team, we are building skills to do experiments and ablations and evaluations and how the research team can use all of these agentic tools as part of their research process is also quite interesting.”
Omar Sanseviero May 24, 2026 ▶ 22:44
Assertion Supported
Sanseviero: AI labs republished model merging techniques previously created on Reddit
“Yeah, like all of the FrankenMoe stuff, like all of the Axolotl library, like all of these tools, and there were papers published by different companies and research labs one or two years later that were rediscovering what was already done by The Reddit or Dis…”
Omar Sanseviero May 24, 2026 ▶ 23:45
Prediction Not checkable as stated
Sanseviero: Next generation of AI fine-tuners will not write code
“I do think the next generation of fine tuners will not be, I mean, will be people that are not coding at all, right? Like one year ago, we had to write like our own Colab with Transformers or Oncelot or whichever library of your choice. I do think as we like k…”
Omar Sanseviero May 24, 2026 ▶ 25:14
Prediction Not checkable as stated
Sanseviero: Deep architecture research will not be automated within two years
“If you want to do, like, deeper research in the architecture, my hunch is that most likely this will not be, like, automatable, at least in the next one or two years.”
Omar Sanseviero May 24, 2026 ▶ 25:47
Assertion Supported
Sanseviero: Kaggle launched an exam-based benchmark leaderboard for AI agents
“Last week, they released a new system for agent evaluation. It's like a very, like, experimental initial benchmark, but pretty much allowing agents to take an exam and compete in a leaderboard, which is always fun.”
Omar Sanseviero May 24, 2026 ▶ 28:31
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.