Insight certainty 3/5 debate potential 2/5

Sanseviero: Embedding offloading suits edge devices; larger models require MoEs or dense architectures

Omar Sanseviero · ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind · May 24, 2026 · at 1:19

Omar Sanseviero of Google DeepMind discusses why Gemma 4's parameter offloading architecture is tailored for edge computing rather than large-scale frontier models.

0:00 / 0:20exact quote · 20.7s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“This is really optimized and designed for, like, on-device. And when I say on-device, I mean, like, running in a phone, Android, Raspberry Pi, and so on, right? When you go larger, you usually want to come back more You want to have more, like, dense architectures or MOEs. So this research, these research decisions were very helpful for this small small use cases.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Omar Sanseviero

Prediction Open · timeframe May 2028
Smartphones will run Gemini 3 Pro-level models locally within two years
“I do think We are heading towards a future in one, two years where imagine like you can run a Gemini three pro powerful model directly in your phone, right?”
Omar Sanseviero May 24, 2026 ▶ 5:51 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Assertion Supported
Sanseviero: AI labs republished model merging techniques previously created on Reddit
“Yeah, like all of the FrankenMoe stuff, like all of the Axolotl library, like all of these tools, and there were papers published by different companies and research labs one or two years later that were rediscovering what was already done by The Reddit or Dis…”
Omar Sanseviero May 24, 2026 ▶ 23:45 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Insight
Sanseviero: Local models handle agentic capabilities well, but world knowledge requires scale
“With local models or models that you can run in your own hardware, you can get capabilities, so you can get agent capabilities, function calling, system instructions, like conversational, and that kind of stuff. Knowledge is much trickier, so for knowledge, yo…”
Omar Sanseviero May 24, 2026 ▶ 5:26 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Insight
Sanseviero: Most Conversational Model Behavior Changes Can Be Done via Prompting
“Just changing how the model behaves, you can do most, most of that via prompting nowadays, and in terms of capabilities, the models are very good out of the box.”
Omar Sanseviero May 24, 2026 ▶ 14:32 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Insight
Sanseviero: MoE models are great for inference but hard to fine-tune
“MOEs are challenging to fine tune. I don't know if we've talked about that in the past, but MOEs in general are like an extremely good architecture. They work great for inference. But when people fine tune them, they struggle a bit. Like they are not as easy t…”
Omar Sanseviero May 24, 2026 ▶ 17:28 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Prediction Not checkable as stated
Sanseviero: Deep architecture research will not be automated within two years
“If you want to do, like, deeper research in the architecture, my hunch is that most likely this will not be, like, automatable, at least in the next one or two years.”
Omar Sanseviero May 24, 2026 ▶ 25:47 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.