Assertion Supported AI assessment confidence: 92% certainty 4/5 debate potential 1/5

Sanseviero: Gemma 4 cannot yet process video and audio simultaneously

Omar Sanseviero · ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind · May 24, 2026 · at 7:28

Omar Sanseviero (Head of Developer Experience at Google DeepMind) details the current multimodal limitations of Gemma 4.

0:00 / 0:14exact quote · 14.8s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“The other thing we do not support yet is video with audio, so we can understand, like, video input or audio input separately, but if you want to pass, like, in the same, from both the visual part and the audio part, we still need to do some improvements around that.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Omar Sanseviero

Prediction Open · timeframe May 2028
Smartphones will run Gemini 3 Pro-level models locally within two years
“I do think We are heading towards a future in one, two years where imagine like you can run a Gemini three pro powerful model directly in your phone, right?”
Omar Sanseviero May 24, 2026 ▶ 5:51 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Assertion Supported
Sanseviero: AI labs republished model merging techniques previously created on Reddit
“Yeah, like all of the FrankenMoe stuff, like all of the Axolotl library, like all of these tools, and there were papers published by different companies and research labs one or two years later that were rediscovering what was already done by The Reddit or Dis…”
Omar Sanseviero May 24, 2026 ▶ 23:45 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Insight
Sanseviero: Local models handle agentic capabilities well, but world knowledge requires scale
“With local models or models that you can run in your own hardware, you can get capabilities, so you can get agent capabilities, function calling, system instructions, like conversational, and that kind of stuff. Knowledge is much trickier, so for knowledge, yo…”
Omar Sanseviero May 24, 2026 ▶ 5:26 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Insight
Sanseviero: Most Conversational Model Behavior Changes Can Be Done via Prompting
“Just changing how the model behaves, you can do most, most of that via prompting nowadays, and in terms of capabilities, the models are very good out of the box.”
Omar Sanseviero May 24, 2026 ▶ 14:32 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Insight
Sanseviero: MoE models are great for inference but hard to fine-tune
“MOEs are challenging to fine tune. I don't know if we've talked about that in the past, but MOEs in general are like an extremely good architecture. They work great for inference. But when people fine tune them, they struggle a bit. Like they are not as easy t…”
Omar Sanseviero May 24, 2026 ▶ 17:28 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Prediction Not checkable as stated
Sanseviero: Deep architecture research will not be automated within two years
“If you want to do, like, deeper research in the architecture, my hunch is that most likely this will not be, like, automatable, at least in the next one or two years.”
Omar Sanseviero May 24, 2026 ▶ 25:47 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.