Omar Sanseviero

Developer Experience Lead, Google DeepMind · 1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

operatorengineerexecutive@osanseviero ↗LinkedIn ↗osanseviero.github.io/hackerllama ↗

Omar Sanseviero leads developer experience and adoption at Google DeepMind across products like Google AI Studio, the Gemini API, and Gemma models. Previously, he served as Head of Platform and Community at Hugging Face, scaling Hugging Face Spaces and directing developer advocacy across the open-source machine learning ecosystem.

23statements → 13claims → 8claims resolved → 88%fully supported → 3.7/5average certainty → 1.78/5average debate potential → ≈4.0/5argument clarity, estimated →

7 supported 0 partly supported 1 contradicted 1 not yet assessed 4 not checkable as stated how the 13 claims stand · each chip opens the sources

3 predictions · 10 assertions · 1 opinion · 5 insights · 4 disclosures · every statement was checked. The predictions and assertions are the 13 claims: statements the public record can support or contradict. 8 are resolved, 1 is not yet assessed, and 4 name no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Omar argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
Sanseviero: AI labs republished model merging techniques previously created on Reddit
“Yeah, like all of the FrankenMoe stuff, like all of the Axolotl library, like all of these tools, and there were papers published by different companies and research labs one or two years later that were rediscovering what was already done by The Reddit or Dis…”
Omar Sanseviero May 24, 2026 ▶ 23:45 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind

Their most notable contradicted claim

Assertion Contradicted
Sanseviero: 31B is the largest quantized model fitting consumer GPUs
“The 31 is really like the largest model size that quantize would fit in a consumer GPU.”
Omar Sanseviero May 24, 2026 ▶ 17:17 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind

How they sound: not measured why? →

We measure speaking style by listening to the audio itself, and a fair number needs at least 2,000 words from one person on tape we have measured. There is too little of Omar Sanseviero on measured tape to publish a rate. This says nothing about how they speak.

Everything Omar Sanseviero said on Latent Space that made the record, most notable first. Filter by type, assessment or year in the ledger →

Prediction Open · timeframe May 2028
Smartphones will run Gemini 3 Pro-level models locally within two years
“I do think We are heading towards a future in one, two years where imagine like you can run a Gemini three pro powerful model directly in your phone, right?”
Omar Sanseviero May 24, 2026 ▶ 5:51 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Assertion Supported
Sanseviero: AI labs republished model merging techniques previously created on Reddit
“Yeah, like all of the FrankenMoe stuff, like all of the Axolotl library, like all of these tools, and there were papers published by different companies and research labs one or two years later that were rediscovering what was already done by The Reddit or Dis…”
Omar Sanseviero May 24, 2026 ▶ 23:45 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Insight
Sanseviero: Local models handle agentic capabilities well, but world knowledge requires scale
“With local models or models that you can run in your own hardware, you can get capabilities, so you can get agent capabilities, function calling, system instructions, like conversational, and that kind of stuff. Knowledge is much trickier, so for knowledge, yo…”
Omar Sanseviero May 24, 2026 ▶ 5:26 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Insight
Sanseviero: Most Conversational Model Behavior Changes Can Be Done via Prompting
“Just changing how the model behaves, you can do most, most of that via prompting nowadays, and in terms of capabilities, the models are very good out of the box.”
Omar Sanseviero May 24, 2026 ▶ 14:32 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Insight
Sanseviero: MoE models are great for inference but hard to fine-tune
“MOEs are challenging to fine tune. I don't know if we've talked about that in the past, but MOEs in general are like an extremely good architecture. They work great for inference. But when people fine tune them, they struggle a bit. Like they are not as easy t…”
Omar Sanseviero May 24, 2026 ▶ 17:28 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Prediction Not checkable as stated
Sanseviero: Deep architecture research will not be automated within two years
“If you want to do, like, deeper research in the architecture, my hunch is that most likely this will not be, like, automatable, at least in the next one or two years.”
Omar Sanseviero May 24, 2026 ▶ 25:47 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Assertion Supported
Gemma 4 E2B loads only 2B of 5B parameters into GPU
“So the GEMA for model is a E to B. That means that it effectively has two billion parameters loaded into the GPU. It actually has almost five billion parameters, but those three billion parameters can be in the CPU, they can be in the disk, which means that yo…”
Omar Sanseviero May 24, 2026 ▶ 0:52 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Insight
Sanseviero: Embedding offloading suits edge devices; larger models require MoEs or dense architectures
“This is really optimized and designed for, like, on-device. And when I say on-device, I mean, like, running in a phone, Android, Raspberry Pi, and so on, right? When you go larger, you usually want to come back more You want to have more, like, dense architect…”
Omar Sanseviero May 24, 2026 ▶ 1:19 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Opinion
Gemma 4 matches frontier AI state of the art from 18 months ago
“I mean, if you look at Gemma, you compare to how we were one year ago, I would say Gemma four is matching state of the art from one and a half years ago for most. Things.”
Omar Sanseviero May 24, 2026 ▶ 5:16 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Assertion Not checkable as stated
Sanseviero: Gemma 3 outperforms stronger general models when fine-tuned on non-English languages
“If you compare Gemma III to other models from back then, maybe the other models were better than Gemma III like as general model, But if you train all of these models for, I don't know, a specific Southeast Asian language, I don't know, Vietnamese, let's say, …”
Omar Sanseviero May 24, 2026 ▶ 8:50 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Assertion Not checkable as stated
Sanseviero: Text diffusion model quality is still worse than autoregressive models
“I think especially like the model quality is still a bit worse from what you would get from the normal autoregressive model.”
Omar Sanseviero May 24, 2026 ▶ 12:40 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Disclosure
Sanseviero: Many Gemma 4 Launch Partners Skipped Fine-Tuning Due to Base Performance
“For Gemma four, we had 50 To 60 partners. And some of them were like, oh yeah, we're going to try and fine tune the 27 B model for this vision task. And they were like, oh, actually the model works too well out of the box. We don't need to fine tune it. Yeah. …”
Omar Sanseviero May 24, 2026 ▶ 14:01 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Insight
Updating mobile base models breaks per-app LoRAs, creating severe maintenance hurdles
“From a developer point of view, I think it will be very tricky because one, you don't want to have 20 different base models in the phone of the users. The battery will just die. You also don't want to have to update 20 LoRa every time you update the base model…”
Omar Sanseviero May 24, 2026 ▶ 15:56 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Assertion Contradicted
Sanseviero: 31B is the largest quantized model fitting consumer GPUs
“The 31 is really like the largest model size that quantize would fit in a consumer GPU.”
Omar Sanseviero May 24, 2026 ▶ 17:17 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Prediction Not checkable as stated
Sanseviero: Next generation of AI fine-tuners will not write code
“I do think the next generation of fine tuners will not be, I mean, will be people that are not coding at all, right? Like one year ago, we had to write like our own Colab with Transformers or Oncelot or whichever library of your choice. I do think as we like k…”
Omar Sanseviero May 24, 2026 ▶ 25:14 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Assertion Supported
Sanseviero: Gemma 4 is Google's most capable open model yet
“Gemma four is just out. It's the most capable open model we've released so far. We already tried to compact as much intelligence per parameter as we could, bring all of these multimodal capabilities.”
Omar Sanseviero May 24, 2026 ▶ 0:10 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Assertion Supported
Sanseviero: Gemini Nano on Pixel and Samsung phones is built on Gemma
“If you buy a Pixel phone or a high-end Samsung, they come with a Gemini Nano, and Gemini Nano is packed into the operating system, and Gemini Nano is really built on top of Gemma.”
Omar Sanseviero May 24, 2026 ▶ 2:13 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Disclosure
DeepMind's Gemma team runs with two to three PMs and one marketer
“The Gemma team is actually relatively small. We have, like, two or three PMs. We have one marketing person, and then there are, like, engineers and researchers working on shipping this.”
Omar Sanseviero May 24, 2026 ▶ 3:21 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Assertion Supported
Sanseviero: Smaller Gemma 4 models process audio and 30-60 second videos
“Multimodal wise, the smaller models can understand audio Images and short videos, so, 30 to 62nd videos and audios.”
Omar Sanseviero May 24, 2026 ▶ 6:43 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Assertion Supported
Sanseviero: Gemma 4 cannot yet process video and audio simultaneously
“The other thing we do not support yet is video with audio, so we can understand, like, video input or audio input separately, but if you want to pass, like, in the same, from both the visual part and the audio part, we still need to do some improvements around…”
Omar Sanseviero May 24, 2026 ▶ 7:28 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Disclosure
Sanseviero: DeepMind is building agentic tools for research ablations and evaluations
“So for example, within the team, we are building skills to do experiments and ablations and evaluations and how the research team can use all of these agentic tools as part of their research process is also quite interesting.”
Omar Sanseviero May 24, 2026 ▶ 22:44 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Assertion Supported
Sanseviero: Kaggle launched an exam-based benchmark leaderboard for AI agents
“Last week, they released a new system for agent evaluation. It's like a very, like, experimental initial benchmark, but pretty much allowing agents to take an exam and compete in a leaderboard, which is always fun.”
Omar Sanseviero May 24, 2026 ▶ 28:31 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Disclosure
Sanseviero: MedGemma 1.5 Is Gemma 3 Fine-Tuned on Google Medical Datasets
“MedGemma, the last MedGemma, which we released three months ago, MedGemma 1.5, it's based on Gemma three. Gemma three. Yeah, Gemma three. So it's pretty much Gemma three and then additional training with some of our medical data sets.”
Omar Sanseviero May 24, 2026 ▶ 15:08 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind

Appearances (1)

EpisodeDateSpeaking time
⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind May 24, 2026 18m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.