Prediction Open · timeframe May 2028
Smartphones will run Gemini 3 Pro-level models locally within two years
“I do think We are heading towards a future in one, two years where imagine like you can run a Gemini three pro powerful model directly in your phone, right?”
Assertion Supported
Sanseviero: AI labs republished model merging techniques previously created on Reddit
“Yeah, like all of the FrankenMoe stuff, like all of the Axolotl library, like all of these tools, and there were papers published by different companies and research labs one or two years later that were rediscovering what was already done by The Reddit or Dis…”
Insight
Sanseviero: Local models handle agentic capabilities well, but world knowledge requires scale
“With local models or models that you can run in your own hardware, you can get capabilities, so you can get agent capabilities, function calling, system instructions, like conversational, and that kind of stuff. Knowledge is much trickier, so for knowledge, yo…”
Insight
Sanseviero: Most Conversational Model Behavior Changes Can Be Done via Prompting
“Just changing how the model behaves, you can do most, most of that via prompting nowadays, and in terms of capabilities, the models are very good out of the box.”
Insight
Sanseviero: MoE models are great for inference but hard to fine-tune
“MOEs are challenging to fine tune. I don't know if we've talked about that in the past, but MOEs in general are like an extremely good architecture. They work great for inference. But when people fine tune them, they struggle a bit. Like they are not as easy t…”
Prediction Not checkable as stated
Sanseviero: Deep architecture research will not be automated within two years
“If you want to do, like, deeper research in the architecture, my hunch is that most likely this will not be, like, automatable, at least in the next one or two years.”
Assertion Supported
Gemma 4 E2B loads only 2B of 5B parameters into GPU
“So the GEMA for model is a E to B. That means that it effectively has two billion parameters loaded into the GPU. It actually has almost five billion parameters, but those three billion parameters can be in the CPU, they can be in the disk, which means that yo…”
Insight
Sanseviero: Embedding offloading suits edge devices; larger models require MoEs or dense architectures
“This is really optimized and designed for, like, on-device. And when I say on-device, I mean, like, running in a phone, Android, Raspberry Pi, and so on, right? When you go larger, you usually want to come back more You want to have more, like, dense architect…”
Opinion
Gemma 4 matches frontier AI state of the art from 18 months ago
“I mean, if you look at Gemma, you compare to how we were one year ago, I would say Gemma four is matching state of the art from one and a half years ago for most. Things.”
Assertion Not checkable as stated
Sanseviero: Gemma 3 outperforms stronger general models when fine-tuned on non-English languages
“If you compare Gemma III to other models from back then, maybe the other models were better than Gemma III like as general model, But if you train all of these models for, I don't know, a specific Southeast Asian language, I don't know, Vietnamese, let's say, …”
Assertion Not checkable as stated
Sanseviero: Text diffusion model quality is still worse than autoregressive models
“I think especially like the model quality is still a bit worse from what you would get from the normal autoregressive model.”
Disclosure
Sanseviero: Many Gemma 4 Launch Partners Skipped Fine-Tuning Due to Base Performance
“For Gemma four, we had 50 To 60 partners. And some of them were like, oh yeah, we're going to try and fine tune the 27 B model for this vision task. And they were like, oh, actually the model works too well out of the box. We don't need to fine tune it. Yeah. …”
Insight
Updating mobile base models breaks per-app LoRAs, creating severe maintenance hurdles
“From a developer point of view, I think it will be very tricky because one, you don't want to have 20 different base models in the phone of the users. The battery will just die. You also don't want to have to update 20 LoRa every time you update the base model…”
Assertion Contradicted
Sanseviero: 31B is the largest quantized model fitting consumer GPUs
“The 31 is really like the largest model size that quantize would fit in a consumer GPU.”
Prediction Not checkable as stated
Sanseviero: Next generation of AI fine-tuners will not write code
“I do think the next generation of fine tuners will not be, I mean, will be people that are not coding at all, right? Like one year ago, we had to write like our own Colab with Transformers or Oncelot or whichever library of your choice. I do think as we like k…”
Assertion Supported
Sanseviero: Gemma 4 is Google's most capable open model yet
“Gemma four is just out. It's the most capable open model we've released so far. We already tried to compact as much intelligence per parameter as we could, bring all of these multimodal capabilities.”
Assertion Supported
Sanseviero: Gemini Nano on Pixel and Samsung phones is built on Gemma
“If you buy a Pixel phone or a high-end Samsung, they come with a Gemini Nano, and Gemini Nano is packed into the operating system, and Gemini Nano is really built on top of Gemma.”
Disclosure
DeepMind's Gemma team runs with two to three PMs and one marketer
“The Gemma team is actually relatively small. We have, like, two or three PMs. We have one marketing person, and then there are, like, engineers and researchers working on shipping this.”
Assertion Supported
Sanseviero: Smaller Gemma 4 models process audio and 30-60 second videos
“Multimodal wise, the smaller models can understand audio Images and short videos, so, 30 to 62nd videos and audios.”
Assertion Supported
Sanseviero: Gemma 4 cannot yet process video and audio simultaneously
“The other thing we do not support yet is video with audio, so we can understand, like, video input or audio input separately, but if you want to pass, like, in the same, from both the visual part and the audio part, we still need to do some improvements around…”
Disclosure
Sanseviero: DeepMind is building agentic tools for research ablations and evaluations
“So for example, within the team, we are building skills to do experiments and ablations and evaluations and how the research team can use all of these agentic tools as part of their research process is also quite interesting.”
Assertion Supported
Sanseviero: Kaggle launched an exam-based benchmark leaderboard for AI agents
“Last week, they released a new system for agent evaluation. It's like a very, like, experimental initial benchmark, but pretty much allowing agents to take an exam and compete in a leaderboard, which is always fun.”
Disclosure
Sanseviero: MedGemma 1.5 Is Gemma 3 Fine-Tuned on Google Medical Datasets
“MedGemma, the last MedGemma, which we released three months ago, MedGemma 1.5, it's based on Gemma three. Gemma three. Yeah, Gemma three. So it's pretty much Gemma three and then additional training with some of our medical data sets.”