Demonstrating that image training lowers text perplexity remains extremely difficult
Mostafa Dehghani · AI is Already Building AI — Google DeepMind’s Mostafa Dehghani · Apr 2, 2026 · at 48:54
Mostafa Dehghani, AI researcher at Google DeepMind, explains the technical challenge of achieving positive capability transfer across modalities during Gemini training.
“So it turned out to be a really, really good model, but it was like really hard to see that. Wow. You know, I train on images and then like Text perplexity goes down. That was hard to see. You know, like the fact that, you know, you train in native model and it's good at, like across all the capabilities is already impressive. But my hope is that, you know, multimodality and model is the way to really push multimodal training to enable like positive transfer across modality.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →