Interleaved multimodal generation overcomes single-shot limits via step-by-step planning
Mostafa Dehghani · AI is Already Building AI — Google DeepMind’s Mostafa Dehghani · Apr 2, 2026 · at 51:37
Mostafa Dehghani, AI research scientist at Google DeepMind, details how interleaved text-and-image generation in Gemini Flash Image (Nano Banana) changes the paradigm from simple text translation to image planning.
“But if you have incremental generation, so if you have text and then an image and text and image, you can get your model to generate these details one by one. So you never expect your model to generate an image, a perfect image in the first shot, right? So, so you expect model, your model to plan about this generation. So it says that, oh, you know, Let me start with big objects because, you know, later I'm going to have a hard time if I put like small objects and the big objects don't fit. Right. So let me just do that. And then like in the next turn, I go with like medium objects and smaller and this like super smart, you know, and you're never bothered by the capability of a single shot image generation because you did planning and then you tune every step difficulty to match the capability of your model to generate a one shot.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →