Oct 28, 2025 · 54m · a16z

Google DeepMind Developers: How Nano Banana Was Made

Oliver Wang · 18m spoken Nicole Brichtova · 14m spoken Yoko Li · 7m spoken Justine Moore · 4m spoken Guido Appenzeller · 4m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The a16z Podcast, hosts Guido Appenzeller and Yoko Li interview Google DeepMind's Nicole Brichtova and Oliver Wang about the development and capabilities of the 'Nano Banana' image generation model, exploring its technical breakthroughs, interface designs, and creative impact.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The host as informed peer 4.7 Guest teaching 3.9 Guest disagreement 1.5 The host pushing back 1.7
05100:0015:0030:0045:002:07–4:51 · The host as informed peer 2/10 Viral Realizations and Personal Avatars Host asks conversational opening questions about viral moments and jokes about variable reward conditioning. Nicole and Oliver share internal avatar demo breakthroughs and server traffic spikes.4:51–7:01 · The host as informed peer 3/10 The Long-Term Vision for AI in Creative Work Host frames a broad question comparing traditional Photoshop manual processes to single-command AI generation and long-term university education. Nicole outlines the spectrum between professional workflow automation and consumer deck/task agents.7:01–9:58 · The host as informed peer 4/10 Defining Art and Intent in the AI Era Host brings a philosophical framing defining art as out-of-distribution samples. Oliver explicitly rejects this definition as too restrictive, re-anchoring art around human intent and state-of-the-art tooling.9:58–12:04 · The host as informed peer 4/10 User Interfaces: Control Knobs vs. Natural Language Host explores UI evolution, drawing on Oliver's background at Adobe to ask about control knobs versus natural language interface. Nicole suggests AI proactive next-step recommendations, prompting the host to offer a counterpoint.12:04–15:37 · The host as informed peer 5/10 Specialized UI Workflows and Model Ensembles Host presents a direct counterpoint that power users tolerate high UI complexity (citing Cursor), challenging simple chatbot interfaces. Oliver agrees by highlighting node-based ComfyUI workflows and model ensembles.15:38–17:51 · The host as informed peer 3/10 AI in Education and Visual Learning Host asks if kindergarteners will learn drawing through tablet-based AI completion. Oliver educates on how model abstraction makes childlike crayon drawings surprisingly difficult to generate.17:51–19:58 · The host as informed peer 5/10 Multimodal Visual Reasoning Host distinguishes between visual approximation and true visual reasoning in diagrams, asking if future models will spend hours reasoning on drafts. Guests endorse this vision, dubbing it visual deep research.19:58–22:26 · The host as informed peer 6/10 Spatial Representations: 2D Projections vs. 3D World Models Host probes whether 3D world models are necessary over 2D projections, bringing domain experience as a cartoonist on how light and shadow trick human 3D perception. Oliver explains latent world representations and projection advantages.22:26–27:31 · The host as informed peer 5/10 Overcoming the Uncanny Valley and Model Evals Host highlights character consistency challenges and uncanny valley friction for familiar faces. Guests outline internal eval techniques on familiar team faces and model deployment trade-offs.27:31–29:59 · The host as informed peer 6/10 Intent Understanding vs. Granular Control Host invokes Rich Sutton's 'Bitter Lesson' and ControlNet sidecar models to question whether structured pose data remains necessary. Oliver and Nicole emphasize that intent understanding supersedes explicit pose extraction.29:59–32:18 · The host as informed peer 6/10 Beyond Pixels: Code and Parametric Generation Host asks if new image representations like SVGs, bezier curves, or brush parameters will replace pixels. Oliver counters by asserting everything is a subset of pixels, though acknowledging code-guided parametric rendering.32:18–35:45 · The host as informed peer 5/10 Product Strategy: Gemini App vs. Developer API Ecosystem Co-host Justine asks about balancing first-party Gemini app features against API ecosystem verticals, citing specific adoption trends in Japan like the Easy Banana Chrome extension for manga.35:45–37:49 · The host as informed peer 5/10 Speed, Latency, and Visual Explainers as Force Multipliers Host inquires about force multipliers that unlock downstream tasks, referencing Neal Stephenson's 'Diamond Age'. Nicole highlights latency thresholds and factual visual explainers as key multipliers.37:49–41:10 · The host as informed peer 5/10 Image Sequences and Interactive Video Frontiers Host conceptualizes image sequences as nodes within a continuous video directed graph. Oliver agrees, explaining frame sequence prediction before detailing personal use cases with family photos.41:10–45:01 · The host as informed peer 5/10 Surprising Community Discoveries and Zero-Shot Reasoning Host asks for surprising emergent model behaviors. Oliver details zero-shot paper figure reconstruction and geometry problem solving, while host follows up on long-context state retention across turns.45:01–50:01 · The host as informed peer 5/10 Addressing Artist Skepticism and the Value of Human Craft Host questions why visual artists express skepticism toward AI tools. Oliver explains early text-to-image lacked granular artistic control, while Nicole highlights taste, craft, and co-designing with artists like Ross Lovegrove.50:01–53:45 · The host as informed peer 5/10 Underrated Features: Interleaved Generation Host asks for underrated features and summarizes model improvement as moving from cherry-picking to lemon-picking. Oliver and Nicole highlight interleaved generation, context window expansion, and self-critique loops.2:07–4:51 · Guest teaching 3/10 Viral Realizations and Personal Avatars Host asks conversational opening questions about viral moments and jokes about variable reward conditioning. Nicole and Oliver share internal avatar demo breakthroughs and server traffic spikes.4:51–7:01 · Guest teaching 4/10 The Long-Term Vision for AI in Creative Work Host frames a broad question comparing traditional Photoshop manual processes to single-command AI generation and long-term university education. Nicole outlines the spectrum between professional workflow automation and consumer deck/task agents.7:01–9:58 · Guest teaching 5/10 Defining Art and Intent in the AI Era Host brings a philosophical framing defining art as out-of-distribution samples. Oliver explicitly rejects this definition as too restrictive, re-anchoring art around human intent and state-of-the-art tooling.9:58–12:04 · Guest teaching 4/10 User Interfaces: Control Knobs vs. Natural Language Host explores UI evolution, drawing on Oliver's background at Adobe to ask about control knobs versus natural language interface. Nicole suggests AI proactive next-step recommendations, prompting the host to offer a counterpoint.12:04–15:37 · Guest teaching 4/10 Specialized UI Workflows and Model Ensembles Host presents a direct counterpoint that power users tolerate high UI complexity (citing Cursor), challenging simple chatbot interfaces. Oliver agrees by highlighting node-based ComfyUI workflows and model ensembles.15:38–17:51 · Guest teaching 4/10 AI in Education and Visual Learning Host asks if kindergarteners will learn drawing through tablet-based AI completion. Oliver educates on how model abstraction makes childlike crayon drawings surprisingly difficult to generate.17:51–19:58 · Guest teaching 3/10 Multimodal Visual Reasoning Host distinguishes between visual approximation and true visual reasoning in diagrams, asking if future models will spend hours reasoning on drafts. Guests endorse this vision, dubbing it visual deep research.19:58–22:26 · Guest teaching 5/10 Spatial Representations: 2D Projections vs. 3D World Models Host probes whether 3D world models are necessary over 2D projections, bringing domain experience as a cartoonist on how light and shadow trick human 3D perception. Oliver explains latent world representations and projection advantages.22:26–27:31 · Guest teaching 4/10 Overcoming the Uncanny Valley and Model Evals Host highlights character consistency challenges and uncanny valley friction for familiar faces. Guests outline internal eval techniques on familiar team faces and model deployment trade-offs.27:31–29:59 · Guest teaching 4/10 Intent Understanding vs. Granular Control Host invokes Rich Sutton's 'Bitter Lesson' and ControlNet sidecar models to question whether structured pose data remains necessary. Oliver and Nicole emphasize that intent understanding supersedes explicit pose extraction.29:59–32:18 · Guest teaching 4/10 Beyond Pixels: Code and Parametric Generation Host asks if new image representations like SVGs, bezier curves, or brush parameters will replace pixels. Oliver counters by asserting everything is a subset of pixels, though acknowledging code-guided parametric rendering.32:18–35:45 · Guest teaching 3/10 Product Strategy: Gemini App vs. Developer API Ecosystem Co-host Justine asks about balancing first-party Gemini app features against API ecosystem verticals, citing specific adoption trends in Japan like the Easy Banana Chrome extension for manga.35:45–37:49 · Guest teaching 3/10 Speed, Latency, and Visual Explainers as Force Multipliers Host inquires about force multipliers that unlock downstream tasks, referencing Neal Stephenson's 'Diamond Age'. Nicole highlights latency thresholds and factual visual explainers as key multipliers.37:49–41:10 · Guest teaching 3/10 Image Sequences and Interactive Video Frontiers Host conceptualizes image sequences as nodes within a continuous video directed graph. Oliver agrees, explaining frame sequence prediction before detailing personal use cases with family photos.41:10–45:01 · Guest teaching 4/10 Surprising Community Discoveries and Zero-Shot Reasoning Host asks for surprising emergent model behaviors. Oliver details zero-shot paper figure reconstruction and geometry problem solving, while host follows up on long-context state retention across turns.45:01–50:01 · Guest teaching 5/10 Addressing Artist Skepticism and the Value of Human Craft Host questions why visual artists express skepticism toward AI tools. Oliver explains early text-to-image lacked granular artistic control, while Nicole highlights taste, craft, and co-designing with artists like Ross Lovegrove.50:01–53:45 · Guest teaching 4/10 Underrated Features: Interleaved Generation Host asks for underrated features and summarizes model improvement as moving from cherry-picking to lemon-picking. Oliver and Nicole highlight interleaved generation, context window expansion, and self-critique loops.2:07–4:51 · Guest disagreement 1/10 Viral Realizations and Personal Avatars Host asks conversational opening questions about viral moments and jokes about variable reward conditioning. Nicole and Oliver share internal avatar demo breakthroughs and server traffic spikes.4:51–7:01 · Guest disagreement 1/10 The Long-Term Vision for AI in Creative Work Host frames a broad question comparing traditional Photoshop manual processes to single-command AI generation and long-term university education. Nicole outlines the spectrum between professional workflow automation and consumer deck/task agents.7:01–9:58 · Guest disagreement 3/10 Defining Art and Intent in the AI Era Host brings a philosophical framing defining art as out-of-distribution samples. Oliver explicitly rejects this definition as too restrictive, re-anchoring art around human intent and state-of-the-art tooling.9:58–12:04 · Guest disagreement 2/10 User Interfaces: Control Knobs vs. Natural Language Host explores UI evolution, drawing on Oliver's background at Adobe to ask about control knobs versus natural language interface. Nicole suggests AI proactive next-step recommendations, prompting the host to offer a counterpoint.12:04–15:37 · Guest disagreement 2/10 Specialized UI Workflows and Model Ensembles Host presents a direct counterpoint that power users tolerate high UI complexity (citing Cursor), challenging simple chatbot interfaces. Oliver agrees by highlighting node-based ComfyUI workflows and model ensembles.15:38–17:51 · Guest disagreement 1/10 AI in Education and Visual Learning Host asks if kindergarteners will learn drawing through tablet-based AI completion. Oliver educates on how model abstraction makes childlike crayon drawings surprisingly difficult to generate.17:51–19:58 · Guest disagreement 1/10 Multimodal Visual Reasoning Host distinguishes between visual approximation and true visual reasoning in diagrams, asking if future models will spend hours reasoning on drafts. Guests endorse this vision, dubbing it visual deep research.19:58–22:26 · Guest disagreement 2/10 Spatial Representations: 2D Projections vs. 3D World Models Host probes whether 3D world models are necessary over 2D projections, bringing domain experience as a cartoonist on how light and shadow trick human 3D perception. Oliver explains latent world representations and projection advantages.22:26–27:31 · Guest disagreement 1/10 Overcoming the Uncanny Valley and Model Evals Host highlights character consistency challenges and uncanny valley friction for familiar faces. Guests outline internal eval techniques on familiar team faces and model deployment trade-offs.27:31–29:59 · Guest disagreement 2/10 Intent Understanding vs. Granular Control Host invokes Rich Sutton's 'Bitter Lesson' and ControlNet sidecar models to question whether structured pose data remains necessary. Oliver and Nicole emphasize that intent understanding supersedes explicit pose extraction.29:59–32:18 · Guest disagreement 2/10 Beyond Pixels: Code and Parametric Generation Host asks if new image representations like SVGs, bezier curves, or brush parameters will replace pixels. Oliver counters by asserting everything is a subset of pixels, though acknowledging code-guided parametric rendering.32:18–35:45 · Guest disagreement 1/10 Product Strategy: Gemini App vs. Developer API Ecosystem Co-host Justine asks about balancing first-party Gemini app features against API ecosystem verticals, citing specific adoption trends in Japan like the Easy Banana Chrome extension for manga.35:45–37:49 · Guest disagreement 1/10 Speed, Latency, and Visual Explainers as Force Multipliers Host inquires about force multipliers that unlock downstream tasks, referencing Neal Stephenson's 'Diamond Age'. Nicole highlights latency thresholds and factual visual explainers as key multipliers.37:49–41:10 · Guest disagreement 1/10 Image Sequences and Interactive Video Frontiers Host conceptualizes image sequences as nodes within a continuous video directed graph. Oliver agrees, explaining frame sequence prediction before detailing personal use cases with family photos.41:10–45:01 · Guest disagreement 1/10 Surprising Community Discoveries and Zero-Shot Reasoning Host asks for surprising emergent model behaviors. Oliver details zero-shot paper figure reconstruction and geometry problem solving, while host follows up on long-context state retention across turns.45:01–50:01 · Guest disagreement 2/10 Addressing Artist Skepticism and the Value of Human Craft Host questions why visual artists express skepticism toward AI tools. Oliver explains early text-to-image lacked granular artistic control, while Nicole highlights taste, craft, and co-designing with artists like Ross Lovegrove.50:01–53:45 · Guest disagreement 1/10 Underrated Features: Interleaved Generation Host asks for underrated features and summarizes model improvement as moving from cherry-picking to lemon-picking. Oliver and Nicole highlight interleaved generation, context window expansion, and self-critique loops.2:07–4:51 · The host pushing back 1/10 Viral Realizations and Personal Avatars Host asks conversational opening questions about viral moments and jokes about variable reward conditioning. Nicole and Oliver share internal avatar demo breakthroughs and server traffic spikes.4:51–7:01 · The host pushing back 0/10 The Long-Term Vision for AI in Creative Work Host frames a broad question comparing traditional Photoshop manual processes to single-command AI generation and long-term university education. Nicole outlines the spectrum between professional workflow automation and consumer deck/task agents.7:01–9:58 · The host pushing back 2/10 Defining Art and Intent in the AI Era Host brings a philosophical framing defining art as out-of-distribution samples. Oliver explicitly rejects this definition as too restrictive, re-anchoring art around human intent and state-of-the-art tooling.9:58–12:04 · The host pushing back 3/10 User Interfaces: Control Knobs vs. Natural Language Host explores UI evolution, drawing on Oliver's background at Adobe to ask about control knobs versus natural language interface. Nicole suggests AI proactive next-step recommendations, prompting the host to offer a counterpoint.12:04–15:37 · The host pushing back 4/10 Specialized UI Workflows and Model Ensembles Host presents a direct counterpoint that power users tolerate high UI complexity (citing Cursor), challenging simple chatbot interfaces. Oliver agrees by highlighting node-based ComfyUI workflows and model ensembles.15:38–17:51 · The host pushing back 1/10 AI in Education and Visual Learning Host asks if kindergarteners will learn drawing through tablet-based AI completion. Oliver educates on how model abstraction makes childlike crayon drawings surprisingly difficult to generate.17:51–19:58 · The host pushing back 1/10 Multimodal Visual Reasoning Host distinguishes between visual approximation and true visual reasoning in diagrams, asking if future models will spend hours reasoning on drafts. Guests endorse this vision, dubbing it visual deep research.19:58–22:26 · The host pushing back 3/10 Spatial Representations: 2D Projections vs. 3D World Models Host probes whether 3D world models are necessary over 2D projections, bringing domain experience as a cartoonist on how light and shadow trick human 3D perception. Oliver explains latent world representations and projection advantages.22:26–27:31 · The host pushing back 2/10 Overcoming the Uncanny Valley and Model Evals Host highlights character consistency challenges and uncanny valley friction for familiar faces. Guests outline internal eval techniques on familiar team faces and model deployment trade-offs.27:31–29:59 · The host pushing back 3/10 Intent Understanding vs. Granular Control Host invokes Rich Sutton's 'Bitter Lesson' and ControlNet sidecar models to question whether structured pose data remains necessary. Oliver and Nicole emphasize that intent understanding supersedes explicit pose extraction.29:59–32:18 · The host pushing back 2/10 Beyond Pixels: Code and Parametric Generation Host asks if new image representations like SVGs, bezier curves, or brush parameters will replace pixels. Oliver counters by asserting everything is a subset of pixels, though acknowledging code-guided parametric rendering.32:18–35:45 · The host pushing back 1/10 Product Strategy: Gemini App vs. Developer API Ecosystem Co-host Justine asks about balancing first-party Gemini app features against API ecosystem verticals, citing specific adoption trends in Japan like the Easy Banana Chrome extension for manga.35:45–37:49 · The host pushing back 1/10 Speed, Latency, and Visual Explainers as Force Multipliers Host inquires about force multipliers that unlock downstream tasks, referencing Neal Stephenson's 'Diamond Age'. Nicole highlights latency thresholds and factual visual explainers as key multipliers.37:49–41:10 · The host pushing back 1/10 Image Sequences and Interactive Video Frontiers Host conceptualizes image sequences as nodes within a continuous video directed graph. Oliver agrees, explaining frame sequence prediction before detailing personal use cases with family photos.41:10–45:01 · The host pushing back 1/10 Surprising Community Discoveries and Zero-Shot Reasoning Host asks for surprising emergent model behaviors. Oliver details zero-shot paper figure reconstruction and geometry problem solving, while host follows up on long-context state retention across turns.45:01–50:01 · The host pushing back 2/10 Addressing Artist Skepticism and the Value of Human Craft Host questions why visual artists express skepticism toward AI tools. Oliver explains early text-to-image lacked granular artistic control, while Nicole highlights taste, craft, and co-designing with artists like Ross Lovegrove.50:01–53:45 · The host pushing back 1/10 Underrated Features: Interleaved Generation Host asks for underrated features and summarizes model improvement as moving from cherry-picking to lemon-picking. Oliver and Nicole highlight interleaved generation, context window expansion, and self-critique loops.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 0% · guest 100%0:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%36:00 · the host 0% · guest 100%36:00 · the host 0% · guest 100%39:00 · the host 0% · guest 100%39:00 · the host 0% · guest 100%42:00 · the host 0% · guest 100%42:00 · the host 0% · guest 100%45:00 · the host 0% · guest 100%45:00 · the host 0% · guest 100%48:00 · the host 0% · guest 100%48:00 · the host 0% · guest 100%51:00 · the host 0% · guest 100%51:00 · the host 0% · guest 100%54:00 · the host 0% · guest 0%54:00 · the host 0% · guest 0%
Sharpest disagreement ▶ 7:19 Oliver rejects out-of-distribution definition of art

Oliver directly pushes back on the host's quoted definition of art as an out-of-distribution sample, calling it too restrictive and explaining why great art is often in-distribution.

Hardest push from the host ▶ 12:09 Host pushes back on simplified natural language UI vision

The host explicitly announces taking a counterpoint, challenging the guest's simplified conversational UI vision by bringing up power users who demand high complexity and citing Cursor as evidence.

Biggest teaching moment ▶ 30:42 Oliver reframes representations around pixel completeness

When the host asks if net-new representations like SVGs or bezier curves will replace pixels, Oliver firmly educates that everything—including text—is ultimately a subset of pixels.

The host holds their own ▶ 27:31 Host citing Bitter Lesson and ControlNet sidecar architecture

The host demonstrates strong technical depth by referencing Rich Sutton's 'Bitter Lesson', sidecar model architectures like ControlNet, and open pose parameter passing.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
Viral Realizations and Personal Avatars 2311 Host asks conversational opening questions about viral moments and jokes about variable reward conditioning. Nicole and Oliver share internal avatar demo breakthroughs and server traffic spikes.
The Long-Term Vision for AI in Creative Work 3410 Host frames a broad question comparing traditional Photoshop manual processes to single-command AI generation and long-term university education. Nicole outlines the spectrum between professional workflow automation and consumer deck/task agents.
Defining Art and Intent in the AI Era 4532 Host brings a philosophical framing defining art as out-of-distribution samples. Oliver explicitly rejects this definition as too restrictive, re-anchoring art around human intent and state-of-the-art tooling.
User Interfaces: Control Knobs vs. Natural Language 4423 Host explores UI evolution, drawing on Oliver's background at Adobe to ask about control knobs versus natural language interface. Nicole suggests AI proactive next-step recommendations, prompting the host to offer a counterpoint.
Specialized UI Workflows and Model Ensembles 5424 Host presents a direct counterpoint that power users tolerate high UI complexity (citing Cursor), challenging simple chatbot interfaces. Oliver agrees by highlighting node-based ComfyUI workflows and model ensembles.
AI in Education and Visual Learning 3411 Host asks if kindergarteners will learn drawing through tablet-based AI completion. Oliver educates on how model abstraction makes childlike crayon drawings surprisingly difficult to generate.
Multimodal Visual Reasoning 5311 Host distinguishes between visual approximation and true visual reasoning in diagrams, asking if future models will spend hours reasoning on drafts. Guests endorse this vision, dubbing it visual deep research.
Spatial Representations: 2D Projections vs. 3D World Models 6523 Host probes whether 3D world models are necessary over 2D projections, bringing domain experience as a cartoonist on how light and shadow trick human 3D perception. Oliver explains latent world representations and projection advantages.
Overcoming the Uncanny Valley and Model Evals 5412 Host highlights character consistency challenges and uncanny valley friction for familiar faces. Guests outline internal eval techniques on familiar team faces and model deployment trade-offs.
Intent Understanding vs. Granular Control 6423 Host invokes Rich Sutton's 'Bitter Lesson' and ControlNet sidecar models to question whether structured pose data remains necessary. Oliver and Nicole emphasize that intent understanding supersedes explicit pose extraction.
Beyond Pixels: Code and Parametric Generation 6422 Host asks if new image representations like SVGs, bezier curves, or brush parameters will replace pixels. Oliver counters by asserting everything is a subset of pixels, though acknowledging code-guided parametric rendering.
Product Strategy: Gemini App vs. Developer API Ecosystem 5311 Co-host Justine asks about balancing first-party Gemini app features against API ecosystem verticals, citing specific adoption trends in Japan like the Easy Banana Chrome extension for manga.
Speed, Latency, and Visual Explainers as Force Multipliers 5311 Host inquires about force multipliers that unlock downstream tasks, referencing Neal Stephenson's 'Diamond Age'. Nicole highlights latency thresholds and factual visual explainers as key multipliers.
Image Sequences and Interactive Video Frontiers 5311 Host conceptualizes image sequences as nodes within a continuous video directed graph. Oliver agrees, explaining frame sequence prediction before detailing personal use cases with family photos.
Surprising Community Discoveries and Zero-Shot Reasoning 5411 Host asks for surprising emergent model behaviors. Oliver details zero-shot paper figure reconstruction and geometry problem solving, while host follows up on long-context state retention across turns.
Addressing Artist Skepticism and the Value of Human Craft 5522 Host questions why visual artists express skepticism toward AI tools. Oliver explains early text-to-image lacked granular artistic control, while Nicole highlights taste, craft, and co-designing with artists like Ross Lovegrove.
Underrated Features: Interleaved Generation 5411 Host asks for underrated features and summarizes model improvement as moving from cherry-picking to lemon-picking. Oliver and Nicole highlight interleaved generation, context window expansion, and self-critique loops.

Statements from this episode (32)

Assertion Not checkable as stated
Oliver Wang: Nano Banana exceeded initial LMSYS Chatbot Arena query capacity
“What we saw was that we budgeted like, you know, a comparable amount of queries per second as we had for our previous models that were on Ella Marina. And we had to keep upping that number as people were going to Ella Marina to use the model.”
Oliver Wang Oct 28, 2025 ▶ 2:28
Insight
Oliver Wang: Great art is often in-distribution relative to prior art
“I think that out of distribution sample, that is a little bit too restrictive. I think a lot of great art is actually in distribution for art that occurred before it.”
Oliver Wang Oct 28, 2025 ▶ 7:19
Insight
Oliver Wang: Human intent is the defining element of art
“Like to me, I think that the most important thing for art is intent. And so the, what is generated from these models is, is a tool to allow people to create art.”
Oliver Wang Oct 28, 2025 ▶ 7:34
Prediction Not checkable as stated
Oliver Wang: Professional creatives will always adopt state-of-the-art AI tools
“The high end, and the professionals, and the creatives, like, they'll always use state-of-the-art tools, and this is like another tool in the tool belt for people to make cool things.”
Oliver Wang Oct 28, 2025 ▶ 8:05
Assertion Supported
Wang: Nano Banana follows instructions less effectively in long conversations
“Once you get into real long conversations, like it starts to follow your instructions a little bit worse”
Oliver Wang Oct 28, 2025 ▶ 9:42
Assertion Not checkable as stated
Oliver Wang: DeepMind has not solved combining simple and professional AI UIs
“We haven't exactly figured out how to enable both of those yet.”
Oliver Wang Oct 28, 2025 ▶ 10:58
Insight
Brichtova: Massive opportunity exists for mid-tier AI tools between chatbots and pro software
“And for them, I do think that there's a space of, like, that you need more control than the chatbot gives you, but you don't need as much control as what the professional tools give you, and like, what's that kind of in-between state? There's a ton of opportun…”
Nicole Brichtova Oct 28, 2025 ▶ 14:03
Prediction Not checkable as stated
Wang: A single AI model will never satisfy all use cases
“I definitely don't think that, that the broad amount of use cases will be fully satisfied by one model at any point. So I think that there will always be a diversity of models. Some, like I'll give you an example, but some, you know, we could optimize for inst…”
Oliver Wang Oct 28, 2025 ▶ 14:55
Assertion Not checkable as stated
Oliver Wang: AI models struggle to generate childlike crayon drawings
“We've been trying to get the model to create like, childlike crayon drawings, which is actually quite challenging. Ironically, you know, sometimes the things that are hard to make are, because the level of abstraction is very large. So it's actually quite diff…”
Oliver Wang Oct 28, 2025 ▶ 16:49
Prediction Not checkable as stated
Oliver Wang: Multimodal AI will transform education through visual learning
“I'm in general, I'm very, Optimistic about AI for education, and part of the reason is I think that most of us are visual learners, right? So the AI right now as a tutor, basically all it can do is talk to you or give you text to read, and that's definitely no…”
Oliver Wang Oct 28, 2025 ▶ 17:14
Prediction Not checkable as stated
Wang: Visual modality will remain critical for AI agents collaborating with humans
“A hundred percent. I definitely think so. The future for these AI models that I'm most excited by is where they are tools for people to accomplish more things. Like, I think if you imagine a future where you have these agentic models that just talk to each oth…”
Oliver Wang Oct 28, 2025 ▶ 18:23
Prediction Open · timeframe Oct 2030
Wang: AI image models will adopt multi-hour test-time compute when needed
“Yeah, absolutely. If it's necessary, yeah.”
Oliver Wang Oct 28, 2025 ▶ 19:07
Insight
Wang: Generative AI can solve visual problems using 2D projections
“I think we can solve almost all the problems, if not all the problems, working on the projection of the three-D world directly, and letting the models learn the latent world representations.”
Oliver Wang Oct 28, 2025 ▶ 20:47
Assertion Not checkable as stated
Wang: AI-generated videos yield highly accurate 3D reconstructions
“Video models have very good three-D understanding. You can run reconstruction algorithms over the videos you generate, and they're very accurate.”
Oliver Wang Oct 28, 2025 ▶ 20:56
Insight
Brichtova: AI character consistency evals require testing on familiar faces
“So when we're developing this model, we actually started out doing character consistency evals and faces we didn't know, and it doesn't tell you anything. And then we started testing it on ourselves and quickly realized like, okay, this is what you need to do …”
Nicole Brichtova Oct 28, 2025 ▶ 23:19
Insight
Wang: Image model adoption accelerates once character consistency crosses quality threshold
“Once the quality gets above a certain level for character consistency, it can kind of just take off because it becomes useful for so much more.”
Oliver Wang Oct 28, 2025 ▶ 24:23
Insight
Wang: Generative AI model output reflects research lab preferences over objective standards
“There is no right answer. So actually there's quite a lot of, I don't know if it's taste, but it's like preference that goes into the models. And I think you can kind of see the difference in preferences of the different research labs in the models that they r…”
Oliver Wang Oct 28, 2025 ▶ 25:42
Disclosure
Brichtova: Nano Banana text rendering underperformed at initial release
“For this first release, the model's not as good as text rendering at, as we would like it to be, and that's something that we want to fix in the future”
Nicole Brichtova Oct 28, 2025 ▶ 27:04
Insight
Wang: Artists want AI models to infer intent, which models now do better
“Really what an artist wants when they want to do something is they want the intent to be understood. And I think that, that these AI models are getting better at understanding the intent of users. So often when you ask text queries now, the model gets what you…”
Oliver Wang Oct 28, 2025 ▶ 28:17
Prediction Not checkable as stated
Wang: AI models will solve complex multi-person scenes without dedicated pose interfaces
“Yeah, I think in that case, I wouldn't spend a ton of time building a custom interface for making this picture of 46 people. It seems like the kind of thing that we can solve.”
Oliver Wang Oct 28, 2025 ▶ 29:46
Insight
Oliver Wang: Editability is the main reason to leave pixel representations
“The primary reason I think you would want to leave the pixel domain is for editability.”
Oliver Wang Oct 28, 2025 ▶ 30:58
Insight
Brichtova: Fun consumer AI features serve as a gateway to utility
“Fun is kind of a gateway to utility where, you know, people come to make a figurine image of themselves, but then they stay because it helps them with their math homework, or it helps them write something, right?”
Nicole Brichtova Oct 28, 2025 ▶ 33:25
Prediction Not checkable as stated
Brichtova: DeepMind will not build specialized architecture software
“We're probably not going to go build a software for an architecture firm. My dad is an architect and he would probably love that. But I don't think that's something that we will do, but somebody should go and do that.”
Nicole Brichtova Oct 28, 2025 ▶ 34:24
Insight
Brichtova: Speed is an AI force multiplier only after meeting quality thresholds
“There has to be some quality bar because if it's just fast and the quality isn't there, then it also doesn't matter, right? Like you have to hit a quality bar and then speed becomes a force multiplier.”
Nicole Brichtova Oct 28, 2025 ▶ 36:40
Prediction Not checkable as stated
Oliver Wang: Generative AI is heading toward fully interactive real-time environments
“Making something that's like fully interactive and real time and is the direction this field is headed.”
Oliver Wang Oct 28, 2025 ▶ 39:08
Insight
Wang: Core value of edit models is extreme personalization
“The real beauty of the edit models is that you can make it about the one thing that matters most to you.”
Oliver Wang Oct 28, 2025 ▶ 40:05
Assertion Not checkable as stated
Oliver Wang: Nano Banana renders functional webpages from HTML code images
“I've seen examples where people give it, like an image of HTML code and have the model render the webpage.”
Oliver Wang Oct 28, 2025 ▶ 42:37
Opinion
Brichtova: AI image generation models lack artistic taste
“But it is, there's a lot of craft, and there's a lot of taste, right, that you accumulate sometimes over decades, right, and I don't think these models really have taste, right”
Nicole Brichtova Oct 28, 2025 ▶ 47:41
Disclosure
Brichtova: DeepMind prototyped physical chair designed via fine-tuned model
“We just worked with Russ Lovegrove on fine tuning a model on his sketches so that he can then create something new out of that, and then we design an actual physical chair that we, like, have a prototype of.”
Nicole Brichtova Oct 28, 2025 ▶ 48:17
Insight
Wang: Optimizing AI models for average user preferences dilutes creative output
“So when we're, you know, when we're optimizing these models, like, one thing we could do is we could optimize for Like the average preference of everybody. But I don't think you end up with interesting things by doing that. You end up with something that every…”
Oliver Wang Oct 28, 2025 ▶ 49:32
Assertion Not checkable as stated
Wang: Users have not yet discovered Nano Banana's interleaved generation feature
“I think we've always been amazed that nobody ever posts anything about, so interleaved generation is what we call the model's ability to generate more than one image for a specific prompt. So you can ask for like, I want a story, like a bedtime story or someth…”
Oliver Wang Oct 28, 2025 ▶ 50:16
Insight
Oliver Wang: AI image models must raise worst-case output quality
“Every model can cherry pick images that look perfect. So like now I think the real question is like how expressible is this model and what's the worst image you would get given what you're trying to do? So I think by raising the quality of the worst image We r…”
Oliver Wang Oct 28, 2025 ▶ 51:10
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.