Jun 2, 2025 · 24m · latent-space

[AIEWF Preview] Gemini in 2025 and Realtime Voice AI

Shrestha Basu Mallick · 8m spoken Logan Kilpatrick · 5m spoken Shawn Wang · 2m spoken Sam Charrington · 1m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Recorded on-site at Google I/O, hosts from TWIML AI and Latent Space interview Google DeepMind product leaders and Daily's CEO to examine recent Gemini platform upgrades, infrastructure trade-offs, and the architectural evolution of real-time voice AI.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 11.8% of the talking time here. How this is scored →

The hosts as informed peer 4.6 Guest teaching 4.4 Guest disagreement 1.3 The hosts pushing back 1.9
05100:0010:0020:000:04–3:52 · The hosts as informed peer 4/10 Introductions and Roles at Google AI Studio The conversation opens with lighthearted introductions and IO recap highlights. Logan and Shrestha share insights on thinking budgets, thought summaries, and URL context tooling while the hosts facilitate smoothly.3:52–6:59 · The hosts as informed peer 5/10 Implicit Context Caching and Infrastructure Trade-offs Swyx probes into the infrastructure complexities of context caching and highlights Gemini diffusion. Logan elaborates on the latency-versus-cost trade-offs and explains the potential for real-time generative UI.7:00–11:03 · The hosts as informed peer 4/10 Gemini Live API Challenges, Constraints, and Multi-State Workflows Sam and Swyx ask about developer friction with the Live API. Shrestha and Logan educate on session length limits, multi-state system instruction changes, and high provider lock-in due to bespoke infrastructure.11:04–14:29 · The hosts as informed peer 6/10 Unified Gemini Foundation Models Versus Modular Architectures Logan explains DeepMind's single-model philosophy versus modular architectures. Swyx pushes back when image generation is conflated, pointing out the architectural difference between autoregressive and diffusion models.14:29–17:02 · The hosts as informed peer 4/10 Quinn Joins to Discuss Daily, Pipecat, and Voice Architecture Evolution Quinn joins to discuss Daily and Pipecat's orchestration role. Shrestha details the initial architectural trade-off of using NotebookLM's TTS before graduating to native audio-to-audio streaming.17:03–20:17 · The hosts as informed peer 4/10 Voice Activity Detection, Frameworks, and Low-Latency Networking Sam asks about the voice-specific infrastructure needed around models. Shrestha and Quinn detail server-side VAD tuning, sub-500ms latency requirements, and the division of labor between frameworks and native model APIs.20:18–22:35 · The hosts as informed peer 5/10 Proactive Audio, Speaker Identification, and Asynchronous Tool Calling Shrestha and Quinn detail cutting-edge capabilities like proactive audio semantic filtering, unreleased multi-speaker identification, and non-blocking asynchronous function execution.0:04–3:52 · Guest teaching 3/10 Introductions and Roles at Google AI Studio The conversation opens with lighthearted introductions and IO recap highlights. Logan and Shrestha share insights on thinking budgets, thought summaries, and URL context tooling while the hosts facilitate smoothly.3:52–6:59 · Guest teaching 3/10 Implicit Context Caching and Infrastructure Trade-offs Swyx probes into the infrastructure complexities of context caching and highlights Gemini diffusion. Logan elaborates on the latency-versus-cost trade-offs and explains the potential for real-time generative UI.7:00–11:03 · Guest teaching 5/10 Gemini Live API Challenges, Constraints, and Multi-State Workflows Sam and Swyx ask about developer friction with the Live API. Shrestha and Logan educate on session length limits, multi-state system instruction changes, and high provider lock-in due to bespoke infrastructure.11:04–14:29 · Guest teaching 5/10 Unified Gemini Foundation Models Versus Modular Architectures Logan explains DeepMind's single-model philosophy versus modular architectures. Swyx pushes back when image generation is conflated, pointing out the architectural difference between autoregressive and diffusion models.14:29–17:02 · Guest teaching 4/10 Quinn Joins to Discuss Daily, Pipecat, and Voice Architecture Evolution Quinn joins to discuss Daily and Pipecat's orchestration role. Shrestha details the initial architectural trade-off of using NotebookLM's TTS before graduating to native audio-to-audio streaming.17:03–20:17 · Guest teaching 5/10 Voice Activity Detection, Frameworks, and Low-Latency Networking Sam asks about the voice-specific infrastructure needed around models. Shrestha and Quinn detail server-side VAD tuning, sub-500ms latency requirements, and the division of labor between frameworks and native model APIs.20:18–22:35 · Guest teaching 6/10 Proactive Audio, Speaker Identification, and Asynchronous Tool Calling Shrestha and Quinn detail cutting-edge capabilities like proactive audio semantic filtering, unreleased multi-speaker identification, and non-blocking asynchronous function execution.0:04–3:52 · Guest disagreement 1/10 Introductions and Roles at Google AI Studio The conversation opens with lighthearted introductions and IO recap highlights. Logan and Shrestha share insights on thinking budgets, thought summaries, and URL context tooling while the hosts facilitate smoothly.3:52–6:59 · Guest disagreement 1/10 Implicit Context Caching and Infrastructure Trade-offs Swyx probes into the infrastructure complexities of context caching and highlights Gemini diffusion. Logan elaborates on the latency-versus-cost trade-offs and explains the potential for real-time generative UI.7:00–11:03 · Guest disagreement 1/10 Gemini Live API Challenges, Constraints, and Multi-State Workflows Sam and Swyx ask about developer friction with the Live API. Shrestha and Logan educate on session length limits, multi-state system instruction changes, and high provider lock-in due to bespoke infrastructure.11:04–14:29 · Guest disagreement 2/10 Unified Gemini Foundation Models Versus Modular Architectures Logan explains DeepMind's single-model philosophy versus modular architectures. Swyx pushes back when image generation is conflated, pointing out the architectural difference between autoregressive and diffusion models.14:29–17:02 · Guest disagreement 1/10 Quinn Joins to Discuss Daily, Pipecat, and Voice Architecture Evolution Quinn joins to discuss Daily and Pipecat's orchestration role. Shrestha details the initial architectural trade-off of using NotebookLM's TTS before graduating to native audio-to-audio streaming.17:03–20:17 · Guest disagreement 1/10 Voice Activity Detection, Frameworks, and Low-Latency Networking Sam asks about the voice-specific infrastructure needed around models. Shrestha and Quinn detail server-side VAD tuning, sub-500ms latency requirements, and the division of labor between frameworks and native model APIs.20:18–22:35 · Guest disagreement 2/10 Proactive Audio, Speaker Identification, and Asynchronous Tool Calling Shrestha and Quinn detail cutting-edge capabilities like proactive audio semantic filtering, unreleased multi-speaker identification, and non-blocking asynchronous function execution.0:04–3:52 · The hosts pushing back 1/10 Introductions and Roles at Google AI Studio The conversation opens with lighthearted introductions and IO recap highlights. Logan and Shrestha share insights on thinking budgets, thought summaries, and URL context tooling while the hosts facilitate smoothly.3:52–6:59 · The hosts pushing back 2/10 Implicit Context Caching and Infrastructure Trade-offs Swyx probes into the infrastructure complexities of context caching and highlights Gemini diffusion. Logan elaborates on the latency-versus-cost trade-offs and explains the potential for real-time generative UI.7:00–11:03 · The hosts pushing back 2/10 Gemini Live API Challenges, Constraints, and Multi-State Workflows Sam and Swyx ask about developer friction with the Live API. Shrestha and Logan educate on session length limits, multi-state system instruction changes, and high provider lock-in due to bespoke infrastructure.11:04–14:29 · The hosts pushing back 4/10 Unified Gemini Foundation Models Versus Modular Architectures Logan explains DeepMind's single-model philosophy versus modular architectures. Swyx pushes back when image generation is conflated, pointing out the architectural difference between autoregressive and diffusion models.14:29–17:02 · The hosts pushing back 1/10 Quinn Joins to Discuss Daily, Pipecat, and Voice Architecture Evolution Quinn joins to discuss Daily and Pipecat's orchestration role. Shrestha details the initial architectural trade-off of using NotebookLM's TTS before graduating to native audio-to-audio streaming.17:03–20:17 · The hosts pushing back 1/10 Voice Activity Detection, Frameworks, and Low-Latency Networking Sam asks about the voice-specific infrastructure needed around models. Shrestha and Quinn detail server-side VAD tuning, sub-500ms latency requirements, and the division of labor between frameworks and native model APIs.20:18–22:35 · The hosts pushing back 2/10 Proactive Audio, Speaker Identification, and Asynchronous Tool Calling Shrestha and Quinn detail cutting-edge capabilities like proactive audio semantic filtering, unreleased multi-speaker identification, and non-blocking asynchronous function execution.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 23.9% · guest 76.1%0:00 · the hosts 23.9% · guest 76.1%3:00 · the hosts 15.1% · guest 84.9%3:00 · the hosts 15.1% · guest 84.9%6:00 · the hosts 6.7% · guest 93.3%6:00 · the hosts 6.7% · guest 93.3%9:00 · the hosts 1.1% · guest 98.9%9:00 · the hosts 1.1% · guest 98.9%12:00 · the hosts 15% · guest 85%12:00 · the hosts 15% · guest 85%15:00 · the hosts 9.4% · guest 90.6%15:00 · the hosts 9.4% · guest 90.6%18:00 · the hosts 1% · guest 99%18:00 · the hosts 1% · guest 99%21:00 · the hosts 20.3% · guest 79.7%21:00 · the hosts 20.3% · guest 79.7%24:00 · the hosts 45.7% · guest 54.3%24:00 · the hosts 45.7% · guest 54.3%
Sharpest disagreement ▶ 14:10 Technical clarification on Imagen vs Gemini models

When Shrestha groups interleaved image generation under Gemini's capabilities, Swyx interjects to enforce the technical boundary between autoregressive and diffusion architectures.

Hardest push from the hosts ▶ 14:10 Swyx challenges model unification claim

Swyx refuses the implied premise that image generation in Gemini is the exact same underlying architecture, pressing that one is autoregressive and the other is diffusion.

Biggest teaching moment ▶ 12:04 Logan explains DeepMind's unified model philosophy

Logan explains the high-level philosophy from DeepMind leadership on training one core Gemini model and how merging specialized research branches unlocked unexpected multimodal reasoning capabilities.

The host holds their own ▶ 14:10 Swyx demonstrates technical distinction on model types

Swyx demonstrates his technical grasp by immediately distinguishing the diffusion backbone of Imagen from the autoregressive backbone of Gemini.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Introductions and Roles at Google AI Studio 4311 The conversation opens with lighthearted introductions and IO recap highlights. Logan and Shrestha share insights on thinking budgets, thought summaries, and URL context tooling while the hosts facilitate smoothly.
Implicit Context Caching and Infrastructure Trade-offs 5312 Swyx probes into the infrastructure complexities of context caching and highlights Gemini diffusion. Logan elaborates on the latency-versus-cost trade-offs and explains the potential for real-time generative UI.
Gemini Live API Challenges, Constraints, and Multi-State Workflows 4512 Sam and Swyx ask about developer friction with the Live API. Shrestha and Logan educate on session length limits, multi-state system instruction changes, and high provider lock-in due to bespoke infrastructure.
Unified Gemini Foundation Models Versus Modular Architectures 6524 Logan explains DeepMind's single-model philosophy versus modular architectures. Swyx pushes back when image generation is conflated, pointing out the architectural difference between autoregressive and diffusion models.
Quinn Joins to Discuss Daily, Pipecat, and Voice Architecture Evolution 4411 Quinn joins to discuss Daily and Pipecat's orchestration role. Shrestha details the initial architectural trade-off of using NotebookLM's TTS before graduating to native audio-to-audio streaming.
Voice Activity Detection, Frameworks, and Low-Latency Networking 4511 Sam asks about the voice-specific infrastructure needed around models. Shrestha and Quinn detail server-side VAD tuning, sub-500ms latency requirements, and the division of labor between frameworks and native model APIs.
Proactive Audio, Speaker Identification, and Asynchronous Tool Calling 5622 Shrestha and Quinn detail cutting-edge capabilities like proactive audio semantic filtering, unreleased multi-speaker identification, and non-blocking asynchronous function execution.

Statements from this episode (19)

Disclosure
Kilpatrick: Gemini 2.5 Pro will allow disabling thinking in early June
“Thinking budgets coming to 2.5 pro. So, and you can also, you'll be able to disable thinking as well. So if you just want 2.5 pro is like a raw non reasoning model, we'll have that hopefully in early June”
Logan Kilpatrick Jun 2, 2025 ▶ 1:30
Prediction Not checkable as stated
Mallick: Google's URL Context tool will enable developer research agents
“And I think that this will unlock new use cases. Like if people want to build their own version of a research agent, which is something developers ask us for a lot.”
Shrestha Basu Mallick Jun 2, 2025 ▶ 3:43
Assertion Supported
Kilpatrick: Google's implicit caching automatically passes cost savings to developers
“Explicit caching is nice. Like there's definitely use cases where it makes sense, but people want implicit caching. So I'm happy passing the cost saving on to developers. You don't have to do anything. It just works right now and you're saving money.”
Logan Kilpatrick Jun 2, 2025 ▶ 4:09
Prediction Not checkable as stated
Kilpatrick: Generative UI will be the killer use case for diffusion LLMs
“But I do think that's going to be the killer use case will be like this generative UI experience that doesn't exist today because the models just take too long to generate tokens.”
Logan Kilpatrick Jun 2, 2025 ▶ 6:51
Assertion Not checkable as stated
Basu Mallick: Transcription was one of Gemini Live API's biggest early use cases
“Transcription actually, even before we released native audio, now of course you get text and audio interleaved in the output, but transcription used to be one of the biggest use cases we had on the live API.”
Shrestha Basu Mallick Jun 2, 2025 ▶ 7:25
Assertion Supported
Mallick: Google was first to market with live video API input
“So we were actually the first to market with also video input”
Shrestha Basu Mallick Jun 2, 2025 ▶ 7:53
Assertion Partly supported
Basu Mallick: Early Gemini Live sessions were capped at 20m audio, 5m video
“Like when we started, you could do like 15 to 20 minutes of audio, I'm sorry, and about five minutes of video.”
Shrestha Basu Mallick Jun 2, 2025 ▶ 8:05
Assertion Supported
Basu Mallick: Google was first to introduce tool chaining for live APIs
“Again, we were very proud because we introduced tool chaining first, so you could change search and code execution to all kinds of analysis.”
Shrestha Basu Mallick Jun 2, 2025 ▶ 8:33
Insight
Kilpatrick: Live AI APIs create severe vendor lock-in due to bespoke infra
“I think if you look at a lot of the live API infrastructure right now, like you really do need to commit that you're like gonna, you know, there's, it's not easily interoperable between different model providers. Like everyone's infrastructure is all bespoke a…”
Logan Kilpatrick Jun 2, 2025 ▶ 9:16
Prediction Not checkable as stated
Kilpatrick: Model-agnostic infrastructure will emerge for real-time live APIs
“I think hopefully there'll be like some level of like similarity and you'll get some model agnostic infrastructure to help make that, you know, make developers feel a little bit easier about being able to move between models potentially.”
Logan Kilpatrick Jun 2, 2025 ▶ 9:43
Prediction Not checkable as stated
Basu Mallick: Most voice use cases will transition to audio-to-audio models
“I do think perhaps eventually For most use cases, as these audio to audio architecture models get better, a lot of use cases will probably transition to that.”
Shrestha Basu Mallick Jun 2, 2025 ▶ 11:29
Disclosure
Kilpatrick: Google's strategy is building Gemini as one single unified model
“Like we're here to make one model and like that model is Gemini. And like, I think you do need to just trust this point, like to make the capabilities work in some cases, like you do need to have these forks that like go off and make that capability and harde…”
Logan Kilpatrick Jun 2, 2025 ▶ 12:25
Assertion Not checkable as stated
Kilpatrick: Gemini's SOTA video performance resulted from reasoning, not video engineering
“With reasoning is a great example of this where like multimodal with video understanding ended up like having this huge, like it's having this beautiful moment. The model is like soda out of the box because of all the reasoning capabilities that were baked in…”
Logan Kilpatrick Jun 2, 2025 ▶ 13:19
Assertion Supported
Mallick: Gemini Live API allows custom VAD tuning and third-party integration
“Now developers can actually tune the sensitivity on our voice activity detection model as well as, you know, how much of the prefix, like how much of a time duration at the beginning, at the start or stop of saying things. And we also have a mode where now you…”
Shrestha Basu Mallick Jun 2, 2025 ▶ 17:30
Disclosure
Mallick: Google released experimental proactive audio for Gemini native audio
“One of the features that we've pushed out A little more experimental, but would love for people to test it is what we're calling proactive audio, and it's available only in the native audio, in the audio to audio architecture right now. And what this feature d…”
Shrestha Basu Mallick Jun 2, 2025 ▶ 20:19
Assertion Contradicted
Mallick: Gemini recognizes distinct voices as an unsupported emergent behavior
“This is not officially supported yet. The model just does it.”
Shrestha Basu Mallick Jun 2, 2025 ▶ 21:31
Disclosure
Mallick: Google launched async function calling for cascaded voice models
“One thing that we launched on the cascaded architecture that we hope to eventually bring to the native audio as well is asynchronous function calling.”
Shrestha Basu Mallick Jun 2, 2025 ▶ 22:07
Assertion Supported
Mallick: Gemini officially supports 24 languages but responds in Klingon
“We officially support 24 languages, but you can try talking to the model and cling on and it'll respond to you.”
Shrestha Basu Mallick Jun 2, 2025 ▶ 23:28
Prediction Not checkable as stated
Mallick: Gemini will expand global language support well before next I/O
“So I think we'll get there way before next IO, but I just think more and more capabilities into the main model.”
Shrestha Basu Mallick Jun 2, 2025 ▶ 23:37
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.