Jun 2, 2025 · 24m · latent-space
[AIEWF Preview] Gemini in 2025 and Realtime Voice AI
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
Recorded on-site at Google I/O, hosts from TWIML AI and Latent Space interview Google DeepMind product leaders and Daily's CEO to examine recent Gemini platform upgrades, infrastructure trade-offs, and the architectural evolution of real-time voice AI.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 11.8% of the talking time here. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
When Shrestha groups interleaved image generation under Gemini's capabilities, Swyx interjects to enforce the technical boundary between autoregressive and diffusion architectures.
Hardest push from the hosts ▶ 14:10 Swyx challenges model unification claimSwyx refuses the implied premise that image generation in Gemini is the exact same underlying architecture, pressing that one is autoregressive and the other is diffusion.
Biggest teaching moment ▶ 12:04 Logan explains DeepMind's unified model philosophyLogan explains the high-level philosophy from DeepMind leadership on training one core Gemini model and how merging specialized research branches unlocked unexpected multimodal reasoning capabilities.
The host holds their own ▶ 14:10 Swyx demonstrates technical distinction on model typesSwyx demonstrates his technical grasp by immediately distinguishing the diffusion backbone of Imagen from the autoregressive backbone of Gemini.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Introductions and Roles at Google AI Studio | 4 | 3 | 1 | 1 | The conversation opens with lighthearted introductions and IO recap highlights. Logan and Shrestha share insights on thinking budgets, thought summaries, and URL context tooling while the hosts facilitate smoothly. | |
| Implicit Context Caching and Infrastructure Trade-offs | 5 | 3 | 1 | 2 | Swyx probes into the infrastructure complexities of context caching and highlights Gemini diffusion. Logan elaborates on the latency-versus-cost trade-offs and explains the potential for real-time generative UI. | |
| Gemini Live API Challenges, Constraints, and Multi-State Workflows | 4 | 5 | 1 | 2 | Sam and Swyx ask about developer friction with the Live API. Shrestha and Logan educate on session length limits, multi-state system instruction changes, and high provider lock-in due to bespoke infrastructure. | |
| Unified Gemini Foundation Models Versus Modular Architectures | 6 | 5 | 2 | 4 | Logan explains DeepMind's single-model philosophy versus modular architectures. Swyx pushes back when image generation is conflated, pointing out the architectural difference between autoregressive and diffusion models. | |
| Quinn Joins to Discuss Daily, Pipecat, and Voice Architecture Evolution | 4 | 4 | 1 | 1 | Quinn joins to discuss Daily and Pipecat's orchestration role. Shrestha details the initial architectural trade-off of using NotebookLM's TTS before graduating to native audio-to-audio streaming. | |
| Voice Activity Detection, Frameworks, and Low-Latency Networking | 4 | 5 | 1 | 1 | Sam asks about the voice-specific infrastructure needed around models. Shrestha and Quinn detail server-side VAD tuning, sub-500ms latency requirements, and the division of labor between frameworks and native model APIs. | |
| Proactive Audio, Speaker Identification, and Asynchronous Tool Calling | 5 | 6 | 2 | 2 | Shrestha and Quinn detail cutting-edge capabilities like proactive audio semantic filtering, unreleased multi-speaker identification, and non-blocking asynchronous function execution. |