Feb 28, 2025 · 28m · latent-space

Gemini 2.0 Flash and Flash Thinking: the new SOTA models for the agentic era

Logan Kilpatrick · 17m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Google AI Studio Product Lead Logan Kilpatrick joins hosts Alessio and Swix to break down the Gemini 2.0 ecosystem, detailing reasoning scaling in Flash Thinking, real-time multimodal APIs, and developer platform strategies.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 6.1 Guest teaching 4.7 Guest disagreement 1.6 The hosts pushing back 3.0
05100:0010:0020:000:03–5:47 · The hosts as informed peer 6/10 Logan Kilpatrick Returns and Introduces Google Product Role Swix introduces Logan and articulates the emerging meta-pattern of frontier distillation pipelines and MoE teacher models. Logan provides inside context on Google's pricing strategy and the shift from input-tiered pricing to flat-rate tokens.5:47–8:23 · The hosts as informed peer 5/10 Google's Experimental Releases and Coding with Cursor Composer Alessio presses on the operational meaning of Google's 'experimental' model labels and shares real-world developer feedback regarding Cursor Composer support. Logan clarifies Google's release-valve strategy and hot-swapping practices.8:23–12:31 · The hosts as informed peer 6/10 The Reasoning Frontier and Native Scaling in Flash Thinking Swix and Logan discuss the shift from parameter scaling to inference-time compute scaling. Swix highlights the 2T parameter wall while Logan explains the co-scaling dynamics between core pre-trained base models and reinforcement learning.12:31–15:02 · The hosts as informed peer 7/10 Long-Context Benchmarks and Abstraction vs Transparency in AI Swix presents empirical findings from his daily regression benchmarks where Gemini Flash outperforms rival models in long-context summarization. When Logan proposes turning it into an open leaderboard, Swix pushes back with product design rationale around interface abstraction.15:02–20:38 · The hosts as informed peer 7/10 DeepSeek R1 Insights and Convergence of Reasoning Methods Swix cuts through generic high-level discussion on DeepSeek to probe specific technical architectural shifts like discarding Monte Carlo tree search and process reward models. Alessio then questions whether conventional chat interfaces represent legacy technical debt.20:39–24:42 · The hosts as informed peer 7/10 Gemini Nano Integration and Local On-Device Model Adoption Swix explains the technical blocker behind Gemini Nano adoption, noting it remains gated behind Chrome Canary flags. They transition to discussing Project Astra, where Swix raises architectural trade-offs regarding KV caching, deletion constraints, and attention mechanisms versus RAG.24:42–25:47 · The hosts as informed peer 5/10 Search as a Tool, Deep Research, and Online LLMs Logan outlines Google's Search as a Tool initiative and its role in deep research workflows. Swix categorizes the emerging online LLM market landscape before wrapping up with podcast cross-promotions.0:03–5:47 · Guest teaching 4/10 Logan Kilpatrick Returns and Introduces Google Product Role Swix introduces Logan and articulates the emerging meta-pattern of frontier distillation pipelines and MoE teacher models. Logan provides inside context on Google's pricing strategy and the shift from input-tiered pricing to flat-rate tokens.5:47–8:23 · Guest teaching 5/10 Google's Experimental Releases and Coding with Cursor Composer Alessio presses on the operational meaning of Google's 'experimental' model labels and shares real-world developer feedback regarding Cursor Composer support. Logan clarifies Google's release-valve strategy and hot-swapping practices.8:23–12:31 · Guest teaching 5/10 The Reasoning Frontier and Native Scaling in Flash Thinking Swix and Logan discuss the shift from parameter scaling to inference-time compute scaling. Swix highlights the 2T parameter wall while Logan explains the co-scaling dynamics between core pre-trained base models and reinforcement learning.12:31–15:02 · Guest teaching 4/10 Long-Context Benchmarks and Abstraction vs Transparency in AI Swix presents empirical findings from his daily regression benchmarks where Gemini Flash outperforms rival models in long-context summarization. When Logan proposes turning it into an open leaderboard, Swix pushes back with product design rationale around interface abstraction.15:02–20:38 · Guest teaching 5/10 DeepSeek R1 Insights and Convergence of Reasoning Methods Swix cuts through generic high-level discussion on DeepSeek to probe specific technical architectural shifts like discarding Monte Carlo tree search and process reward models. Alessio then questions whether conventional chat interfaces represent legacy technical debt.20:39–24:42 · Guest teaching 6/10 Gemini Nano Integration and Local On-Device Model Adoption Swix explains the technical blocker behind Gemini Nano adoption, noting it remains gated behind Chrome Canary flags. They transition to discussing Project Astra, where Swix raises architectural trade-offs regarding KV caching, deletion constraints, and attention mechanisms versus RAG.24:42–25:47 · Guest teaching 4/10 Search as a Tool, Deep Research, and Online LLMs Logan outlines Google's Search as a Tool initiative and its role in deep research workflows. Swix categorizes the emerging online LLM market landscape before wrapping up with podcast cross-promotions.0:03–5:47 · Guest disagreement 1/10 Logan Kilpatrick Returns and Introduces Google Product Role Swix introduces Logan and articulates the emerging meta-pattern of frontier distillation pipelines and MoE teacher models. Logan provides inside context on Google's pricing strategy and the shift from input-tiered pricing to flat-rate tokens.5:47–8:23 · Guest disagreement 1/10 Google's Experimental Releases and Coding with Cursor Composer Alessio presses on the operational meaning of Google's 'experimental' model labels and shares real-world developer feedback regarding Cursor Composer support. Logan clarifies Google's release-valve strategy and hot-swapping practices.8:23–12:31 · Guest disagreement 2/10 The Reasoning Frontier and Native Scaling in Flash Thinking Swix and Logan discuss the shift from parameter scaling to inference-time compute scaling. Swix highlights the 2T parameter wall while Logan explains the co-scaling dynamics between core pre-trained base models and reinforcement learning.12:31–15:02 · Guest disagreement 2/10 Long-Context Benchmarks and Abstraction vs Transparency in AI Swix presents empirical findings from his daily regression benchmarks where Gemini Flash outperforms rival models in long-context summarization. When Logan proposes turning it into an open leaderboard, Swix pushes back with product design rationale around interface abstraction.15:02–20:38 · Guest disagreement 3/10 DeepSeek R1 Insights and Convergence of Reasoning Methods Swix cuts through generic high-level discussion on DeepSeek to probe specific technical architectural shifts like discarding Monte Carlo tree search and process reward models. Alessio then questions whether conventional chat interfaces represent legacy technical debt.20:39–24:42 · Guest disagreement 2/10 Gemini Nano Integration and Local On-Device Model Adoption Swix explains the technical blocker behind Gemini Nano adoption, noting it remains gated behind Chrome Canary flags. They transition to discussing Project Astra, where Swix raises architectural trade-offs regarding KV caching, deletion constraints, and attention mechanisms versus RAG.24:42–25:47 · Guest disagreement 0/10 Search as a Tool, Deep Research, and Online LLMs Logan outlines Google's Search as a Tool initiative and its role in deep research workflows. Swix categorizes the emerging online LLM market landscape before wrapping up with podcast cross-promotions.0:03–5:47 · The hosts pushing back 2/10 Logan Kilpatrick Returns and Introduces Google Product Role Swix introduces Logan and articulates the emerging meta-pattern of frontier distillation pipelines and MoE teacher models. Logan provides inside context on Google's pricing strategy and the shift from input-tiered pricing to flat-rate tokens.5:47–8:23 · The hosts pushing back 2/10 Google's Experimental Releases and Coding with Cursor Composer Alessio presses on the operational meaning of Google's 'experimental' model labels and shares real-world developer feedback regarding Cursor Composer support. Logan clarifies Google's release-valve strategy and hot-swapping practices.8:23–12:31 · The hosts pushing back 3/10 The Reasoning Frontier and Native Scaling in Flash Thinking Swix and Logan discuss the shift from parameter scaling to inference-time compute scaling. Swix highlights the 2T parameter wall while Logan explains the co-scaling dynamics between core pre-trained base models and reinforcement learning.12:31–15:02 · The hosts pushing back 5/10 Long-Context Benchmarks and Abstraction vs Transparency in AI Swix presents empirical findings from his daily regression benchmarks where Gemini Flash outperforms rival models in long-context summarization. When Logan proposes turning it into an open leaderboard, Swix pushes back with product design rationale around interface abstraction.15:02–20:38 · The hosts pushing back 4/10 DeepSeek R1 Insights and Convergence of Reasoning Methods Swix cuts through generic high-level discussion on DeepSeek to probe specific technical architectural shifts like discarding Monte Carlo tree search and process reward models. Alessio then questions whether conventional chat interfaces represent legacy technical debt.20:39–24:42 · The hosts pushing back 4/10 Gemini Nano Integration and Local On-Device Model Adoption Swix explains the technical blocker behind Gemini Nano adoption, noting it remains gated behind Chrome Canary flags. They transition to discussing Project Astra, where Swix raises architectural trade-offs regarding KV caching, deletion constraints, and attention mechanisms versus RAG.24:42–25:47 · The hosts pushing back 1/10 Search as a Tool, Deep Research, and Online LLMs Logan outlines Google's Search as a Tool initiative and its role in deep research workflows. Swix categorizes the emerging online LLM market landscape before wrapping up with podcast cross-promotions.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 15:45 Interrupting high-level framing to demand technical specifics

Swix interrupts Logan's geopolitical overview of Chinese AI development to aggressively redirect the discussion to concrete algorithmic findings like the rejection of MCTS.

Hardest push from the hosts ▶ 13:56 Refusing public leaderboard suggestion in favor of product abstraction

Swix rejects Logan's suggestion to build an LMSYS-style leaderboard, citing design lessons from NotebookLM about hiding internal complexity to optimize user experience.

Biggest teaching moment ▶ 20:56 Explaining the real-world deployment state of Gemini Nano

When Alessio and Logan question why Gemini Nano adoption is slow, Swix informs them that the model remains restricted behind experimental browser feature flags in Chrome Canary.

The host holds their own ▶ 12:30 Demonstrating daily multi-model regression benchmarking

Swix outlines his proprietary daily regression testing framework across frontier models, demonstrating deep empirical knowledge of long-context token utilization differences between Gemini Flash and o3-mini.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Logan Kilpatrick Returns and Introduces Google Product Role 6412 Swix introduces Logan and articulates the emerging meta-pattern of frontier distillation pipelines and MoE teacher models. Logan provides inside context on Google's pricing strategy and the shift from input-tiered pricing to flat-rate tokens.
Google's Experimental Releases and Coding with Cursor Composer 5512 Alessio presses on the operational meaning of Google's 'experimental' model labels and shares real-world developer feedback regarding Cursor Composer support. Logan clarifies Google's release-valve strategy and hot-swapping practices.
The Reasoning Frontier and Native Scaling in Flash Thinking 6523 Swix and Logan discuss the shift from parameter scaling to inference-time compute scaling. Swix highlights the 2T parameter wall while Logan explains the co-scaling dynamics between core pre-trained base models and reinforcement learning.
Long-Context Benchmarks and Abstraction vs Transparency in AI 7425 Swix presents empirical findings from his daily regression benchmarks where Gemini Flash outperforms rival models in long-context summarization. When Logan proposes turning it into an open leaderboard, Swix pushes back with product design rationale around interface abstraction.
DeepSeek R1 Insights and Convergence of Reasoning Methods 7534 Swix cuts through generic high-level discussion on DeepSeek to probe specific technical architectural shifts like discarding Monte Carlo tree search and process reward models. Alessio then questions whether conventional chat interfaces represent legacy technical debt.
Gemini Nano Integration and Local On-Device Model Adoption 7624 Swix explains the technical blocker behind Gemini Nano adoption, noting it remains gated behind Chrome Canary flags. They transition to discussing Project Astra, where Swix raises architectural trade-offs regarding KV caching, deletion constraints, and attention mechanisms versus RAG.
Search as a Tool, Deep Research, and Online LLMs 5401 Logan outlines Google's Search as a Tool initiative and its role in deep research workflows. Swix categorizes the emerging online LLM market landscape before wrapping up with podcast cross-promotions.

Statements from this episode (10)

Assertion Supported
Google simplifies Gemini Flash pricing to flat 10 cents per million tokens
“So it went from seven and a half cents per million tokens to 10 cents. The sort of way that this was offset was we used to distinguish. I don't know if folks are familiar with this, but we used to distinguish based on input token volume. So it was like over a …”
Logan Kilpatrick Feb 28, 2025 ▶ 2:16
Prediction Not checkable as stated
Kilpatrick: Flash models will continue surpassing previous Pro model capabilities cheaply
“We keep seeing this jump where the capabilities of the pro models get superseded by the flash models as the next generation comes. So if you look at our two point O flash model is actually on like every dimension, a stronger and better model than the previous …”
Logan Kilpatrick Feb 28, 2025 ▶ 3:29
Prediction Not checkable as stated
Kilpatrick: Reasoning and test-time compute will scale faster medium-term
“And it feels like that is. More likely in the medium term going to be the thing that like continues to just like rapidly scale up relative to the other capabilities.”
Logan Kilpatrick Feb 28, 2025 ▶ 5:36
Disclosure
Kilpatrick: Google hot-swaps experimental models without allowing developers to pin old versions
“Framing it as experimental is, is really to try to signal that like one, you shouldn't use this thing in production. like it's You know, rate limits are usually the limitation to stop people from doing this, but to like the model might change. And in certain …”
Logan Kilpatrick Feb 28, 2025 ▶ 6:21
Assertion Supported
Noam Shazeer and Jack Rae co-lead Google DeepMind's reasoning effort
“Jack Ray. Yeah. He's been a long time deep mind research scientist, was previously a pre-training person. We actually overlapped at open AI together a little bit, and then is now back at deep mind with no co-leading the reasoning effort.”
Logan Kilpatrick Feb 28, 2025 ▶ 8:34
Prediction Not checkable as stated
Kilpatrick: Reasoning will solve multi-item retrieval in long context
“And like, it feels like, again, like back to this, the thread around these capabilities, like it feels like long context with reasoning is like finally going to be that thing where like, it actually just like blows the lid off of it. And like, it makes the use…”
Logan Kilpatrick Feb 28, 2025 ▶ 11:43
Prediction Not checkable as stated
Kilpatrick: Future browsers and IDEs will integrate live multimodal AI screen sharing
“If every browser in the world doesn't have this experience in the future, I'd be surprised if every IDE in the future doesn't have the ability to like, here, let me just show the model what's happening here, share my screen, talk to it live. Like, I think that…”
Logan Kilpatrick Feb 28, 2025 ▶ 18:31
Prediction Not checkable as stated
Next billions of AI users will onboard via SMS and phone
“I don't think those people are going to come in through, you know, some front end website somewhere. Like those people are going to come in through audio from a telephone to texting to email. Like that's just the most obvious outcome.”
Logan Kilpatrick Feb 28, 2025 ▶ 20:15
Opinion
Kilpatrick: Multimodal Live API provides 80% of Project Astra experience
“I think most probably like 80% of that experience you could get out of the box with the multimodal live API.”
Logan Kilpatrick Feb 28, 2025 ▶ 22:38
Opinion
Embeddings will not scale for cross-session memory in real-time AI systems
“I don't think that like embeddings are going to be able to scale to, I think they work well for some of this as like kind of the MVP version of the experience, but I think you're going to need a different experience and it's Probably something like really smar…”
Logan Kilpatrick Feb 28, 2025 ▶ 23:41
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.