Aug 16, 2025 · 42m · a16z

Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building

Jack Parker-Holder · 17m spoken Shlomi Fruchter · 12m spoken Justine Moore · 2m spoken Erik Torenberg · 17s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The a16z Podcast, Google DeepMind researchers Shlomi Fruchter and Jack Parker-Holder discuss Genie 3, a groundbreaking real-time foundation world model capable of generating interactive 3D environments from simple text prompts. They explore its architectural breakthroughs, emergent physical simulation capabilities, persistent spatial memory, and broad potential applications ranging from creative software to robotics and embodied AI.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The host as informed peer 2.7 Guest teaching 3.6 Guest disagreement 0.4 The host pushing back 0.1
05100:0015:0030:000:00–3:30 · The host as informed peer 2/10 Genie 3 World Generation Capabilities Highlight Marco asks about interactive memory and data collection strategies. Jack explains how Genie 3 synthesized progress from Genie II, GameNGen, and Veo II.3:30–6:28 · The host as informed peer 2/10 The Magic of Real-Time Control and Low Latency Justine highlights the perfect release timing relative to viral X/Reddit trends and asks about primary use cases. Shlomi emphasizes that core generative world capabilities underpin all applications equally.6:28–12:33 · The host as informed peer 3/10 Reinforcement Learning Origins and Unlimited Environments Mike points out a specific visual detail in the blog post regarding paint persistence. Shlomi educates the hosts on why they avoided explicit 3D representations like NeRFs in favor of pure generative frame-by-frame modeling.12:33–15:16 · The host as informed peer 2/10 Memory Duration Trade-offs and Real-Time Latency Marco asks if scaling Genie brings emergent LLM-like reasoning. Shlomi gently corrects the framing, noting that while world understanding improves, 'reasoning' is not the accurate term for this architecture.15:16–18:25 · The host as informed peer 3/10 Simulating Environmental Physics and Unlikely Prompts Justine notes specific physics dynamics like character swimming. Shlomi explains the technical tension between world consistency and following unlikely user text prompts.18:25–21:56 · The host as informed peer 3/10 Direct Text Controllability and World Generation Mike asks why this was branded Genie 3 instead of Veo 3 Real-Time. Shlomi highlights functional differences, such as interactive navigation and lack of audio, justifying separate research tracks.21:56–26:19 · The host as informed peer 4/10 The Convergence vs. Divergence of Generative Modalities Mike sets up a high-level technical question regarding whether video generation and world models will converge or diverge. Jack and Shlomi break down why separate technical trade-offs make distinct models optimal for now.26:19–28:19 · The host as informed peer 2/10 Research-Driven Innovation Versus Downstream Use Cases Justine asks whether research direction is guided by downstream applications or raw technical exploration. Shlomi clarifies that pushing capability boundaries comes first, with downstream applications following organically.28:19–32:22 · The host as informed peer 2/10 Future Horizons: Multiplayer Worlds and Experiential AI Marco proposes potential future directions like multiplayer world models. Shlomi discusses broader future horizons including therapeutic simulations for phobias and public speaking.32:22–37:58 · The host as informed peer 4/10 Transforming Robotics and Embodied AI Training Mike cites Demis Hassabis's recent comments regarding SIMA and Genie agent composability in robotics. Jack provides a deep dive on why generative world models bridge the sim-to-real gap far better than standard lab simulations.37:58–40:41 · The host as informed peer 3/10 Developer Access Timelines and the World Model S-Curve Justine introduces the framework of S-curves to evaluate where world models sit relative to LLMs. Jack and Shlomi contextualize current progress and answer a humorous closing question about living in a simulation.0:00–3:30 · Guest teaching 3/10 Genie 3 World Generation Capabilities Highlight Marco asks about interactive memory and data collection strategies. Jack explains how Genie 3 synthesized progress from Genie II, GameNGen, and Veo II.3:30–6:28 · Guest teaching 2/10 The Magic of Real-Time Control and Low Latency Justine highlights the perfect release timing relative to viral X/Reddit trends and asks about primary use cases. Shlomi emphasizes that core generative world capabilities underpin all applications equally.6:28–12:33 · Guest teaching 4/10 Reinforcement Learning Origins and Unlimited Environments Mike points out a specific visual detail in the blog post regarding paint persistence. Shlomi educates the hosts on why they avoided explicit 3D representations like NeRFs in favor of pure generative frame-by-frame modeling.12:33–15:16 · Guest teaching 5/10 Memory Duration Trade-offs and Real-Time Latency Marco asks if scaling Genie brings emergent LLM-like reasoning. Shlomi gently corrects the framing, noting that while world understanding improves, 'reasoning' is not the accurate term for this architecture.15:16–18:25 · Guest teaching 4/10 Simulating Environmental Physics and Unlikely Prompts Justine notes specific physics dynamics like character swimming. Shlomi explains the technical tension between world consistency and following unlikely user text prompts.18:25–21:56 · Guest teaching 4/10 Direct Text Controllability and World Generation Mike asks why this was branded Genie 3 instead of Veo 3 Real-Time. Shlomi highlights functional differences, such as interactive navigation and lack of audio, justifying separate research tracks.21:56–26:19 · Guest teaching 4/10 The Convergence vs. Divergence of Generative Modalities Mike sets up a high-level technical question regarding whether video generation and world models will converge or diverge. Jack and Shlomi break down why separate technical trade-offs make distinct models optimal for now.26:19–28:19 · Guest teaching 3/10 Research-Driven Innovation Versus Downstream Use Cases Justine asks whether research direction is guided by downstream applications or raw technical exploration. Shlomi clarifies that pushing capability boundaries comes first, with downstream applications following organically.28:19–32:22 · Guest teaching 3/10 Future Horizons: Multiplayer Worlds and Experiential AI Marco proposes potential future directions like multiplayer world models. Shlomi discusses broader future horizons including therapeutic simulations for phobias and public speaking.32:22–37:58 · Guest teaching 5/10 Transforming Robotics and Embodied AI Training Mike cites Demis Hassabis's recent comments regarding SIMA and Genie agent composability in robotics. Jack provides a deep dive on why generative world models bridge the sim-to-real gap far better than standard lab simulations.37:58–40:41 · Guest teaching 3/10 Developer Access Timelines and the World Model S-Curve Justine introduces the framework of S-curves to evaluate where world models sit relative to LLMs. Jack and Shlomi contextualize current progress and answer a humorous closing question about living in a simulation.0:00–3:30 · Guest disagreement 0/10 Genie 3 World Generation Capabilities Highlight Marco asks about interactive memory and data collection strategies. Jack explains how Genie 3 synthesized progress from Genie II, GameNGen, and Veo II.3:30–6:28 · Guest disagreement 0/10 The Magic of Real-Time Control and Low Latency Justine highlights the perfect release timing relative to viral X/Reddit trends and asks about primary use cases. Shlomi emphasizes that core generative world capabilities underpin all applications equally.6:28–12:33 · Guest disagreement 0/10 Reinforcement Learning Origins and Unlimited Environments Mike points out a specific visual detail in the blog post regarding paint persistence. Shlomi educates the hosts on why they avoided explicit 3D representations like NeRFs in favor of pure generative frame-by-frame modeling.12:33–15:16 · Guest disagreement 1/10 Memory Duration Trade-offs and Real-Time Latency Marco asks if scaling Genie brings emergent LLM-like reasoning. Shlomi gently corrects the framing, noting that while world understanding improves, 'reasoning' is not the accurate term for this architecture.15:16–18:25 · Guest disagreement 0/10 Simulating Environmental Physics and Unlikely Prompts Justine notes specific physics dynamics like character swimming. Shlomi explains the technical tension between world consistency and following unlikely user text prompts.18:25–21:56 · Guest disagreement 1/10 Direct Text Controllability and World Generation Mike asks why this was branded Genie 3 instead of Veo 3 Real-Time. Shlomi highlights functional differences, such as interactive navigation and lack of audio, justifying separate research tracks.21:56–26:19 · Guest disagreement 1/10 The Convergence vs. Divergence of Generative Modalities Mike sets up a high-level technical question regarding whether video generation and world models will converge or diverge. Jack and Shlomi break down why separate technical trade-offs make distinct models optimal for now.26:19–28:19 · Guest disagreement 0/10 Research-Driven Innovation Versus Downstream Use Cases Justine asks whether research direction is guided by downstream applications or raw technical exploration. Shlomi clarifies that pushing capability boundaries comes first, with downstream applications following organically.28:19–32:22 · Guest disagreement 0/10 Future Horizons: Multiplayer Worlds and Experiential AI Marco proposes potential future directions like multiplayer world models. Shlomi discusses broader future horizons including therapeutic simulations for phobias and public speaking.32:22–37:58 · Guest disagreement 1/10 Transforming Robotics and Embodied AI Training Mike cites Demis Hassabis's recent comments regarding SIMA and Genie agent composability in robotics. Jack provides a deep dive on why generative world models bridge the sim-to-real gap far better than standard lab simulations.37:58–40:41 · Guest disagreement 0/10 Developer Access Timelines and the World Model S-Curve Justine introduces the framework of S-curves to evaluate where world models sit relative to LLMs. Jack and Shlomi contextualize current progress and answer a humorous closing question about living in a simulation.0:00–3:30 · The host pushing back 0/10 Genie 3 World Generation Capabilities Highlight Marco asks about interactive memory and data collection strategies. Jack explains how Genie 3 synthesized progress from Genie II, GameNGen, and Veo II.3:30–6:28 · The host pushing back 0/10 The Magic of Real-Time Control and Low Latency Justine highlights the perfect release timing relative to viral X/Reddit trends and asks about primary use cases. Shlomi emphasizes that core generative world capabilities underpin all applications equally.6:28–12:33 · The host pushing back 0/10 Reinforcement Learning Origins and Unlimited Environments Mike points out a specific visual detail in the blog post regarding paint persistence. Shlomi educates the hosts on why they avoided explicit 3D representations like NeRFs in favor of pure generative frame-by-frame modeling.12:33–15:16 · The host pushing back 0/10 Memory Duration Trade-offs and Real-Time Latency Marco asks if scaling Genie brings emergent LLM-like reasoning. Shlomi gently corrects the framing, noting that while world understanding improves, 'reasoning' is not the accurate term for this architecture.15:16–18:25 · The host pushing back 0/10 Simulating Environmental Physics and Unlikely Prompts Justine notes specific physics dynamics like character swimming. Shlomi explains the technical tension between world consistency and following unlikely user text prompts.18:25–21:56 · The host pushing back 1/10 Direct Text Controllability and World Generation Mike asks why this was branded Genie 3 instead of Veo 3 Real-Time. Shlomi highlights functional differences, such as interactive navigation and lack of audio, justifying separate research tracks.21:56–26:19 · The host pushing back 0/10 The Convergence vs. Divergence of Generative Modalities Mike sets up a high-level technical question regarding whether video generation and world models will converge or diverge. Jack and Shlomi break down why separate technical trade-offs make distinct models optimal for now.26:19–28:19 · The host pushing back 0/10 Research-Driven Innovation Versus Downstream Use Cases Justine asks whether research direction is guided by downstream applications or raw technical exploration. Shlomi clarifies that pushing capability boundaries comes first, with downstream applications following organically.28:19–32:22 · The host pushing back 0/10 Future Horizons: Multiplayer Worlds and Experiential AI Marco proposes potential future directions like multiplayer world models. Shlomi discusses broader future horizons including therapeutic simulations for phobias and public speaking.32:22–37:58 · The host pushing back 0/10 Transforming Robotics and Embodied AI Training Mike cites Demis Hassabis's recent comments regarding SIMA and Genie agent composability in robotics. Jack provides a deep dive on why generative world models bridge the sim-to-real gap far better than standard lab simulations.37:58–40:41 · The host pushing back 0/10 Developer Access Timelines and the World Model S-Curve Justine introduces the framework of S-curves to evaluate where world models sit relative to LLMs. Jack and Shlomi contextualize current progress and answer a humorous closing question about living in a simulation.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 0% · guest 100%0:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%36:00 · the host 0% · guest 100%36:00 · the host 0% · guest 100%39:00 · the host 0% · guest 100%39:00 · the host 0% · guest 100%42:00 · the host 0% · guest 100%42:00 · the host 0% · guest 100%
Sharpest disagreement ▶ 13:45 Rejecting LLM Reasoning Premise

Shlomi politely rejects Marco's premise that scaling Genie creates emergent LLM-style reasoning, clarifying that world understanding is fundamentally distinct.

Hardest push from the host ▶ 20:48 Challenging Genie 3 vs. Veo 3 Branding

Host Mike directly presses the guests on why the project wasn't branded as Veo 3 Real-Time given the overlap in capabilities.

Biggest teaching moment ▶ 11:33 Explaining Generative vs. NeRF Representations

Shlomi educates the host on why avoiding explicit 3D geometric representations like NeRFs was crucial for maintaining spatial generalization across generated frames.

The host holds their own ▶ 32:50 Citing Hassabis on SIMA Agent Composability

Host Mike demonstrates strong background knowledge by referencing Demis Hassabis's recent remarks on combining SIMA agents with Genie world models for robotics training.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
Genie 3 World Generation Capabilities Highlight 2300 Marco asks about interactive memory and data collection strategies. Jack explains how Genie 3 synthesized progress from Genie II, GameNGen, and Veo II.
The Magic of Real-Time Control and Low Latency 2200 Justine highlights the perfect release timing relative to viral X/Reddit trends and asks about primary use cases. Shlomi emphasizes that core generative world capabilities underpin all applications equally.
Reinforcement Learning Origins and Unlimited Environments 3400 Mike points out a specific visual detail in the blog post regarding paint persistence. Shlomi educates the hosts on why they avoided explicit 3D representations like NeRFs in favor of pure generative frame-by-frame modeling.
Memory Duration Trade-offs and Real-Time Latency 2510 Marco asks if scaling Genie brings emergent LLM-like reasoning. Shlomi gently corrects the framing, noting that while world understanding improves, 'reasoning' is not the accurate term for this architecture.
Simulating Environmental Physics and Unlikely Prompts 3400 Justine notes specific physics dynamics like character swimming. Shlomi explains the technical tension between world consistency and following unlikely user text prompts.
Direct Text Controllability and World Generation 3411 Mike asks why this was branded Genie 3 instead of Veo 3 Real-Time. Shlomi highlights functional differences, such as interactive navigation and lack of audio, justifying separate research tracks.
The Convergence vs. Divergence of Generative Modalities 4410 Mike sets up a high-level technical question regarding whether video generation and world models will converge or diverge. Jack and Shlomi break down why separate technical trade-offs make distinct models optimal for now.
Research-Driven Innovation Versus Downstream Use Cases 2300 Justine asks whether research direction is guided by downstream applications or raw technical exploration. Shlomi clarifies that pushing capability boundaries comes first, with downstream applications following organically.
Future Horizons: Multiplayer Worlds and Experiential AI 2300 Marco proposes potential future directions like multiplayer world models. Shlomi discusses broader future horizons including therapeutic simulations for phobias and public speaking.
Transforming Robotics and Embodied AI Training 4510 Mike cites Demis Hassabis's recent comments regarding SIMA and Genie agent composability in robotics. Jack provides a deep dive on why generative world models bridge the sim-to-real gap far better than standard lab simulations.
Developer Access Timelines and the World Model S-Curve 3300 Justine introduces the framework of S-curves to evaluate where world models sit relative to LLMs. Jack and Shlomi contextualize current progress and answer a humorous closing question about living in a simulation.

Statements from this episode (18)

Assertion Not checkable as stated
Parker-Holder: Genie 3 outputs look real to non-expert viewers
“And it, it's at the point where, like, a human who is not an expert, like, will watch it and think it looks real, right?”
Jack Parker-Holder Aug 16, 2025 ▶ 0:16
Assertion Supported
Fruchter: DeepMind achieved real-time interactive environment generation
“We weren't sure, like, how, how big it's gonna be but we definitely felt, I felt definitely that we have something that was kind of like for a long time coming, basically being able to generate environments in real time. I think a lot of work that was done in,…”
Shlomi Fruchter Aug 16, 2025 ▶ 0:40
Assertion Partly supported
Parker-Holder: Veo 2 launched one week after Genie 2
“So we had this Genie II project that was much more sort of like, three environments that it could generate, and it wasn't super high quality. It felt like it coming from Genie I, but it wasn't the same quality as things like VO II, which you know, the Saints o…”
Jack Parker-Holder Aug 16, 2025 ▶ 2:09
Disclosure
Parker-Holder: DeepMind targeted minute-plus spatial memory for Genie 3
“And then for Genie three, we basically went much more ambitious on the same sort of approach, right? And we made it like a headline goal for ourselves. It's like, can we make the memory be what it is, right? We said we want minute plus memory and real time and…”
Jack Parker-Holder Aug 16, 2025 ▶ 10:36
Disclosure
Fruchter: Genie 3 generates environments without explicit 3D representations
“One thing that we didn't want to do, and we didn't want to build an explicit representation, right? So there are definitely methods that are able to achieve consistency, and they did that through an explicit, some three, you know, their nerves, and other metho…”
Shlomi Fruchter Aug 16, 2025 ▶ 11:33
Assertion Supported
Fruchter: Genie 3 currently caps spatial memory retention at one minute
“There was not like, there is no like fundamental limitation, but Currently the current design we limited to one minute of this type of memories.”
Shlomi Fruchter Aug 16, 2025 ▶ 12:46
Insight
Parker-Holder: Genie 3 terrain physics emerge from scale, not explicit logic
“And I think that that really is the property of scale and breadth of training. So this is very much like an emergent thing. I don't think there's anything Like really specific we do for this, right?”
Jack Parker-Holder Aug 16, 2025 ▶ 16:16
Assertion Supported
Parker-Holder: Genie 2 suffered transfer issues from image prompting
“So I think that that's actually a really important capability that we didn't have with Genie too as well, right? Because we relied on image prompting. And so, there was some transfer issue, like, where you rely on imagine to generate the image, and that often …”
Jack Parker-Holder Aug 16, 2025 ▶ 19:14
Assertion Supported
Fruchter: Veo cannot currently navigate environments or take actions like Genie
“Genie allows you to navigate environment and then maybe take actions. And that's not something that Veo at this point can do.”
Shlomi Fruchter Aug 16, 2025 ▶ 21:05
Disclosure
Fruchter: Genie 3 is a research preview, not a commercial product
“Gini Free is pretty much a research preview, right? It's not something we are releasing at this point.”
Shlomi Fruchter Aug 16, 2025 ▶ 21:51
Disclosure
Parker-Holder: DeepMind kept Veo 3 and Genie 3 as separate projects
“We obviously made a choice that we want VO three and Gini three to be separate projects this year, right?”
Jack Parker-Holder Aug 16, 2025 ▶ 24:39
Assertion Not checkable as stated
Parker-Holder: Veo 3 achieves higher visual quality than Genie 3
“VO three is clearly a higher quality threshold than Gene three, right?”
Jack Parker-Holder Aug 16, 2025 ▶ 24:57
Opinion
Parker-Holder: World models are the fastest path to embodied AGI
“I still think, honestly, for my, what I'm excited about for AGI, and which is more embodied agents I really believe this is the fastest path to getting these agents, like, in the real world”
Jack Parker-Holder Aug 16, 2025 ▶ 29:40
Opinion
Fruchter: Current world models remain far from accurately simulating reality
“I think they are very far from actually simulating the world accurately and being able to do, to kind of put a person in there and then do whatever they want.”
Shlomi Fruchter Aug 16, 2025 ▶ 30:39
Disclosure
Parker-Holder: Genie 3 is an environment simulator, not an autonomous agent
“We designed it to be an environment rather than an agent, right? So Genie three is very much like an environment model. Like we don't see it as like an agent itself that can like think and act in the world. It's more just a general purpose sort of simulator in…”
Jack Parker-Holder Aug 16, 2025 ▶ 33:44
Assertion Not checkable as stated
Parker-Holder: Top robotics simulators still face significant sim-to-real gaps
“The robotic simulations are even the best ones, and we have some of the best ones at DeepMind and Majoko, right, which we work with. They're still quite far away from the real world, right? And so you have the sim to real gap.”
Jack Parker-Holder Aug 16, 2025 ▶ 34:32
Insight
Parker-Holder: Genie 3 combines real-world data with simulation for robotics
“Really what we think with Genie three is it's the best of both, right? Because you're taking a real world data driven approach, right? But then you've got the ability to learn in simulation. So It kind of combines the good parts of each of those. And so that's…”
Jack Parker-Holder Aug 16, 2025 ▶ 36:12
Opinion
Fruchter: A simulated universe would run on continuous, analog hardware
“If we live in a simulation, my take is that it doesn't run on our current hardware because it's analog and not like, you know, it's continuous.”
Shlomi Fruchter Aug 16, 2025 ▶ 40:54
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.