Feb 17, 2024 · 39m · a16z

Text to Video: The Next Leap in AI Generation

Robin Rombach · 13m spoken Andreas Blattman · 10m spoken Anshin Mita · 9m spoken Anjney Midha · 2m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the a16z podcast, host Anshin Mita interviews Stability AI researchers Robin Rombach and Andreas Blattman on the technical breakthroughs, open-source impact, and physical modeling challenges behind Stable Video Diffusion.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The host as informed peer 3.3 Guest teaching 4.6 Guest disagreement 0.2 The host pushing back 0.0
05100:0010:0020:0030:002:12–4:35 · The host as informed peer 2/10 Legal Disclaimers and Compliance Information The host opens with high-level introductory questions regarding Stable Diffusion and Stable Video Diffusion. Guest Robin Rombach provides a brief technical baseline explaining diffusion models and university origins.4:35–8:05 · The host as informed peer 3/10 Diffusion Models vs Autoregressive Models and Distillation The host asks open-ended clarifying questions about diffusion versus autoregressive models. Guests educate the host on noise transformation versus sequence generation, as well as single-step distillation.8:05–12:00 · The host as informed peer 3/10 Evolution of Image Models and Open Source The host synthesizes the growth of open-source Lego block developments. Guests share early development details regarding latent diffusion training constraints and classifier-free guidance.12:00–15:19 · The host as informed peer 5/10 Physics, 3D Understanding, and Video Modeling The host demonstrates insight by linking the guests physics and engineering backgrounds to implicit 3D world modeling in video systems. Guests fully agree and expand on world simulation.15:19–19:40 · The host as informed peer 3/10 Data Pipelines and Training Infrastructure for Video The host prompts the guests on hardware and infrastructure differences between image and video models. Guests detailed CPU codec bottlenecks, batch size requirements, and cluster failure risks.19:40–24:11 · The host as informed peer 4/10 Three-Stage Training and Multi-View 3D Synthesis The host raises the problem of 3D structural consistency in video generation. Guest Robin Rombach breaks down the three-stage training process and fine-tuning on multi-view orbit data.24:11–28:47 · The host as informed peer 4/10 Fine-Grained Control with LoRAs and Future Video Editing The host queries whether future video production will rely on managing hundreds of LoRAs. Robin gently reframes this, pointing out that managing extensive LoRA libraries is unscalable compared to direct text and motion prompts.28:47–31:12 · The host as informed peer 2/10 Community Exploration and Favorite Generated Memes A brief lighthearted segment where the host asks about community creations. Guests share favorite examples such as animated memes and classical artwork.31:12–36:14 · The host as informed peer 4/10 Future Frontiers, Real-Time Inference, and Compute Constraints The host asks about hardware wishlists and research bottlenecks. Guests explain how compute constraints drive architectural breakthroughs, which the host summarizes as no constraints, no creativity.36:14–39:06 · The host as informed peer 3/10 Research Philosophy, Open Source Impact, and Conclusion The host asks how a compact team competes with major tech monopolies. Guests express their passion for open-source research, and the host points out OpenAI citing their research in DALL-E 3.2:12–4:35 · Guest teaching 3/10 Legal Disclaimers and Compliance Information The host opens with high-level introductory questions regarding Stable Diffusion and Stable Video Diffusion. Guest Robin Rombach provides a brief technical baseline explaining diffusion models and university origins.4:35–8:05 · Guest teaching 6/10 Diffusion Models vs Autoregressive Models and Distillation The host asks open-ended clarifying questions about diffusion versus autoregressive models. Guests educate the host on noise transformation versus sequence generation, as well as single-step distillation.8:05–12:00 · Guest teaching 5/10 Evolution of Image Models and Open Source The host synthesizes the growth of open-source Lego block developments. Guests share early development details regarding latent diffusion training constraints and classifier-free guidance.12:00–15:19 · Guest teaching 5/10 Physics, 3D Understanding, and Video Modeling The host demonstrates insight by linking the guests physics and engineering backgrounds to implicit 3D world modeling in video systems. Guests fully agree and expand on world simulation.15:19–19:40 · Guest teaching 6/10 Data Pipelines and Training Infrastructure for Video The host prompts the guests on hardware and infrastructure differences between image and video models. Guests detailed CPU codec bottlenecks, batch size requirements, and cluster failure risks.19:40–24:11 · Guest teaching 6/10 Three-Stage Training and Multi-View 3D Synthesis The host raises the problem of 3D structural consistency in video generation. Guest Robin Rombach breaks down the three-stage training process and fine-tuning on multi-view orbit data.24:11–28:47 · Guest teaching 5/10 Fine-Grained Control with LoRAs and Future Video Editing The host queries whether future video production will rely on managing hundreds of LoRAs. Robin gently reframes this, pointing out that managing extensive LoRA libraries is unscalable compared to direct text and motion prompts.28:47–31:12 · Guest teaching 3/10 Community Exploration and Favorite Generated Memes A brief lighthearted segment where the host asks about community creations. Guests share favorite examples such as animated memes and classical artwork.31:12–36:14 · Guest teaching 5/10 Future Frontiers, Real-Time Inference, and Compute Constraints The host asks about hardware wishlists and research bottlenecks. Guests explain how compute constraints drive architectural breakthroughs, which the host summarizes as no constraints, no creativity.36:14–39:06 · Guest teaching 2/10 Research Philosophy, Open Source Impact, and Conclusion The host asks how a compact team competes with major tech monopolies. Guests express their passion for open-source research, and the host points out OpenAI citing their research in DALL-E 3.2:12–4:35 · Guest disagreement 0/10 Legal Disclaimers and Compliance Information The host opens with high-level introductory questions regarding Stable Diffusion and Stable Video Diffusion. Guest Robin Rombach provides a brief technical baseline explaining diffusion models and university origins.4:35–8:05 · Guest disagreement 0/10 Diffusion Models vs Autoregressive Models and Distillation The host asks open-ended clarifying questions about diffusion versus autoregressive models. Guests educate the host on noise transformation versus sequence generation, as well as single-step distillation.8:05–12:00 · Guest disagreement 1/10 Evolution of Image Models and Open Source The host synthesizes the growth of open-source Lego block developments. Guests share early development details regarding latent diffusion training constraints and classifier-free guidance.12:00–15:19 · Guest disagreement 0/10 Physics, 3D Understanding, and Video Modeling The host demonstrates insight by linking the guests physics and engineering backgrounds to implicit 3D world modeling in video systems. Guests fully agree and expand on world simulation.15:19–19:40 · Guest disagreement 0/10 Data Pipelines and Training Infrastructure for Video The host prompts the guests on hardware and infrastructure differences between image and video models. Guests detailed CPU codec bottlenecks, batch size requirements, and cluster failure risks.19:40–24:11 · Guest disagreement 0/10 Three-Stage Training and Multi-View 3D Synthesis The host raises the problem of 3D structural consistency in video generation. Guest Robin Rombach breaks down the three-stage training process and fine-tuning on multi-view orbit data.24:11–28:47 · Guest disagreement 1/10 Fine-Grained Control with LoRAs and Future Video Editing The host queries whether future video production will rely on managing hundreds of LoRAs. Robin gently reframes this, pointing out that managing extensive LoRA libraries is unscalable compared to direct text and motion prompts.28:47–31:12 · Guest disagreement 0/10 Community Exploration and Favorite Generated Memes A brief lighthearted segment where the host asks about community creations. Guests share favorite examples such as animated memes and classical artwork.31:12–36:14 · Guest disagreement 0/10 Future Frontiers, Real-Time Inference, and Compute Constraints The host asks about hardware wishlists and research bottlenecks. Guests explain how compute constraints drive architectural breakthroughs, which the host summarizes as no constraints, no creativity.36:14–39:06 · Guest disagreement 0/10 Research Philosophy, Open Source Impact, and Conclusion The host asks how a compact team competes with major tech monopolies. Guests express their passion for open-source research, and the host points out OpenAI citing their research in DALL-E 3.2:12–4:35 · The host pushing back 0/10 Legal Disclaimers and Compliance Information The host opens with high-level introductory questions regarding Stable Diffusion and Stable Video Diffusion. Guest Robin Rombach provides a brief technical baseline explaining diffusion models and university origins.4:35–8:05 · The host pushing back 0/10 Diffusion Models vs Autoregressive Models and Distillation The host asks open-ended clarifying questions about diffusion versus autoregressive models. Guests educate the host on noise transformation versus sequence generation, as well as single-step distillation.8:05–12:00 · The host pushing back 0/10 Evolution of Image Models and Open Source The host synthesizes the growth of open-source Lego block developments. Guests share early development details regarding latent diffusion training constraints and classifier-free guidance.12:00–15:19 · The host pushing back 0/10 Physics, 3D Understanding, and Video Modeling The host demonstrates insight by linking the guests physics and engineering backgrounds to implicit 3D world modeling in video systems. Guests fully agree and expand on world simulation.15:19–19:40 · The host pushing back 0/10 Data Pipelines and Training Infrastructure for Video The host prompts the guests on hardware and infrastructure differences between image and video models. Guests detailed CPU codec bottlenecks, batch size requirements, and cluster failure risks.19:40–24:11 · The host pushing back 0/10 Three-Stage Training and Multi-View 3D Synthesis The host raises the problem of 3D structural consistency in video generation. Guest Robin Rombach breaks down the three-stage training process and fine-tuning on multi-view orbit data.24:11–28:47 · The host pushing back 0/10 Fine-Grained Control with LoRAs and Future Video Editing The host queries whether future video production will rely on managing hundreds of LoRAs. Robin gently reframes this, pointing out that managing extensive LoRA libraries is unscalable compared to direct text and motion prompts.28:47–31:12 · The host pushing back 0/10 Community Exploration and Favorite Generated Memes A brief lighthearted segment where the host asks about community creations. Guests share favorite examples such as animated memes and classical artwork.31:12–36:14 · The host pushing back 0/10 Future Frontiers, Real-Time Inference, and Compute Constraints The host asks about hardware wishlists and research bottlenecks. Guests explain how compute constraints drive architectural breakthroughs, which the host summarizes as no constraints, no creativity.36:14–39:06 · The host pushing back 0/10 Research Philosophy, Open Source Impact, and Conclusion The host asks how a compact team competes with major tech monopolies. Guests express their passion for open-source research, and the host points out OpenAI citing their research in DALL-E 3.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 0% · guest 100%0:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%36:00 · the host 0% · guest 100%36:00 · the host 0% · guest 100%39:00 · the host 0% · guest 0%39:00 · the host 0% · guest 0%
Sharpest disagreement ▶ 27:09 Gentle reframe on LoRA library scalability

Robin gently counters the host premise that directors will rely on hundreds of LoRAs, explaining that managing such libraries is not a scalable creation approach.

Hardest push from the host ▶ 13:00 Host reframes discussion through physics perspective

The host actively reframes the discussion around the guests academic backgrounds in physics and engineering to pivot into 3D world modeling.

Biggest teaching moment ▶ 15:45 Detailed breakdown of video training hardware bottlenecks

Andreas educates the host on how CPU decoding speeds unexpectedly became a critical bottleneck during high-resolution multi-GPU video model training.

The host holds their own ▶ 34:00 Host connects guests constrained history to broader AI ecosystem

The host demonstrates deep domain awareness by highlighting how the guests early single-GPU constraints inspired broader university research groups.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
Legal Disclaimers and Compliance Information 2300 The host opens with high-level introductory questions regarding Stable Diffusion and Stable Video Diffusion. Guest Robin Rombach provides a brief technical baseline explaining diffusion models and university origins.
Diffusion Models vs Autoregressive Models and Distillation 3600 The host asks open-ended clarifying questions about diffusion versus autoregressive models. Guests educate the host on noise transformation versus sequence generation, as well as single-step distillation.
Evolution of Image Models and Open Source 3510 The host synthesizes the growth of open-source Lego block developments. Guests share early development details regarding latent diffusion training constraints and classifier-free guidance.
Physics, 3D Understanding, and Video Modeling 5500 The host demonstrates insight by linking the guests physics and engineering backgrounds to implicit 3D world modeling in video systems. Guests fully agree and expand on world simulation.
Data Pipelines and Training Infrastructure for Video 3600 The host prompts the guests on hardware and infrastructure differences between image and video models. Guests detailed CPU codec bottlenecks, batch size requirements, and cluster failure risks.
Three-Stage Training and Multi-View 3D Synthesis 4600 The host raises the problem of 3D structural consistency in video generation. Guest Robin Rombach breaks down the three-stage training process and fine-tuning on multi-view orbit data.
Fine-Grained Control with LoRAs and Future Video Editing 4510 The host queries whether future video production will rely on managing hundreds of LoRAs. Robin gently reframes this, pointing out that managing extensive LoRA libraries is unscalable compared to direct text and motion prompts.
Community Exploration and Favorite Generated Memes 2300 A brief lighthearted segment where the host asks about community creations. Guests share favorite examples such as animated memes and classical artwork.
Future Frontiers, Real-Time Inference, and Compute Constraints 4500 The host asks about hardware wishlists and research bottlenecks. Guests explain how compute constraints drive architectural breakthroughs, which the host summarizes as no constraints, no creativity.
Research Philosophy, Open Source Impact, and Conclusion 3200 The host asks how a compact team competes with major tech monopolies. Guests express their passion for open-source research, and the host points out OpenAI citing their research in DALL-E 3.

Statements from this episode (13)

Assertion Supported
Rombach: Stable Diffusion's core technique was developed during university research
“It's based on a technique that we developed while we were still at the university.”
Robin Rombach Feb 17, 2024 ▶ 3:23
Disclosure
Rombach: Stability AI distillation research enables single-step diffusion generation
“We ourselves, we have published a distillation work a week ago that actually shows that you can go as low as one sampling step, which is, I would say like a big advantage of these diffusion models.”
Robin Rombach Feb 17, 2024 ▶ 7:05
What-if
Blattman: AI image progress depended on open-sourcing Stable Diffusion
“Open sourcing, like a foundation model as stable diffusion initially, that led to a whole lot of research on these models, which was, yeah, it was in, in retrospect, extremely important to do this. I think otherwise we wouldn't have seen the improvements we sa…”
Andreas Blattman Feb 17, 2024 ▶ 9:26
Insight
Blattman: Video AI models must learn physics and 3D geometry
“I think video is, is like an awesome kind of data because you, to solve that task, to solve video generation, a model needs to Learn much about like physical properties of the world of like the physical foundations of the world that there is so much without kn…”
Andreas Blattman Feb 17, 2024 ▶ 12:01
Prediction Not checkable as stated
Rombach: Video AI models scaled like LLMs will gain world understanding
“Having something like we are seeing in language modeling, but trained on pixels on videos will probably give like super interesting downstream behavior to not, not only like generating videos, but also understanding of the world.”
Robin Rombach Feb 17, 2024 ▶ 14:05
Disclosure
Rombach: Stable Video Diffusion took six months to develop
“I would say roughly half a year and like for this model that we just put out, I think the main challenge was that we actually, Yeah, I had to scale the data set and the data loading.”
Robin Rombach Feb 17, 2024 ▶ 15:20
Insight
Blattman: High batch sizes are critical for training diffusion models
“So for diffusion models, it's really important to have a high batch size, because the gradients gets, like, you can approximate the gradient, which thrives the learning much better if the batch size is higher. And especially for diffusion models, it's like rea…”
Andreas Blattman Feb 17, 2024 ▶ 17:43
Assertion Supported
Rombach: Pre-trained video models learn 3D synthesis faster than image models
“We showed that it's actually like helpful to incorporate like this implicit three D knowledge that Knowledge that is captured in all of the videos into the model, and then the model can learn much quicker than if you start from the pure image model.”
Robin Rombach Feb 17, 2024 ▶ 23:07
Opinion
Rombach: Managing hundreds of LoRAs is unscalable for video control
“Maintaining like a library of hundreds of LoRa's is maybe not like the, Most scalable approach.”
Robin Rombach Feb 17, 2024 ▶ 27:09
Prediction Not checkable as stated
Rombach: AI video generation will evolve into real-time interactive experiences
“Because then this will become more like, I don't know, sometimes I think about this as like a video game, right? You type your prompt, and you immediately see what happens given your input view, and I think this might be a super nice user experience, actually.”
Robin Rombach Feb 17, 2024 ▶ 28:23
Assertion Supported
Blattman: Stable Video Diffusion acquired 3D reasoning within 2,000 iterations
“We showed that by our three D fine tuning. This was, by the way, this was completely surprising for me seeing that model after a 1002 thousand iterations, like already getting what is like, three D reasoning or like explicit three D reasoning.”
Andreas Blattman Feb 17, 2024 ▶ 29:27
Insight
Rombach: Compute constraints drive AI innovation more than scaling hardware
“If you only rely on like more compute it's a bit boring. I think like compute constraints can also Drive innovation, right? So for example, the latent diffusion framework, we developed it at the university because we just, like, we had, like, single GPUs where…”
Robin Rombach Feb 17, 2024 ▶ 34:02
Assertion Supported
Rombach: DALL-E 3 uses an autoencoder trained on a single GPU
“Dolly three uses a model that, like, the autoencoder that was trained on a single GPU.”
Robin Rombach Feb 17, 2024 ▶ 34:34
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.