Robin Rombach, lead AI researcher at Stability AI, discusses distillation techniques that dramatically reduce the number of sampling steps required by diffusion models.
“We ourselves, we have published a distillation work a week ago that actually shows that you can go as low as one sampling step, which is, I would say like a big advantage of these diffusion models.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Robin Rombach
Insight
Rombach: Compute constraints drive AI innovation more than scaling hardware
“If you only rely on like more compute it's a bit boring. I think like compute constraints can also Drive innovation, right? So for example, the latent diffusion framework, we developed it at the university because we just, like, we had, like, single GPUs where…”
Robin RombachFeb 17, 2024▶ 34:02Text to Video: The Next Leap in AI Generation
PredictionNot checkable as stated
Rombach: Video AI models scaled like LLMs will gain world understanding
“Having something like we are seeing in language modeling, but trained on pixels on videos will probably give like super interesting downstream behavior to not, not only like generating videos, but also understanding of the world.”
Robin RombachFeb 17, 2024▶ 14:05Text to Video: The Next Leap in AI Generation
AssertionSupported
Rombach: Pre-trained video models learn 3D synthesis faster than image models
“We showed that it's actually like helpful to incorporate like this implicit three D knowledge that Knowledge that is captured in all of the videos into the model, and then the model can learn much quicker than if you start from the pure image model.”
Robin RombachFeb 17, 2024▶ 23:07Text to Video: The Next Leap in AI Generation
PredictionNot checkable as stated
Rombach: AI video generation will evolve into real-time interactive experiences
“Because then this will become more like, I don't know, sometimes I think about this as like a video game, right? You type your prompt, and you immediately see what happens given your input view, and I think this might be a super nice user experience, actually.”
Robin RombachFeb 17, 2024▶ 28:23Text to Video: The Next Leap in AI Generation
AssertionSupported
Rombach: DALL-E 3 uses an autoencoder trained on a single GPU
“Dolly three uses a model that, like, the autoencoder that was trained on a single GPU.”
Robin RombachFeb 17, 2024▶ 34:34Text to Video: The Next Leap in AI Generation
Opinion
Rombach: Managing hundreds of LoRAs is unscalable for video control
“Maintaining like a library of hundreds of LoRa's is maybe not like the, Most scalable approach.”
Robin RombachFeb 17, 2024▶ 27:09Text to Video: The Next Leap in AI Generation
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 1,000 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.