Robin Rombach, co-creator of Stable Diffusion and researcher at Stability AI, discusses the training efficiency of Stable Video Diffusion when adapted for 3D multi-view object synthesis.
Insight
Rombach: Compute constraints drive AI innovation more than scaling hardware
“If you only rely on like more compute it's a bit boring. I think like compute constraints can also Drive innovation, right? So for example, the latent diffusion framework, we developed it at the university because we just, like, we had, like, single GPUs where…”
Prediction Not checkable as stated
Rombach: Video AI models scaled like LLMs will gain world understanding
“Having something like we are seeing in language modeling, but trained on pixels on videos will probably give like super interesting downstream behavior to not, not only like generating videos, but also understanding of the world.”
Prediction Not checkable as stated
Rombach: AI video generation will evolve into real-time interactive experiences
“Because then this will become more like, I don't know, sometimes I think about this as like a video game, right? You type your prompt, and you immediately see what happens given your input view, and I think this might be a super nice user experience, actually.”
Assertion Supported
Rombach: DALL-E 3 uses an autoencoder trained on a single GPU
“Dolly three uses a model that, like, the autoencoder that was trained on a single GPU.”
Disclosure
Rombach: Stability AI distillation research enables single-step diffusion generation
“We ourselves, we have published a distillation work a week ago that actually shows that you can go as low as one sampling step, which is, I would say like a big advantage of these diffusion models.”
Opinion
Rombach: Managing hundreds of LoRAs is unscalable for video control
“Maintaining like a library of hundreds of LoRa's is maybe not like the, Most scalable approach.”