Everything Robin Rombach said on any show that made the record, most notable first. Each card names its show and opens the statement there.
Rombach: Pre-training AI models on video gives implicit understanding of physics
“Pre-training on videos gives, like, implicit understanding of the physics of interactions with the real world, and then you can get stuff like action prediction, like robotics out of the same model.”
Rombach: Same generative AI models making movies can power robot brains
“I think what really excites me is you can use the same kind of AI model to make a movie and deploy that as a brain on a robot.”
Rombach: Compute constraints drive AI innovation more than scaling hardware
“If you only rely on like more compute it's a bit boring. I think like compute constraints can also Drive innovation, right? So for example, the latent diffusion framework, we developed it at the university because we just, like, we had, like, single GPUs where…”
Robin Rombach: Most Valuable AI Video Uses Require Human-in-the-Loop
“But I think ultimately. Like the real interesting use cases, they come when you have like a human in the loop who iterates and uses it as a medium.”
Rombach: Black Forest Labs builds custom AI models with major IP owners
“We do work with certain IP holders to Develop models together with them. Some of them based on our open source models, some of them based on like our more powerful proprietary models.”
Rombach: Video AI models scaled like LLMs will gain world understanding
“Having something like we are seeing in language modeling, but trained on pixels on videos will probably give like super interesting downstream behavior to not, not only like generating videos, but also understanding of the world.”
Rombach: Pre-trained video models learn 3D synthesis faster than image models
“We showed that it's actually like helpful to incorporate like this implicit three D knowledge that Knowledge that is captured in all of the videos into the model, and then the model can learn much quicker than if you start from the pure image model.”
Rombach: AI video generation will evolve into real-time interactive experiences
“Because then this will become more like, I don't know, sometimes I think about this as like a video game, right? You type your prompt, and you immediately see what happens given your input view, and I think this might be a super nice user experience, actually.”
Rombach: DALL-E 3 uses an autoencoder trained on a single GPU
“Dolly three uses a model that, like, the autoencoder that was trained on a single GPU.”
Rombach: Black Forest Labs builds multimodal models for physical AI
“We are building multimodal visual models for content creation and for now physical AI.”
Rombach: Adapting visual AI for robotics requires mere hours of data
“In practice what you do is, You have like all this like visual understanding in the models. And then you need only a very little bit of like a few hours of fine tuning data to adjust the model on that specific task.”
Rombach: Stability AI distillation research enables single-step diffusion generation
“We ourselves, we have published a distillation work a week ago that actually shows that you can go as low as one sampling step, which is, I would say like a big advantage of these diffusion models.”
Rombach: Managing hundreds of LoRAs is unscalable for video control
“Maintaining like a library of hundreds of LoRa's is maybe not like the, Most scalable approach.”
Rombach: Stable Video Diffusion took six months to develop
“I would say roughly half a year and like for this model that we just put out, I think the main challenge was that we actually, Yeah, I had to scale the data set and the data loading.”
Rombach: Stable Diffusion's core technique was developed during university research
“It's based on a technique that we developed while we were still at the university.”