Video Models
topic on 3 shows · 19 statements across 10 episodes
Latent Space
No Priors
the a16z Podcast
19 statements about Video Models, every show
Anandkumar: Existing video and vision world models incorrectly assume fixed resolutions
“That immediately distinguishes us from other so-called world models, whether it's video models, vision models, they all assume during training and inference, it's a fixed resolution.”
Open-source LLMs have reached closed parity, but video models lag far behind
“The difference between the best open source LLM and the best open closed source LLM is very small. Like it used to be six months. I don't think it's six months anymore. I think it's like basically almost unparative. Video models are definitely not, there's a h…”
Long-form AI video generation must switch to autoregressive architectures over diffusion
“Autoregressive video seems to me like that is the bet that the future is going to be making. But there are no good open source autoregressive video models out there today. And that seems to be, if you want to get like an hour movie, if you want to see video mo…”
Ethan He: Video models require image foundations and 100% synthetic caption pairs
“Building a video model. You actually need to build a image model first and building, building these two models. The data you need is a hundred percent synthetic pair of language and image or language to video because on the internet, actually the videos Don't …”
Ethan He: Video Models Must Bootstrap From Image Diffusion Models for Semantic Understanding
“After you train such model, such image model, the reason it's a foundation for video models is that image, image models are Cheaper to train and they have much denser connection between language and text. So, sorry, language and images. For example, you train …”
Ethan He: Training video models costs roughly the same as medium-scale LLMs
“So surprisingly video models is like the cost is very, is comparable to language models. And obviously the largest scale is language model. Maybe like a medium scale language models.”
He: Manual reference video conditioning is a workaround, not true long context
“It doesn't need to have a very long context, but it's, I feel like it's an intermediate solution. It's cheating. Yeah, the model should Be able to like selectively know, like where, where should I draw references?”
Ethan He: Long context management in video models leads LLM context work
“I feel this is actually, this part of long contacts is a little bit ahead of the LLM part.”
Ethan He: RLMs and video models will dynamically pull context like humans
“But humans' contacts can, like, attention can work because we can dynamically pull in contacts from different places. The same mechanism I think it's going to happen for RLMs and video models.”
He: Video Agents Are Inherently Costlier Due to Iterative Multi-Sample Generation
“I think the enterprise will have much more budget for video models because the agents are inherently more expensive than the other video models themselves because they do this iterative process. They generate many, many variations.”
World models are significantly more complex than traditional generative video models
“What world models do is they actually have to understand the full range of possibilities and outcomes from the current states and based on the action that you take generates the next states, right? So the next frame. And so it is a much more sort of complex pr…”
Wang: AI-generated videos yield highly accurate 3D reconstructions
“Video models have very good three-D understanding. You can run reconstruction algorithms over the videos you generate, and they're very accurate.”
Altman: Video AI models capable of deepfaking anyone will arrive soon
“So like, very soon, the world is going to have to contend with incredible video models that can deepfake anyone or kind of show anything you want, and that will mostly be great.”
Altman: Rights holders reacted differently to video AI than image AI
“And we saw an example of a different, like video models got a very different response from rights holders than ImageGen does.”
Yurtseven: Video models now account for over 50% of Fal's revenue
“That was February, so now, now it's probably over 80. No, 50%. 50? Yeah, okay. It's like over 50. Yeah, yeah, a hundred percent.”
Lingelbach: LLMs, not video models, limit real-time AI actors
“Oh, a hundred percent. But I think that the bigger limitation is not the video model right now. I think the bigger limitation is actually that I think large language models still have a lot of work in terms of making people feel very authentic.”
Song: Video Models Unexpectedly Learn 3D Spatial Reasoning From Raw Video
“It turns out from these type of videos, we are able to show that a video model is able to reason about three D quite well, which in some sense is unexpected before in, in the community.”
Song: Dream Machine Spontaneously Generates Consistent Cinematic Cuts Without Prompting
“Even though we haven't really asked the model to do anything about it, the model is able to reason about the second shot being a cut of the first shot that happens to have the same things in the first shot. So this is curse definitely some kind of a non-trivia…”
Doshi: Video AI models are pre-trained on a billion images first
“And then the other thing with video is like videos is just extraordinarily computationally expensive to do inference or even training on. And a lot of the video models first train, like pre-train with like a billion images first anyway, to like have a rich, Se…”