video models

also referred to as: video model

12 statements across 5 episodes · 6 bullish · 1 bearish · 5 people on the record · first statement Sep 8, 2025 by Gorkem Yurtseven · across every show →

Everything said about video models, oldest first

Sep 8, 2025 bullish
Disclosure
Yurtseven: Video models now account for over 50% of Fal's revenue
“That was February, so now, now it's probably over 80. No, 50%. 50? Yeah, okay. It's like over 50. Yeah, yeah, a hundred percent.”
Gorkem Yurtseven Sep 8, 2025 ▶ 35:27 A Technical History of Generative Media
Dec 6, 2025 positive
Insight
World models are significantly more complex than traditional generative video models
“What world models do is they actually have to understand the full range of possibilities and outcomes from the current states and based on the action that you take generates the next states, right? So the next frame. And so it is a much more sort of complex pr…”
Pim de Witte Dec 6, 2025 ▶ 41:14 World Models & General Intuition: Khosla's largest bet since LLMs & OpenAI
Jun 1, 2026 neutral
Insight
Ethan He: Training video models costs roughly the same as medium-scale LLMs
“So surprisingly video models is like the cost is very, is comparable to language models. And obviously the largest scale is language model. Maybe like a medium scale language models.”
Ethan He Jun 1, 2026 ▶ 34:15 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 1, 2026
Insight
Ethan He: Video Models Must Bootstrap From Image Diffusion Models for Semantic Understanding
“After you train such model, such image model, the reason it's a foundation for video models is that image, image models are Cheaper to train and they have much denser connection between language and text. So, sorry, language and images. For example, you train …”
Ethan He Jun 1, 2026 ▶ 18:54 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 1, 2026 bullish
Prediction Not checkable as stated
Ethan He: RLMs and video models will dynamically pull context like humans
“But humans' contacts can, like, attention can work because we can dynamically pull in contacts from different places. The same mechanism I think it's going to happen for RLMs and video models.”
Ethan He Jun 1, 2026 ▶ 1:05:58 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 1, 2026 neutral
Insight
He: Manual reference video conditioning is a workaround, not true long context
“It doesn't need to have a very long context, but it's, I feel like it's an intermediate solution. It's cheating. Yeah, the model should Be able to like selectively know, like where, where should I draw references?”
Ethan He Jun 1, 2026 ▶ 1:00:40 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 1, 2026
Insight
Ethan He: Video models require image foundations and 100% synthetic caption pairs
“Building a video model. You actually need to build a image model first and building, building these two models. The data you need is a hundred percent synthetic pair of language and image or language to video because on the internet, actually the videos Don't …”
Ethan He Jun 1, 2026 ▶ 11:55 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 1, 2026
Insight
He: Video Agents Are Inherently Costlier Due to Iterative Multi-Sample Generation
“I think the enterprise will have much more budget for video models because the agents are inherently more expensive than the other video models themselves because they do this iterative process. They generate many, many variations.”
Ethan He Jun 1, 2026 ▶ 1:31:20 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 1, 2026 positive
Opinion
Ethan He: Long context management in video models leads LLM context work
“I feel this is actually, this part of long contacts is a little bit ahead of the LLM part.”
Ethan He Jun 1, 2026 ▶ 1:03:42 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Aug 3, 2026 positive
Opinion
Open-source LLMs have reached closed parity, but video models lag far behind
“The difference between the best open source LLM and the best open closed source LLM is very small. Like it used to be six months. I don't think it's six months anymore. I think it's like basically almost unparative. Video models are definitely not, there's a h…”
Ali Taha Aug 3, 2026 ▶ 1:14:57 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Aug 3, 2026 bullish
Prediction Not checkable as stated
Long-form AI video generation must switch to autoregressive architectures over diffusion
“Autoregressive video seems to me like that is the bet that the future is going to be making. But there are no good open source autoregressive video models out there today. And that seems to be, if you want to get like an hour movie, if you want to see video mo…”
Ali Taha Aug 3, 2026 ▶ 1:18:07 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Sep 4, 2026 negative
Assertion Contradicted
Anandkumar: Existing video and vision world models incorrectly assume fixed resolutions
“That immediately distinguishes us from other so-called world models, whether it's video models, vision models, they all assume during training and inference, it's a fixed resolution.”
Anima Anandkumar Sep 4, 2026 ▶ 10:15 Faster Chips That Don't Melt — Anima Anandkumar & Benedikt Jenik, Accelerated Understanding
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.