Former xAI and NVIDIA researcher Ethan He discusses the architectural trade-offs between frame-by-frame spatial compression and temporal compression in video models.
“The difference is if you compress the temporal dimension, you get a much higher compression rate. Because there is temporal redundancy between frames, because this frame and the last frame, likely they are mostly similar. So there's only some small difference. For example, like I think in one, 2.1, they have like an eight by eight by four compression rate. So the four temporal tokens are compressed into one tokens that can save a lot of Save, save a lot of the context lens. If you do it frame by frame, you have to do maybe like eight by eight by one. Your context lens will be four times larger.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Ethan He
Insight
Ethan He: Pipeline bug fixes drive more model gains than new algorithms
“And often I find that this is kind of boring, but like a lot of the improvements does not come from new algorithms. It comes from finding small bugs here and there in the data pipeline, in the model training pipeline. Those gave the biggest boost to the model …”
Ethan HeJun 1, 2026▶ 7:40Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
PredictionNot checkable as stated
Ethan He: Falling inference costs will enable generative UIs for everything
“So I think as a inference cost come down, we are going to have generative UI for everything.”
Ethan HeJun 1, 2026▶ 25:46Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Insight
Ethan He: Training video models costs roughly the same as medium-scale LLMs
“So surprisingly video models is like the cost is very, is comparable to language models. And obviously the largest scale is language model. Maybe like a medium scale language models.”
Ethan HeJun 1, 2026▶ 34:15Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Insight
Ethan He: Visual intelligence in video generation models stems primarily from language models
“The visual intelligence are actually mostly coming from language. Like, these video models, especially from now, since the diffusion model technology is more mature, the, like, every time you see there, there's some improvement on these models, I would say mos…”
Ethan HeJun 1, 2026▶ 1:14:55Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
PredictionNot checkable as stated
Ethan He: LLM Video Agents Will Orchestrate Diffusion Models and Editing Tools
“Video agents, mostly language models, they'll call these generative model, either it's a separate model or a diffusion head or whatever as tool. So this model can iteratively Refine the results or even like you generate longer content through a very long trend…”
Ethan HeJun 1, 2026▶ 1:21:56Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
PredictionHeld up
Ethan He: Video Agents Will Reach Production-Grade Quality by Year-End
“I guess by the end of this year is this is going to be a big hit. So the inflection point will be there and the videos generated by video agents can get to like production great quality. So it can be presented and it can be distributed in, in ads.”
Ethan HeJun 1, 2026▶ 1:30:54Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 200 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.