Ethan He

AI Researcher & Engineer · 2 appearances on the record.

computed by AI from the episodes · how this works → · full disclaimer →

scientistengineeracademic@EthanHe_42 ↗LinkedIn ↗yihui-he.github.io ↗

Ethan He previously worked on NVIDIA's Cosmos world model platform and Megatron-LM MoE architectures before joining xAI to help build Grok Imagine. He has authored research on model compression and currently focuses on LLM context architecture and video agents.

56statements → 21claims → 8claims resolved → 88%fully supported → 3.64/5average certainty → 1.98/5average debate potential → 2said about them ↓

7 supported 0 partly supported 1 contradicted 1 not yet assessed 12 not checkable as stated how the 21 claims stand · each chip opens the sources

10 predictions · 11 assertions · 1 opinion · 28 insights · 6 disclosures · every statement was checked. The predictions and assertions are the 21 claims: statements the public record can support or contradict. 8 are resolved, 1 is not yet assessed, and 12 name no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Ethan argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Prediction Held up
Ethan He: Video Agents Will Reach Production-Grade Quality by Year-End
“I guess by the end of this year is this is going to be a big hit. So the inflection point will be there and the videos generated by video agents can get to like production great quality. So it can be presented and it can be distributed in, in ads.”
Ethan He Jun 1, 2026 ▶ 1:30:54 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He

Their most notable contradicted claim

Assertion Contradicted
He: Grok Imagine 0.9 was first large-scale joint audio-video model deployed
“So Grok Imagine, there were .9, I believe it's is a first first audio video trends model deployed at a large scale.”
Ethan He Jun 1, 2026 ▶ 42:45 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He

Expressed certainty vs assessment result

none yet certainty 1
none yet certainty 2
75% certainty 3
100% certainty 4
none yet certainty 5

weighted support: a fully supported claim counts one, a partly supported claim counts half. Each filled bar is clickable and opens exactly those claims; "none yet" means nothing said at that certainty level has resolved yet

How they sound: not measured why? →

We measure speaking style by listening to the audio itself, and a fair number needs at least 2,000 words from one person on tape we have measured. There is too little of Ethan He on measured tape to publish a rate. This says nothing about how they speak.

Everything Ethan He said on Latent Space that made the record, most notable first. Filter by type, assessment or year in the ledger →

Insight
Ethan He: Pipeline bug fixes drive more model gains than new algorithms
“And often I find that this is kind of boring, but like a lot of the improvements does not come from new algorithms. It comes from finding small bugs here and there in the data pipeline, in the model training pipeline. Those gave the biggest boost to the model …”
Ethan He Jun 1, 2026 ▶ 7:40 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Prediction Not checkable as stated
Ethan He: Falling inference costs will enable generative UIs for everything
“So I think as a inference cost come down, we are going to have generative UI for everything.”
Ethan He Jun 1, 2026 ▶ 25:46 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Insight
Ethan He: Training video models costs roughly the same as medium-scale LLMs
“So surprisingly video models is like the cost is very, is comparable to language models. And obviously the largest scale is language model. Maybe like a medium scale language models.”
Ethan He Jun 1, 2026 ▶ 34:15 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Insight
Ethan He: Visual intelligence in video generation models stems primarily from language models
“The visual intelligence are actually mostly coming from language. Like, these video models, especially from now, since the diffusion model technology is more mature, the, like, every time you see there, there's some improvement on these models, I would say mos…”
Ethan He Jun 1, 2026 ▶ 1:14:55 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Prediction Not checkable as stated
Ethan He: LLM Video Agents Will Orchestrate Diffusion Models and Editing Tools
“Video agents, mostly language models, they'll call these generative model, either it's a separate model or a diffusion head or whatever as tool. So this model can iteratively Refine the results or even like you generate longer content through a very long trend…”
Ethan He Jun 1, 2026 ▶ 1:21:56 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Prediction Held up
Ethan He: Video Agents Will Reach Production-Grade Quality by Year-End
“I guess by the end of this year is this is going to be a big hit. So the inflection point will be there and the videos generated by video agents can get to like production great quality. So it can be presented and it can be distributed in, in ads.”
Ethan He Jun 1, 2026 ▶ 1:30:54 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Insight
Ethan He: Peak pre-training learning rate works best for MoE upcycling
“We found that the best is to use the original highest peak learning rate from pre-training, which works the best.”
Ethan He Oct 29, 2024 ▶ 33:34 [Paper Club] Upcycling Large Language Models into Mixture of Experts
Assertion Not checkable as stated
Ethan He: Small xAI Team Built Grok Imagine in Three Months
“There were no, no infra, no data, and no model. And it just a few engineers, we built it in three months and released the first model, Grok Imagine,”
Ethan He Jun 1, 2026 ▶ 3:59 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Insight
Ethan He: Video models require image foundations and 100% synthetic caption pairs
“Building a video model. You actually need to build a image model first and building, building these two models. The data you need is a hundred percent synthetic pair of language and image or language to video because on the internet, actually the videos Don't …”
Ethan He Jun 1, 2026 ▶ 11:55 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Prediction Not checkable as stated
Ethan He: Neural OS models can synthesize novel user interfaces
“So if you train your neural OS or neural computer on the standard screen recordings on the entire internet, the model can imagine completely new interface to interact with the computer.”
Ethan He Jun 1, 2026 ▶ 31:45 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Assertion Supported
Ethan He: Storing and moving video datasets costs millions per month
“So, so it's like just storing, storing the network, those costs, it's just I guess it would be a few millions per month to just storing everything, not to mention the GPU costs.”
Ethan He Jun 1, 2026 ▶ 35:49 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Assertion Open · timeframe Jun 2029
Ethan He: Grok Imagine Video Extension Tracks Full Historical Context
“So the Glock Imagine video extension, it has historical context of all of the previous generated videos. It can it has a context of who is speaking and what objects have appeared and everything having that to generate the next video.”
Ethan He Jun 1, 2026 ▶ 55:32 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Opinion
Ethan He: Long context management in video models leads LLM context work
“I feel this is actually, this part of long contacts is a little bit ahead of the LLM part.”
Ethan He Jun 1, 2026 ▶ 1:03:42 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Prediction Not checkable as stated
Ethan He: Powerful video AI will naturally learn to control physical robots
“Once these models can use computers and understand the future state of computer extremely well, the robots might be Might be one of the tools a very powerful AI can use. So the powerful AI might just be able to control the physical embodiment naturally.”
Ethan He Jun 1, 2026 ▶ 1:33:20 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Insight
He: Upcycling dense models to MoE beats continuing dense training per FLOP
“By training these upcycled models, you can achieve better accuracy than simply training the dense model further for the same number of flops.”
Ethan He Oct 29, 2024 ▶ 20:32 [Paper Club] Upcycling Large Language Models into Mixture of Experts
Insight
He: Mixtral's top-k before softmax routing hurts MoE upcycling performance
“We actually found the mix-throughs approach didn't work as well as expected, because the original model, the original switch transformer from Google uses a softmax and topk for a reason. And because of upcycling, if you switch to topk, then softmax, it actuall…”
Ethan He Oct 29, 2024 ▶ 24:34 [Paper Club] Upcycling Large Language Models into Mixture of Experts
Insight
Ethan He: Video Foundation Models Follow Scaling Laws Like LLMs
“There, once I built the Cosmos one, I realized as this thing also has a scaling law similar to language model.”
Ethan He Jun 1, 2026 ▶ 3:11 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Insight
Ethan He: Daily iteration speed is the top factor in model training
“When I look at like training models, I don't so actually the top important thing is like how many how many iterations can you do like per, per day? And the more iteration can you do, you can train the model much faster. So if you have a very strong infra and y…”
Ethan He Jun 1, 2026 ▶ 6:21 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Insight
Ethan He: Coding Models Shift Research Bottlenecks Back to Compute
“Compute might become a bottleneck again, because previously, like if you want to train a new model, say you want to generate new synthetic data and then, or write a new algorithm, it might take a few weeks. And during that period of time, you don't, you might …”
Ethan He Jun 1, 2026 ▶ 9:14 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Insight
Ethan He: Training generative video models on unlabeled data aids generalization
“For the generative model training, there's also really like a small percentage of unlabeled data. So, so the model is instructed to generate a video without any text instruction. That, that can also help the model generalize.”
Ethan He Jun 1, 2026 ▶ 15:00 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Insight
Ethan He: Training transformers directly on raw image pixels is impossible
“If you're trying, if you can technically, theoretically train image or video models on pure pixels, but the problem is that the, it's a lot of tokens. So like one image, like it's a thousand by a thousand is like one million tokens, one million pixels. It's im…”
Ethan He Jun 1, 2026 ▶ 15:30 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Insight
Ethan He: Diffusion transformer training closely mirrors LLM training architecture
“So now the training, training of the diffusion transformer, you already generated models use diffusion transformers. It is actually quite standard. It's very similar to how you train a language transformer models. It's not that much difference. It's just the t…”
Ethan He Jun 1, 2026 ▶ 17:40 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Insight
Ethan He: Video Models Must Bootstrap From Image Diffusion Models for Semantic Understanding
“After you train such model, such image model, the reason it's a foundation for video models is that image, image models are Cheaper to train and they have much denser connection between language and text. So, sorry, language and images. For example, you train …”
Ethan He Jun 1, 2026 ▶ 18:54 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Insight
Ethan He: Frame-by-frame video compression enables real-time interactivity, temporal compression adds lag
“That being said, the benefit of the frame per frame compression, we might come back to this later, is real timeliness and interactivity. Because if you strain the output of the model frame by frame, you can As a model can respond to any user request immediatel…”
Ethan He Jun 1, 2026 ▶ 22:54 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He

Show 24statements(32 left)

The other half of the tape: Ethan He's own voice is left out of every number here. Other people bring the name up 2 times in 1 episode on Latent Space. every mention, with the transcript →

Who brings them up most Shawn Wang 2

Every mention by year

tap a year for its mentions
0011212026episodesmentions
0112026episodes it came up in
0010.5212026episodesmentions per episode

Appearances (2)

EpisodeDateSpeaking time
Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Jun 1, 2026 1h 2m
[Paper Club] Upcycling Large Language Models into Mixture of Experts Oct 29, 2024 31m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.