Assertion Supported AI assessment confidence: 88% certainty 4/5 debate potential 2/5

Ethan He: Megatron MoE was first to train trillion-parameter MoEs at 40% MFU

Ethan He · Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He · Jun 1, 2026 · at 1:41:50

Ethan He, former researcher at NVIDIA and xAI, reflects on his work developing open-source training infrastructure for large-scale Mixture-of-Experts models at NVIDIA.

0:00 / 0:17exact quote · 17.6s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“The Megatron MOEs was the first It was the first framework open source to be able to train these MOEs at very large scales, like a hundred billion parameters to even trillion parameters efficiently at like 40% MFU.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Ethan He

Insight
Ethan He: Pipeline bug fixes drive more model gains than new algorithms
“And often I find that this is kind of boring, but like a lot of the improvements does not come from new algorithms. It comes from finding small bugs here and there in the data pipeline, in the model training pipeline. Those gave the biggest boost to the model …”
Ethan He Jun 1, 2026 ▶ 7:40 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Prediction Not checkable as stated
Ethan He: Falling inference costs will enable generative UIs for everything
“So I think as a inference cost come down, we are going to have generative UI for everything.”
Ethan He Jun 1, 2026 ▶ 25:46 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Insight
Ethan He: Training video models costs roughly the same as medium-scale LLMs
“So surprisingly video models is like the cost is very, is comparable to language models. And obviously the largest scale is language model. Maybe like a medium scale language models.”
Ethan He Jun 1, 2026 ▶ 34:15 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Insight
Ethan He: Visual intelligence in video generation models stems primarily from language models
“The visual intelligence are actually mostly coming from language. Like, these video models, especially from now, since the diffusion model technology is more mature, the, like, every time you see there, there's some improvement on these models, I would say mos…”
Ethan He Jun 1, 2026 ▶ 1:14:55 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Prediction Not checkable as stated
Ethan He: LLM Video Agents Will Orchestrate Diffusion Models and Editing Tools
“Video agents, mostly language models, they'll call these generative model, either it's a separate model or a diffusion head or whatever as tool. So this model can iteratively Refine the results or even like you generate longer content through a very long trend…”
Ethan He Jun 1, 2026 ▶ 1:21:56 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Prediction Held up
Ethan He: Video Agents Will Reach Production-Grade Quality by Year-End
“I guess by the end of this year is this is going to be a big hit. So the inflection point will be there and the videos generated by video agents can get to like production great quality. So it can be presented and it can be distributed in, in ads.”
Ethan He Jun 1, 2026 ▶ 1:30:54 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.