Ethan He

16 statements across 1 episodes · 7 bullish · 0 bearish · 1 people on the record · first statement Jun 1, 2026 by Ethan He · said 2 times in 1 episodes since 2026 · across every show →

On the record as a speaker too: Ethan He's record, appearances and statements → this page counts the times other people say the name.

Mentions by year

brought up most by Shawn Wang (2)

tap a year for its mentions
0011212026episodesmentions
0112026episodes it came up in
0010.5212026episodesmentions per episode

every mention, scene by scene, with the transcript →

Everything said about Ethan He, oldest first

Jun 1, 2026 positive
Insight
He: AI first-principles planning calculates the theoretical minimum days to ship
“If you think about some limitation, for example, the current data, like how, how fast can we acquire the videos? And if you think about training the models, like what's the iteration speed? For training a model end-to-end and how, how would adding more GPUs ac…”
Ethan He Jun 1, 2026 ▶ 1:08:22 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 1, 2026 bullish
Prediction Not checkable as stated
Ethan He: Falling inference costs will enable generative UIs for everything
“So I think as a inference cost come down, we are going to have generative UI for everything.”
Ethan He Jun 1, 2026 ▶ 25:46 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 1, 2026 positive
Insight
Ethan He: Pipeline bug fixes drive more model gains than new algorithms
“And often I find that this is kind of boring, but like a lot of the improvements does not come from new algorithms. It comes from finding small bugs here and there in the data pipeline, in the model training pipeline. Those gave the biggest boost to the model …”
Ethan He Jun 1, 2026 ▶ 7:40 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 1, 2026
Insight
Ethan He: Video Models Must Bootstrap From Image Diffusion Models for Semantic Understanding
“After you train such model, such image model, the reason it's a foundation for video models is that image, image models are Cheaper to train and they have much denser connection between language and text. So, sorry, language and images. For example, you train …”
Ethan He Jun 1, 2026 ▶ 18:54 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 1, 2026 neutral
Disclosure
Ethan He left xAI because changing corporate priorities limited LLM research
“For me there's a lot of research you want to do that you cannot do at, as a company. And also like the priorities and objective, the, for company typically can change very fast. It is, it's also the same for XAI. So, so now it's kind of like the time to, there…”
Ethan He Jun 1, 2026 ▶ 1:33:55 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 1, 2026 positive
Assertion Not checkable as stated
He: Elon Musk is very hands-on and works closely with xAI teams
“He also worked very closely with people like people imagine online, like he, he's very hands-on.”
Ethan He Jun 1, 2026 ▶ 1:09:40 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 1, 2026 neutral
Insight
Ethan He: Training transformers directly on raw image pixels is impossible
“If you're trying, if you can technically, theoretically train image or video models on pure pixels, but the problem is that the, it's a lot of tokens. So like one image, like it's a thousand by a thousand is like one million tokens, one million pixels. It's im…”
Ethan He Jun 1, 2026 ▶ 15:30 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 1, 2026 neutral
Insight
He: Manual reference video conditioning is a workaround, not true long context
“It doesn't need to have a very long context, but it's, I feel like it's an intermediate solution. It's cheating. Yeah, the model should Be able to like selectively know, like where, where should I draw references?”
Ethan He Jun 1, 2026 ▶ 1:00:40 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 1, 2026 neutral
Insight
Ethan He: Diffusion transformer training closely mirrors LLM training architecture
“So now the training, training of the diffusion transformer, you already generated models use diffusion transformers. It is actually quite standard. It's very similar to how you train a language transformer models. It's not that much difference. It's just the t…”
Ethan He Jun 1, 2026 ▶ 17:40 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 1, 2026
Insight
Ethan He: Video models require image foundations and 100% synthetic caption pairs
“Building a video model. You actually need to build a image model first and building, building these two models. The data you need is a hundred percent synthetic pair of language and image or language to video because on the internet, actually the videos Don't …”
Ethan He Jun 1, 2026 ▶ 11:55 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 1, 2026
Disclosure
Ethan He: NVIDIA Cosmos required labelers to describe videos for blind reconstruction
“So that's in the protocol of Cosmos labeling. We required the objective we gave to the labelers was that you have to describe the video as detailed as possible, such that a blind person hears a blob of text, can reconstruct what the video is like from their he…”
Ethan He Jun 1, 2026 ▶ 13:39 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 1, 2026
Insight
Ethan He: Coding Models Shift Research Bottlenecks Back to Compute
“Compute might become a bottleneck again, because previously, like if you want to train a new model, say you want to generate new synthetic data and then, or write a new algorithm, it might take a few weeks. And during that period of time, you don't, you might …”
Ethan He Jun 1, 2026 ▶ 9:14 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 1, 2026 neutral
Assertion Not checkable as stated
Ethan He: NVIDIA spent about a year building the Cosmos model
“One thing I say, like, thanks to my experience at NVIDIA, because first time when we were building Cosmos together, we built it for about a year.”
Ethan He Jun 1, 2026 ▶ 5:13 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 1, 2026 positive
Prediction Not checkable as stated
Ethan He: Neural OS models can synthesize novel user interfaces
“So if you train your neural OS or neural computer on the standard screen recordings on the entire internet, the model can imagine completely new interface to interact with the computer.”
Ethan He Jun 1, 2026 ▶ 31:45 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 1, 2026 positive
Prediction Not checkable as stated
Ethan He: Powerful video AI will naturally learn to control physical robots
“Once these models can use computers and understand the future state of computer extremely well, the robots might be Might be one of the tools a very powerful AI can use. So the powerful AI might just be able to control the physical embodiment naturally.”
Ethan He Jun 1, 2026 ▶ 1:33:20 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 1, 2026 positive
Insight
Ethan He: Daily iteration speed is the top factor in model training
“When I look at like training models, I don't so actually the top important thing is like how many how many iterations can you do like per, per day? And the more iteration can you do, you can train the model much faster. So if you have a very strong infra and y…”
Ethan He Jun 1, 2026 ▶ 6:21 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.