The Ledger, every show

Every statement that passed quotation and attribution checks, across all 44 shows. Pick shows below, then mix any filter with any other.

shows every show 44 of 44
every show
clear all ✕
LATENT SPACE Prediction Not checkable as stated
Ethan He: Falling inference costs will enable generative UIs for everything
“So I think as a inference cost come down, we are going to have generative UI for everything.”
Ethan He Jun 1, 2026 ▶ 25:46 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
LATENT SPACE Prediction Not checkable as stated
Ethan He: LLM Video Agents Will Orchestrate Diffusion Models and Editing Tools
“Video agents, mostly language models, they'll call these generative model, either it's a separate model or a diffusion head or whatever as tool. So this model can iteratively Refine the results or even like you generate longer content through a very long trend…”
Ethan He Jun 1, 2026 ▶ 1:21:56 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
LATENT SPACE Prediction Held up
Ethan He: Video Agents Will Reach Production-Grade Quality by Year-End
“I guess by the end of this year is this is going to be a big hit. So the inflection point will be there and the videos generated by video agents can get to like production great quality. So it can be presented and it can be distributed in, in ads.”
Ethan He Jun 1, 2026 ▶ 1:30:54 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
LATENT SPACE Assertion Not checkable as stated
Ethan He: Small xAI Team Built Grok Imagine in Three Months
“There were no, no infra, no data, and no model. And it just a few engineers, we built it in three months and released the first model, Grok Imagine,”
Ethan He Jun 1, 2026 ▶ 3:59 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
LATENT SPACE Prediction Not checkable as stated
Ethan He: Neural OS models can synthesize novel user interfaces
“So if you train your neural OS or neural computer on the standard screen recordings on the entire internet, the model can imagine completely new interface to interact with the computer.”
Ethan He Jun 1, 2026 ▶ 31:45 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
LATENT SPACE Assertion Supported
Ethan He: Storing and moving video datasets costs millions per month
“So, so it's like just storing, storing the network, those costs, it's just I guess it would be a few millions per month to just storing everything, not to mention the GPU costs.”
Ethan He Jun 1, 2026 ▶ 35:49 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
LATENT SPACE Assertion Open · timeframe Jun 2029
Ethan He: Grok Imagine Video Extension Tracks Full Historical Context
“So the Glock Imagine video extension, it has historical context of all of the previous generated videos. It can it has a context of who is speaking and what objects have appeared and everything having that to generate the next video.”
Ethan He Jun 1, 2026 ▶ 55:32 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
LATENT SPACE Prediction Not checkable as stated
Ethan He: Powerful video AI will naturally learn to control physical robots
“Once these models can use computers and understand the future state of computer extremely well, the robots might be Might be one of the tools a very powerful AI can use. So the powerful AI might just be able to control the physical embodiment naturally.”
Ethan He Jun 1, 2026 ▶ 1:33:20 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
LATENT SPACE Assertion Contradicted
He: Grok Imagine 0.9 was first large-scale joint audio-video model deployed
“So Grok Imagine, there were .9, I believe it's is a first first audio video trends model deployed at a large scale.”
Ethan He Jun 1, 2026 ▶ 42:45 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
LATENT SPACE Prediction Not checkable as stated
Ethan He predicts world models will culminate in real-time neural computers
“I think the final state will be, for example, like a video version of Playbook where you can interact with a neural computer. You move your mouse and you click on the generative interface. And it will reply to you through, through pixels generally in real time…”
Ethan He Jun 1, 2026 ▶ 53:06 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
LATENT SPACE Prediction Not checkable as stated
Ethan He: RLMs and video models will dynamically pull context like humans
“But humans' contacts can, like, attention can work because we can dynamically pull in contacts from different places. The same mechanism I think it's going to happen for RLMs and video models.”
Ethan He Jun 1, 2026 ▶ 1:05:58 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
LATENT SPACE Assertion Not checkable as stated
He: Elon Musk is very hands-on and works closely with xAI teams
“He also worked very closely with people like people imagine online, like he, he's very hands-on.”
Ethan He Jun 1, 2026 ▶ 1:09:40 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
LATENT SPACE Prediction Not checkable as stated
Ethan He: AI watermarking will remain vulnerable to reverse-engineering
“As a limitation is like the technology is, as a paper, Was out there and people can reverse engineer that how to get rid of it. And it's, I think even as it advance, it's still, still possible to reverse engineer it.”
Ethan He Jun 1, 2026 ▶ 1:12:11 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
LATENT SPACE Prediction Not checkable as stated
Ethan He: Video Agents Will Transition to Fully Automated Video Production
“So in, in Asian, in Gorky Imagine agent mode, you can still go in there and do, do stuff by yourself. Gradually, as the model capability increase, it will be able to do everything fully automated.”
Ethan He Jun 1, 2026 ▶ 1:25:30 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
LATENT SPACE Prediction Not checkable as stated
Ethan He: LLMs will soon become context-aware and manage context
“I think one thing pretty, pretty interesting. I think might be happening soon is the language models will be like context aware and manage its own context.”
Ethan He Jun 1, 2026 ▶ 1:35:33 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
LATENT SPACE Assertion Supported
Ethan He: Megatron MoE was first to train trillion-parameter MoEs at 40% MFU
“The Megatron MOEs was the first It was the first framework open source to be able to train these MOEs at very large scales, like a hundred billion parameters to even trillion parameters efficiently at like 40% MFU.”
Ethan He Jun 1, 2026 ▶ 1:41:50 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
LATENT SPACE Assertion Supported
Ethan He: Hugging Face's sequential GEMM loop for Mixtral is inefficient
“Let's also look at the implementation of Mixtro eight by seven on Hagen-Phys transformer. You will soon notice the, in the expert operation there, You would iterate over all of the experts and compute each of the gem operations one by one. We found that this i…”
Ethan He Oct 29, 2024 ▶ 12:37 [Paper Club] Upcycling Large Language Models into Mixture of Experts
LATENT SPACE Assertion Supported
Ethan He: MoE experts do not cleanly specialize into semantic domains
“Unfortunately, people didn't find, like, a significant interpretability inside these experts. Say, one expert focus on math, the other focus on literature. I think the problem is that neural network hidden states are already very entangled. So, when hidden sta…”
Ethan He Oct 29, 2024 ▶ 16:19 [Paper Club] Upcycling Large Language Models into Mixture of Experts
LATENT SPACE Assertion Supported
He: Upcycling a 15B model on 1T tokens yielded 4% MMLU gain
“On other scaling experiments, we tried on 15 B models upcycling and applied on one trillion tokens and achieved roughly about five percent improvement in terms of the validation loss and four percent improvement on MMLU.”
Ethan He Oct 29, 2024 ▶ 21:12 [Paper Club] Upcycling Large Language Models into Mixture of Experts
LATENT SPACE Assertion Not checkable as stated
Ethan He: NVIDIA spent about a year building the Cosmos model
“One thing I say, like, thanks to my experience at NVIDIA, because first time when we were building Cosmos together, we built it for about a year.”
Ethan He Jun 1, 2026 ▶ 5:13 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
LATENT SPACE Assertion Supported
NVIDIA Cosmos uses 50,000 to 60,000 tokens for five seconds of video
“Yeah, for example, like in Cosmos, I think just five seconds of video is like a 50, 50 K or a 60 K number of tokens. So like, if you do 50 seconds as a 500 K tokens, if you do longer than that, easily explode.”
Ethan He Jun 1, 2026 ▶ 56:16 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.