Disclosure certainty 4/5 debate potential 1/5

Sun: Moonlake is training a unified multimodal latent representation

Fan-yun Sun · Moonlake: Interactive, Multimodal World Models — with Chris Manning and Fan-yun Sun · Apr 2, 2026 · at 56:28

Moonlake co-founder Fan-yun Sun explains the architectural objective of their world model regarding multimodal reasoning across vision and audio.

0:00 / 0:13exact quote · 13.4s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“We do want to basically like, we, our model model, like the one we're training is basically Towards the goal of having a combined latent representation across all these different modalities, right? Such that you can like reason across these different modalities.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Fan-yun Sun

Prediction Open · timeframe Apr 2031
Sun: Neural rendering with world priors will replace rasterizers and DLSS
“We actually believe that this is going to be the next paradigm of rendering. So it's going to replace how rasterizers, it's going to replace DLSS today because it not only has these pixel prior that's learned from the world, such that you can literally play an…”
Fan-yun Sun Apr 2, 2026 ▶ 30:36 Moonlake: Interactive, Multimodal World Models — with Chris Manning and Fan-yun Sun
Opinion
Sun: Pixel-coherent world simulators are overrated for causal reasoning and embodied AI
“Having a world simulator that can produce pixel coherency is very, very useful for games and, you know, marketing and all these things, but it's not as useful as people think when it comes to causal reasoning, when it comes to embodied AI.”
Fan-yun Sun Apr 2, 2026 ▶ 44:36 Moonlake: Interactive, Multimodal World Models — with Chris Manning and Fan-yun Sun
Assertion Supported
Sun: Synthetic data matches real-world data for multimodal model pre-training
“We were actually generating a lot of synthetic data and showing that, hey, you can actually, these synthetic data are actually as useful as real-world data when it comes to multimodal pre-training.”
Fan-yun Sun Apr 2, 2026 ▶ 2:56 Moonlake: Interactive, Multimodal World Models — with Chris Manning and Fan-yun Sun
Insight
Sun: Embodied General Intelligence Requires Interactive Causal Data
“On our way to, let's call it embodied general intelligence, Models need to learn the consequences behind their actions, which means that they need interactive data.”
Fan-yun Sun Apr 2, 2026 ▶ 3:18 Moonlake: Interactive, Multimodal World Models — with Chris Manning and Fan-yun Sun
Insight
Sun: Using structural abstraction in AI does not contradict Bitter Lesson
“I do feel like sometimes people confuse like, oh, like we're taking an, a method with abstraction. That means they don't believe in bitter lesson. Like that's just false, right? Like we are believers in bitter lesson, but then I feel like the question that we …”
Fan-yun Sun Apr 2, 2026 ▶ 14:37 Moonlake: Interactive, Multimodal World Models — with Chris Manning and Fan-yun Sun
Disclosure
Sun: Moonlake splits world modeling into multimodal reasoning and Reverie rendering
“Within our world modeling framework, we think there are two models that we train, right? Like there's the multimodal reasoning model that we just talked about that essentially handles Mainly the causality, the persistency, and logic, determinism, determinism o…”
Fan-yun Sun Apr 2, 2026 ▶ 28:25 Moonlake: Interactive, Multimodal World Models — with Chris Manning and Fan-yun Sun
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.