Dec 6, 2025 · 1h 4m · latent-space
World Models & General Intuition: Khosla's largest bet since LLMs & OpenAI
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this extensive studio interview and live technical demonstration, General Intuition CEO Pim de Witte explains how the startup utilizes Medal.tv's 3.8-billion gameplay dataset to build action-conditioned world models and vision-based foundation agents. De Witte details the technical architecture, fundraising journey, and long-term vision of transferring gaming-derived spatial-temporal intelligence into robotics and physical AI.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Pim directly challenges whether non-interactive 3D Gaussian splats constitute genuine world models, arguing interactivity is the defining criterion.
Hardest push from the hosts ▶ 45:03 Host presses on 3D physics splats vs 2D frame modelsThe host actively pushes back by citing his interview with Justin Johnson, arguing that physical forces applied to splats yield native 3D interactivity.
Biggest teaching moment ▶ 18:12 Pim explains action abstraction over raw keyboard loggingPim educates on why logging raw keystrokes introduces training noise and privacy liabilities, explaining how mapped action spaces solve both problems.
The host holds their own ▶ 47:10 Host cites Noam Brown on implicit world models in LLMsThe host demonstrates deep technical context by citing Noam Brown's counter-argument regarding implicit world model representations inside autoregressive LLMs.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Defining World Models and Action-State Transitions | 4 | 2 | 1 | 0 | The host delivers an introductory monologue summarizing General Intuition's origin from Medal, comparing data capture mechanics to Tesla bug reports, and detailing the recent $134M seed round. The tone is descriptive and supportive, setting up the interview context. | |
| Live Demonstration: Vision-Only Imitation Learning Gameplay | 5 | 4 | 1 | 2 | Pim walks through screen recordings of the vision-only imitation model navigating FPS environments. The host asks clarifying questions about memory horizon, goal-conditioning, and superhuman versus peak-human performance. | |
| Cross-Domain Action Prediction and Real-World Transfer | 5 | 4 | 2 | 2 | Pim demonstrates cross-game and real-world video transfer of action predictions. The host raises the distinction between egocentric (first-person) versus third-person perspectives, which Pim addresses by explaining multi-object control tradeoffs. | |
| World Models, Physical Dynamics, and Partial Observability | 4 | 5 | 1 | 1 | Pim explains world models handling partial observability (navigating through smoke screens) and inheriting real-world camera shake dynamics. The host tracks the visual demonstration with brief observational confirmations. | |
| Model Distillation and Optical Dynamics of Gameplay | 6 | 5 | 1 | 3 | The host probes the purpose of model distillation if the base model already runs in real time, prompting Pim to explain parameter efficiency. Pim then details how gaming input simulates optical dynamics and reduces information loss compared to YouTube pose estimation. | |
| Medal's 3.8B Dataset and Privacy-Preserving Action Mapping | 4 | 6 | 2 | 1 | Pim explains why Medal maps abstracted game actions instead of raw keystrokes (W/A/S/D) to protect user privacy while avoiding noisy training tokens. The host inquires into the origin of this insight and the manual labeling effort. | |
| Medal's Product Mechanics and Selective High-Value Recording | 5 | 4 | 1 | 2 | Pim details how Medal succeeded by focusing on lightweight retroactive in-memory recording rather than a heavy streaming suite. The host connects this to active learning and Tesla's selective data harvesting for corner cases. | |
| Technical Genesis: DIAMOND, SIMA, and Lab Recruitment | 6 | 5 | 2 | 3 | The host questions the tractability of unbounded continuous action spaces when switching from keyboard/mouse to general actions. Pim explains starting with discrete controller spaces before moving to action embeddings. | |
| GI's Strategic Moat and Environment Diversity | 5 | 5 | 2 | 3 | The host asks whether Meta Quest data could replicate Medal's moat. Pim pushes back by highlighting the necessity of public social graph permissions and the vast diversity disparity between PC game catalogs and VR environments. | |
| Frontier Research: Gaia-2, DIAMOND, and SIMA-2 Analysis | 6 | 4 | 1 | 2 | The host cites recent paper club reviews on Gaia-2, SIMA-2, and Genie-3. Pim breaks down SIMA-2's steerability and how an orchestrator VLM like Gemini can act as a high-level policy supervisor. | |
| Inside Vinod Khosla's $134M Seed Investment Process | 5 | 6 | 2 | 1 | Pim recounts Vinod Khosla's rigorous 2030 backward induction pitch process and shares practical advice on valuing proprietary data in AI deals. The host guides the reflection on commercial valuation and lab negotiations. | |
| Mastering Machine Learning Fundamentals: The Fleuret Course | 5 | 6 | 1 | 1 | Pim shares his personal learning journey through Francois Fleuret's deep learning coursework to master mathematical fundamentals. The host affirms the first-principles methodology. | |
| Contrasting World Model Architectures: World Labs vs. GI | 7 | 4 | 3 | 4 | The host directly contrasts GI's frame-based generation with World Labs' 3D Gaussian splat representation, citing his interview with Justin Johnson. Pim critiques non-interactive splats as incomplete world models while noting their representational merits. | |
| Spatial-Temporal Intelligence vs. Autoregressive Text Backbones | 6 | 5 | 3 | 3 | The host and guest debate whether LLMs inherently build world models or if autoregressive text tokens fail to capture continuous spatial-temporal dynamics. Pim positions text as a high-level compression layer rather than the foundational substrate. | |
| Commercial Deployment: Game Engine Bot APIs | 5 | 4 | 1 | 2 | Pim outlines GI's commercial model: providing game developers with an API to replace deterministic behavior trees with human-like bot policies to maintain player retention during off-peak hours. | |
| Cross-Domain Simulation: Truck Simulators, Self-Driving, and Robotics | 6 | 5 | 2 | 2 | Pim notes Medal has more steering wheel hours logged than Waymo fleet miles, positioning simulated driving as an active learning pre-training ground for robotics. The host validates the relevance of active learning and edge-case harvesting. | |
| RL Transition via Playable Clips and Episodic Memory | 5 | 6 | 1 | 1 | Pim articulates the transition from imitation learning to reinforcement learning by converting Medal's recorded episodic highlights into interactive, playable world model environments with reward modeling. | |
| Open Research Partnerships and the 2030 Physical AI Vision | 6 | 4 | 1 | 1 | Pim announces open research partnerships with Kyutai in Paris and outlines the 2030 target of powering 80% of atoms-to-atoms physical AI tasks. The host connects this to Chan Zuckerberg's virtual biology simulation initiatives in closing. |