Dec 6, 2025 · 1h 4m · latent-space

World Models & General Intuition: Khosla's largest bet since LLMs & OpenAI

Pim de Witte · 46m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this extensive studio interview and live technical demonstration, General Intuition CEO Pim de Witte explains how the startup utilizes Medal.tv's 3.8-billion gameplay dataset to build action-conditioned world models and vision-based foundation agents. De Witte details the technical architecture, fundraising journey, and long-term vision of transferring gaming-derived spatial-temporal intelligence into robotics and physical AI.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 5.3 Guest teaching 4.7 Guest disagreement 1.6 The hosts pushing back 1.9
05100:0015:0030:0045:001:00:000:00–2:42 · The hosts as informed peer 4/10 Defining World Models and Action-State Transitions The host delivers an introductory monologue summarizing General Intuition's origin from Medal, comparing data capture mechanics to Tesla bug reports, and detailing the recent $134M seed round. The tone is descriptive and supportive, setting up the interview context.2:42–5:55 · The hosts as informed peer 5/10 Live Demonstration: Vision-Only Imitation Learning Gameplay Pim walks through screen recordings of the vision-only imitation model navigating FPS environments. The host asks clarifying questions about memory horizon, goal-conditioning, and superhuman versus peak-human performance.5:56–9:14 · The hosts as informed peer 5/10 Cross-Domain Action Prediction and Real-World Transfer Pim demonstrates cross-game and real-world video transfer of action predictions. The host raises the distinction between egocentric (first-person) versus third-person perspectives, which Pim addresses by explaining multi-object control tradeoffs.9:14–12:35 · The hosts as informed peer 4/10 World Models, Physical Dynamics, and Partial Observability Pim explains world models handling partial observability (navigating through smoke screens) and inheriting real-world camera shake dynamics. The host tracks the visual demonstration with brief observational confirmations.12:35–17:13 · The hosts as informed peer 6/10 Model Distillation and Optical Dynamics of Gameplay The host probes the purpose of model distillation if the base model already runs in real time, prompting Pim to explain parameter efficiency. Pim then details how gaming input simulates optical dynamics and reduces information loss compared to YouTube pose estimation.17:13–20:34 · The hosts as informed peer 4/10 Medal's 3.8B Dataset and Privacy-Preserving Action Mapping Pim explains why Medal maps abstracted game actions instead of raw keystrokes (W/A/S/D) to protect user privacy while avoiding noisy training tokens. The host inquires into the origin of this insight and the manual labeling effort.20:35–23:56 · The hosts as informed peer 5/10 Medal's Product Mechanics and Selective High-Value Recording Pim details how Medal succeeded by focusing on lightweight retroactive in-memory recording rather than a heavy streaming suite. The host connects this to active learning and Tesla's selective data harvesting for corner cases.23:56–27:24 · The hosts as informed peer 6/10 Technical Genesis: DIAMOND, SIMA, and Lab Recruitment The host questions the tractability of unbounded continuous action spaces when switching from keyboard/mouse to general actions. Pim explains starting with discrete controller spaces before moving to action embeddings.27:26–30:32 · The hosts as informed peer 5/10 GI's Strategic Moat and Environment Diversity The host asks whether Meta Quest data could replicate Medal's moat. Pim pushes back by highlighting the necessity of public social graph permissions and the vast diversity disparity between PC game catalogs and VR environments.30:34–33:28 · The hosts as informed peer 6/10 Frontier Research: Gaia-2, DIAMOND, and SIMA-2 Analysis The host cites recent paper club reviews on Gaia-2, SIMA-2, and Genie-3. Pim breaks down SIMA-2's steerability and how an orchestrator VLM like Gemini can act as a high-level policy supervisor.33:28–39:05 · The hosts as informed peer 5/10 Inside Vinod Khosla's $134M Seed Investment Process Pim recounts Vinod Khosla's rigorous 2030 backward induction pitch process and shares practical advice on valuing proprietary data in AI deals. The host guides the reflection on commercial valuation and lab negotiations.39:05–43:56 · The hosts as informed peer 5/10 Mastering Machine Learning Fundamentals: The Fleuret Course Pim shares his personal learning journey through Francois Fleuret's deep learning coursework to master mathematical fundamentals. The host affirms the first-principles methodology.43:56–47:10 · The hosts as informed peer 7/10 Contrasting World Model Architectures: World Labs vs. GI The host directly contrasts GI's frame-based generation with World Labs' 3D Gaussian splat representation, citing his interview with Justin Johnson. Pim critiques non-interactive splats as incomplete world models while noting their representational merits.47:10–50:31 · The hosts as informed peer 6/10 Spatial-Temporal Intelligence vs. Autoregressive Text Backbones The host and guest debate whether LLMs inherently build world models or if autoregressive text tokens fail to capture continuous spatial-temporal dynamics. Pim positions text as a high-level compression layer rather than the foundational substrate.50:31–54:20 · The hosts as informed peer 5/10 Commercial Deployment: Game Engine Bot APIs Pim outlines GI's commercial model: providing game developers with an API to replace deterministic behavior trees with human-like bot policies to maintain player retention during off-peak hours.54:22–57:39 · The hosts as informed peer 6/10 Cross-Domain Simulation: Truck Simulators, Self-Driving, and Robotics Pim notes Medal has more steering wheel hours logged than Waymo fleet miles, positioning simulated driving as an active learning pre-training ground for robotics. The host validates the relevance of active learning and edge-case harvesting.57:40–1:00:43 · The hosts as informed peer 5/10 RL Transition via Playable Clips and Episodic Memory Pim articulates the transition from imitation learning to reinforcement learning by converting Medal's recorded episodic highlights into interactive, playable world model environments with reward modeling.1:00:43–1:04:40 · The hosts as informed peer 6/10 Open Research Partnerships and the 2030 Physical AI Vision Pim announces open research partnerships with Kyutai in Paris and outlines the 2030 target of powering 80% of atoms-to-atoms physical AI tasks. The host connects this to Chan Zuckerberg's virtual biology simulation initiatives in closing.0:00–2:42 · Guest teaching 2/10 Defining World Models and Action-State Transitions The host delivers an introductory monologue summarizing General Intuition's origin from Medal, comparing data capture mechanics to Tesla bug reports, and detailing the recent $134M seed round. The tone is descriptive and supportive, setting up the interview context.2:42–5:55 · Guest teaching 4/10 Live Demonstration: Vision-Only Imitation Learning Gameplay Pim walks through screen recordings of the vision-only imitation model navigating FPS environments. The host asks clarifying questions about memory horizon, goal-conditioning, and superhuman versus peak-human performance.5:56–9:14 · Guest teaching 4/10 Cross-Domain Action Prediction and Real-World Transfer Pim demonstrates cross-game and real-world video transfer of action predictions. The host raises the distinction between egocentric (first-person) versus third-person perspectives, which Pim addresses by explaining multi-object control tradeoffs.9:14–12:35 · Guest teaching 5/10 World Models, Physical Dynamics, and Partial Observability Pim explains world models handling partial observability (navigating through smoke screens) and inheriting real-world camera shake dynamics. The host tracks the visual demonstration with brief observational confirmations.12:35–17:13 · Guest teaching 5/10 Model Distillation and Optical Dynamics of Gameplay The host probes the purpose of model distillation if the base model already runs in real time, prompting Pim to explain parameter efficiency. Pim then details how gaming input simulates optical dynamics and reduces information loss compared to YouTube pose estimation.17:13–20:34 · Guest teaching 6/10 Medal's 3.8B Dataset and Privacy-Preserving Action Mapping Pim explains why Medal maps abstracted game actions instead of raw keystrokes (W/A/S/D) to protect user privacy while avoiding noisy training tokens. The host inquires into the origin of this insight and the manual labeling effort.20:35–23:56 · Guest teaching 4/10 Medal's Product Mechanics and Selective High-Value Recording Pim details how Medal succeeded by focusing on lightweight retroactive in-memory recording rather than a heavy streaming suite. The host connects this to active learning and Tesla's selective data harvesting for corner cases.23:56–27:24 · Guest teaching 5/10 Technical Genesis: DIAMOND, SIMA, and Lab Recruitment The host questions the tractability of unbounded continuous action spaces when switching from keyboard/mouse to general actions. Pim explains starting with discrete controller spaces before moving to action embeddings.27:26–30:32 · Guest teaching 5/10 GI's Strategic Moat and Environment Diversity The host asks whether Meta Quest data could replicate Medal's moat. Pim pushes back by highlighting the necessity of public social graph permissions and the vast diversity disparity between PC game catalogs and VR environments.30:34–33:28 · Guest teaching 4/10 Frontier Research: Gaia-2, DIAMOND, and SIMA-2 Analysis The host cites recent paper club reviews on Gaia-2, SIMA-2, and Genie-3. Pim breaks down SIMA-2's steerability and how an orchestrator VLM like Gemini can act as a high-level policy supervisor.33:28–39:05 · Guest teaching 6/10 Inside Vinod Khosla's $134M Seed Investment Process Pim recounts Vinod Khosla's rigorous 2030 backward induction pitch process and shares practical advice on valuing proprietary data in AI deals. The host guides the reflection on commercial valuation and lab negotiations.39:05–43:56 · Guest teaching 6/10 Mastering Machine Learning Fundamentals: The Fleuret Course Pim shares his personal learning journey through Francois Fleuret's deep learning coursework to master mathematical fundamentals. The host affirms the first-principles methodology.43:56–47:10 · Guest teaching 4/10 Contrasting World Model Architectures: World Labs vs. GI The host directly contrasts GI's frame-based generation with World Labs' 3D Gaussian splat representation, citing his interview with Justin Johnson. Pim critiques non-interactive splats as incomplete world models while noting their representational merits.47:10–50:31 · Guest teaching 5/10 Spatial-Temporal Intelligence vs. Autoregressive Text Backbones The host and guest debate whether LLMs inherently build world models or if autoregressive text tokens fail to capture continuous spatial-temporal dynamics. Pim positions text as a high-level compression layer rather than the foundational substrate.50:31–54:20 · Guest teaching 4/10 Commercial Deployment: Game Engine Bot APIs Pim outlines GI's commercial model: providing game developers with an API to replace deterministic behavior trees with human-like bot policies to maintain player retention during off-peak hours.54:22–57:39 · Guest teaching 5/10 Cross-Domain Simulation: Truck Simulators, Self-Driving, and Robotics Pim notes Medal has more steering wheel hours logged than Waymo fleet miles, positioning simulated driving as an active learning pre-training ground for robotics. The host validates the relevance of active learning and edge-case harvesting.57:40–1:00:43 · Guest teaching 6/10 RL Transition via Playable Clips and Episodic Memory Pim articulates the transition from imitation learning to reinforcement learning by converting Medal's recorded episodic highlights into interactive, playable world model environments with reward modeling.1:00:43–1:04:40 · Guest teaching 4/10 Open Research Partnerships and the 2030 Physical AI Vision Pim announces open research partnerships with Kyutai in Paris and outlines the 2030 target of powering 80% of atoms-to-atoms physical AI tasks. The host connects this to Chan Zuckerberg's virtual biology simulation initiatives in closing.0:00–2:42 · Guest disagreement 1/10 Defining World Models and Action-State Transitions The host delivers an introductory monologue summarizing General Intuition's origin from Medal, comparing data capture mechanics to Tesla bug reports, and detailing the recent $134M seed round. The tone is descriptive and supportive, setting up the interview context.2:42–5:55 · Guest disagreement 1/10 Live Demonstration: Vision-Only Imitation Learning Gameplay Pim walks through screen recordings of the vision-only imitation model navigating FPS environments. The host asks clarifying questions about memory horizon, goal-conditioning, and superhuman versus peak-human performance.5:56–9:14 · Guest disagreement 2/10 Cross-Domain Action Prediction and Real-World Transfer Pim demonstrates cross-game and real-world video transfer of action predictions. The host raises the distinction between egocentric (first-person) versus third-person perspectives, which Pim addresses by explaining multi-object control tradeoffs.9:14–12:35 · Guest disagreement 1/10 World Models, Physical Dynamics, and Partial Observability Pim explains world models handling partial observability (navigating through smoke screens) and inheriting real-world camera shake dynamics. The host tracks the visual demonstration with brief observational confirmations.12:35–17:13 · Guest disagreement 1/10 Model Distillation and Optical Dynamics of Gameplay The host probes the purpose of model distillation if the base model already runs in real time, prompting Pim to explain parameter efficiency. Pim then details how gaming input simulates optical dynamics and reduces information loss compared to YouTube pose estimation.17:13–20:34 · Guest disagreement 2/10 Medal's 3.8B Dataset and Privacy-Preserving Action Mapping Pim explains why Medal maps abstracted game actions instead of raw keystrokes (W/A/S/D) to protect user privacy while avoiding noisy training tokens. The host inquires into the origin of this insight and the manual labeling effort.20:35–23:56 · Guest disagreement 1/10 Medal's Product Mechanics and Selective High-Value Recording Pim details how Medal succeeded by focusing on lightweight retroactive in-memory recording rather than a heavy streaming suite. The host connects this to active learning and Tesla's selective data harvesting for corner cases.23:56–27:24 · Guest disagreement 2/10 Technical Genesis: DIAMOND, SIMA, and Lab Recruitment The host questions the tractability of unbounded continuous action spaces when switching from keyboard/mouse to general actions. Pim explains starting with discrete controller spaces before moving to action embeddings.27:26–30:32 · Guest disagreement 2/10 GI's Strategic Moat and Environment Diversity The host asks whether Meta Quest data could replicate Medal's moat. Pim pushes back by highlighting the necessity of public social graph permissions and the vast diversity disparity between PC game catalogs and VR environments.30:34–33:28 · Guest disagreement 1/10 Frontier Research: Gaia-2, DIAMOND, and SIMA-2 Analysis The host cites recent paper club reviews on Gaia-2, SIMA-2, and Genie-3. Pim breaks down SIMA-2's steerability and how an orchestrator VLM like Gemini can act as a high-level policy supervisor.33:28–39:05 · Guest disagreement 2/10 Inside Vinod Khosla's $134M Seed Investment Process Pim recounts Vinod Khosla's rigorous 2030 backward induction pitch process and shares practical advice on valuing proprietary data in AI deals. The host guides the reflection on commercial valuation and lab negotiations.39:05–43:56 · Guest disagreement 1/10 Mastering Machine Learning Fundamentals: The Fleuret Course Pim shares his personal learning journey through Francois Fleuret's deep learning coursework to master mathematical fundamentals. The host affirms the first-principles methodology.43:56–47:10 · Guest disagreement 3/10 Contrasting World Model Architectures: World Labs vs. GI The host directly contrasts GI's frame-based generation with World Labs' 3D Gaussian splat representation, citing his interview with Justin Johnson. Pim critiques non-interactive splats as incomplete world models while noting their representational merits.47:10–50:31 · Guest disagreement 3/10 Spatial-Temporal Intelligence vs. Autoregressive Text Backbones The host and guest debate whether LLMs inherently build world models or if autoregressive text tokens fail to capture continuous spatial-temporal dynamics. Pim positions text as a high-level compression layer rather than the foundational substrate.50:31–54:20 · Guest disagreement 1/10 Commercial Deployment: Game Engine Bot APIs Pim outlines GI's commercial model: providing game developers with an API to replace deterministic behavior trees with human-like bot policies to maintain player retention during off-peak hours.54:22–57:39 · Guest disagreement 2/10 Cross-Domain Simulation: Truck Simulators, Self-Driving, and Robotics Pim notes Medal has more steering wheel hours logged than Waymo fleet miles, positioning simulated driving as an active learning pre-training ground for robotics. The host validates the relevance of active learning and edge-case harvesting.57:40–1:00:43 · Guest disagreement 1/10 RL Transition via Playable Clips and Episodic Memory Pim articulates the transition from imitation learning to reinforcement learning by converting Medal's recorded episodic highlights into interactive, playable world model environments with reward modeling.1:00:43–1:04:40 · Guest disagreement 1/10 Open Research Partnerships and the 2030 Physical AI Vision Pim announces open research partnerships with Kyutai in Paris and outlines the 2030 target of powering 80% of atoms-to-atoms physical AI tasks. The host connects this to Chan Zuckerberg's virtual biology simulation initiatives in closing.0:00–2:42 · The hosts pushing back 0/10 Defining World Models and Action-State Transitions The host delivers an introductory monologue summarizing General Intuition's origin from Medal, comparing data capture mechanics to Tesla bug reports, and detailing the recent $134M seed round. The tone is descriptive and supportive, setting up the interview context.2:42–5:55 · The hosts pushing back 2/10 Live Demonstration: Vision-Only Imitation Learning Gameplay Pim walks through screen recordings of the vision-only imitation model navigating FPS environments. The host asks clarifying questions about memory horizon, goal-conditioning, and superhuman versus peak-human performance.5:56–9:14 · The hosts pushing back 2/10 Cross-Domain Action Prediction and Real-World Transfer Pim demonstrates cross-game and real-world video transfer of action predictions. The host raises the distinction between egocentric (first-person) versus third-person perspectives, which Pim addresses by explaining multi-object control tradeoffs.9:14–12:35 · The hosts pushing back 1/10 World Models, Physical Dynamics, and Partial Observability Pim explains world models handling partial observability (navigating through smoke screens) and inheriting real-world camera shake dynamics. The host tracks the visual demonstration with brief observational confirmations.12:35–17:13 · The hosts pushing back 3/10 Model Distillation and Optical Dynamics of Gameplay The host probes the purpose of model distillation if the base model already runs in real time, prompting Pim to explain parameter efficiency. Pim then details how gaming input simulates optical dynamics and reduces information loss compared to YouTube pose estimation.17:13–20:34 · The hosts pushing back 1/10 Medal's 3.8B Dataset and Privacy-Preserving Action Mapping Pim explains why Medal maps abstracted game actions instead of raw keystrokes (W/A/S/D) to protect user privacy while avoiding noisy training tokens. The host inquires into the origin of this insight and the manual labeling effort.20:35–23:56 · The hosts pushing back 2/10 Medal's Product Mechanics and Selective High-Value Recording Pim details how Medal succeeded by focusing on lightweight retroactive in-memory recording rather than a heavy streaming suite. The host connects this to active learning and Tesla's selective data harvesting for corner cases.23:56–27:24 · The hosts pushing back 3/10 Technical Genesis: DIAMOND, SIMA, and Lab Recruitment The host questions the tractability of unbounded continuous action spaces when switching from keyboard/mouse to general actions. Pim explains starting with discrete controller spaces before moving to action embeddings.27:26–30:32 · The hosts pushing back 3/10 GI's Strategic Moat and Environment Diversity The host asks whether Meta Quest data could replicate Medal's moat. Pim pushes back by highlighting the necessity of public social graph permissions and the vast diversity disparity between PC game catalogs and VR environments.30:34–33:28 · The hosts pushing back 2/10 Frontier Research: Gaia-2, DIAMOND, and SIMA-2 Analysis The host cites recent paper club reviews on Gaia-2, SIMA-2, and Genie-3. Pim breaks down SIMA-2's steerability and how an orchestrator VLM like Gemini can act as a high-level policy supervisor.33:28–39:05 · The hosts pushing back 1/10 Inside Vinod Khosla's $134M Seed Investment Process Pim recounts Vinod Khosla's rigorous 2030 backward induction pitch process and shares practical advice on valuing proprietary data in AI deals. The host guides the reflection on commercial valuation and lab negotiations.39:05–43:56 · The hosts pushing back 1/10 Mastering Machine Learning Fundamentals: The Fleuret Course Pim shares his personal learning journey through Francois Fleuret's deep learning coursework to master mathematical fundamentals. The host affirms the first-principles methodology.43:56–47:10 · The hosts pushing back 4/10 Contrasting World Model Architectures: World Labs vs. GI The host directly contrasts GI's frame-based generation with World Labs' 3D Gaussian splat representation, citing his interview with Justin Johnson. Pim critiques non-interactive splats as incomplete world models while noting their representational merits.47:10–50:31 · The hosts pushing back 3/10 Spatial-Temporal Intelligence vs. Autoregressive Text Backbones The host and guest debate whether LLMs inherently build world models or if autoregressive text tokens fail to capture continuous spatial-temporal dynamics. Pim positions text as a high-level compression layer rather than the foundational substrate.50:31–54:20 · The hosts pushing back 2/10 Commercial Deployment: Game Engine Bot APIs Pim outlines GI's commercial model: providing game developers with an API to replace deterministic behavior trees with human-like bot policies to maintain player retention during off-peak hours.54:22–57:39 · The hosts pushing back 2/10 Cross-Domain Simulation: Truck Simulators, Self-Driving, and Robotics Pim notes Medal has more steering wheel hours logged than Waymo fleet miles, positioning simulated driving as an active learning pre-training ground for robotics. The host validates the relevance of active learning and edge-case harvesting.57:40–1:00:43 · The hosts pushing back 1/10 RL Transition via Playable Clips and Episodic Memory Pim articulates the transition from imitation learning to reinforcement learning by converting Medal's recorded episodic highlights into interactive, playable world model environments with reward modeling.1:00:43–1:04:40 · The hosts pushing back 1/10 Open Research Partnerships and the 2030 Physical AI Vision Pim announces open research partnerships with Kyutai in Paris and outlines the 2030 target of powering 80% of atoms-to-atoms physical AI tasks. The host connects this to Chan Zuckerberg's virtual biology simulation initiatives in closing.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 44:04 Pim challenges World Labs 3D splat interactivity

Pim directly challenges whether non-interactive 3D Gaussian splats constitute genuine world models, arguing interactivity is the defining criterion.

Hardest push from the hosts ▶ 45:03 Host presses on 3D physics splats vs 2D frame models

The host actively pushes back by citing his interview with Justin Johnson, arguing that physical forces applied to splats yield native 3D interactivity.

Biggest teaching moment ▶ 18:12 Pim explains action abstraction over raw keyboard logging

Pim educates on why logging raw keystrokes introduces training noise and privacy liabilities, explaining how mapped action spaces solve both problems.

The host holds their own ▶ 47:10 Host cites Noam Brown on implicit world models in LLMs

The host demonstrates deep technical context by citing Noam Brown's counter-argument regarding implicit world model representations inside autoregressive LLMs.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Defining World Models and Action-State Transitions 4210 The host delivers an introductory monologue summarizing General Intuition's origin from Medal, comparing data capture mechanics to Tesla bug reports, and detailing the recent $134M seed round. The tone is descriptive and supportive, setting up the interview context.
Live Demonstration: Vision-Only Imitation Learning Gameplay 5412 Pim walks through screen recordings of the vision-only imitation model navigating FPS environments. The host asks clarifying questions about memory horizon, goal-conditioning, and superhuman versus peak-human performance.
Cross-Domain Action Prediction and Real-World Transfer 5422 Pim demonstrates cross-game and real-world video transfer of action predictions. The host raises the distinction between egocentric (first-person) versus third-person perspectives, which Pim addresses by explaining multi-object control tradeoffs.
World Models, Physical Dynamics, and Partial Observability 4511 Pim explains world models handling partial observability (navigating through smoke screens) and inheriting real-world camera shake dynamics. The host tracks the visual demonstration with brief observational confirmations.
Model Distillation and Optical Dynamics of Gameplay 6513 The host probes the purpose of model distillation if the base model already runs in real time, prompting Pim to explain parameter efficiency. Pim then details how gaming input simulates optical dynamics and reduces information loss compared to YouTube pose estimation.
Medal's 3.8B Dataset and Privacy-Preserving Action Mapping 4621 Pim explains why Medal maps abstracted game actions instead of raw keystrokes (W/A/S/D) to protect user privacy while avoiding noisy training tokens. The host inquires into the origin of this insight and the manual labeling effort.
Medal's Product Mechanics and Selective High-Value Recording 5412 Pim details how Medal succeeded by focusing on lightweight retroactive in-memory recording rather than a heavy streaming suite. The host connects this to active learning and Tesla's selective data harvesting for corner cases.
Technical Genesis: DIAMOND, SIMA, and Lab Recruitment 6523 The host questions the tractability of unbounded continuous action spaces when switching from keyboard/mouse to general actions. Pim explains starting with discrete controller spaces before moving to action embeddings.
GI's Strategic Moat and Environment Diversity 5523 The host asks whether Meta Quest data could replicate Medal's moat. Pim pushes back by highlighting the necessity of public social graph permissions and the vast diversity disparity between PC game catalogs and VR environments.
Frontier Research: Gaia-2, DIAMOND, and SIMA-2 Analysis 6412 The host cites recent paper club reviews on Gaia-2, SIMA-2, and Genie-3. Pim breaks down SIMA-2's steerability and how an orchestrator VLM like Gemini can act as a high-level policy supervisor.
Inside Vinod Khosla's $134M Seed Investment Process 5621 Pim recounts Vinod Khosla's rigorous 2030 backward induction pitch process and shares practical advice on valuing proprietary data in AI deals. The host guides the reflection on commercial valuation and lab negotiations.
Mastering Machine Learning Fundamentals: The Fleuret Course 5611 Pim shares his personal learning journey through Francois Fleuret's deep learning coursework to master mathematical fundamentals. The host affirms the first-principles methodology.
Contrasting World Model Architectures: World Labs vs. GI 7434 The host directly contrasts GI's frame-based generation with World Labs' 3D Gaussian splat representation, citing his interview with Justin Johnson. Pim critiques non-interactive splats as incomplete world models while noting their representational merits.
Spatial-Temporal Intelligence vs. Autoregressive Text Backbones 6533 The host and guest debate whether LLMs inherently build world models or if autoregressive text tokens fail to capture continuous spatial-temporal dynamics. Pim positions text as a high-level compression layer rather than the foundational substrate.
Commercial Deployment: Game Engine Bot APIs 5412 Pim outlines GI's commercial model: providing game developers with an API to replace deterministic behavior trees with human-like bot policies to maintain player retention during off-peak hours.
Cross-Domain Simulation: Truck Simulators, Self-Driving, and Robotics 6522 Pim notes Medal has more steering wheel hours logged than Waymo fleet miles, positioning simulated driving as an active learning pre-training ground for robotics. The host validates the relevance of active learning and edge-case harvesting.
RL Transition via Playable Clips and Episodic Memory 5611 Pim articulates the transition from imitation learning to reinforcement learning by converting Medal's recorded episodic highlights into interactive, playable world model environments with reward modeling.
Open Research Partnerships and the 2030 Physical AI Vision 6411 Pim announces open research partnerships with Kyutai in Paris and outlines the 2030 target of powering 80% of atoms-to-atoms physical AI tasks. The host connects this to Chan Zuckerberg's virtual biology simulation initiatives in closing.

Statements from this episode (34)

Assertion Not checkable as stated
General Intuition's foundation agent runs purely on vision without reinforcement learning
“This is just a base model. There's no RL, no fine tuning. This model sees no game states. It is purely capable, not sequence acceptance. It's purely predicting the actions from the phrase. That's it.”
Pim de Witte Dec 6, 2025 ▶ 4:04
Assertion Not checkable as stated
The baseline of General Intuition's training dataset is peak human gameplay
“The baseline of our data set is PQ and performance.”
Pim de Witte Dec 6, 2025 ▶ 5:49
Assertion Not checkable as stated
General Intuition action models turn internet videos into free training data
“We transferred it over to a real world video, which means that you can use any video on the internet as free training.”
Pim de Witte Dec 6, 2025 ▶ 7:01
Insight
Third-person vision models will be key for controlling multiple environment objects
“The third person I think will be very, very helpful if you're, for instance, trying to control multiple objects in an environment later on. Right now, I think having fully in perception first person is quite helpful.”
Pim de Witte Dec 6, 2025 ▶ 8:11
Disclosure
General Intuition pre-trains from scratch and fine-tunes open-source video models
“We made the decision to Pre-trained world models from scratch, but also we've actually been able to fine tune open source video models to get a better sense of physical transfer.”
Pim de Witte Dec 6, 2025 ▶ 9:30
Assertion Not checkable as stated
General Intuition world models support high-speed mouse sensitivity capabilities
“Our world models have mouse sensitivity, which is something that, like, gamers absolutely want, right? So you can have these, like, very rapid movements, which you couldn't do in any other world model.”
Pim de Witte Dec 6, 2025 ▶ 9:42
Opinion
Gameplay footage teaches spatial reasoning better than YouTube videos, says de Witte
“When you're playing video games, you're actually simulating the optical dynamics with your hand, right? And I think that, like, that's why I think why games are a better representation of switch support reasoning initially than YouTube videos, for instance.”
Pim de Witte Dec 6, 2025 ▶ 14:01
Assertion Not checkable as stated
Medal holds the internet's largest action-labeled video dataset by orders of magnitude
“We have sort of the largest data set of ground truth action labeled video footage on the internet by maybe one or two orders of magnitude.”
Pim de Witte Dec 6, 2025 ▶ 17:48
Disclosure
Medal employed thousands of humans to label game actions over 18 months
“We had thousands of humans label every single action you can take in every single video game over the past year and a half which is an enormous amount of action labels.”
Pim de Witte Dec 6, 2025 ▶ 19:28
Insight
Medal successfully bootstrapped its social network by prioritizing single-player capture tools
“A lot of our competitors were focused on solving the social network and the recorder at the same time, and that never, like, our bet was really that we could get so many people to record with us that we could bootstrap the network on top of that, and that work…”
Pim de Witte Dec 6, 2025 ▶ 20:57
Disclosure
General Intuition will initially model game controller inputs before general action embeddings
“So I think we're going to start with anything you can control using a game controller. But yeah, long term, we want to actually predict maybe like action embeddings and have models sit inside a general action space to be able to transfer out to other inputs as…”
Pim de Witte Dec 6, 2025 ▶ 26:08
Disclosure
Frontier AI research labs aggressively tried to acquire or recruit Medal's team
“So like, right when that happened, a lot of the labs also started on, started understanding what we had. And so we started very aggressively, multiple labs tried to bring us in, in various ways.”
Pim de Witte Dec 6, 2025 ▶ 27:08
Assertion Not checkable as stated
PC gaming offers tens of thousands of scaled environments compared to VR
“The amount of environments in VR that are, that, that have, like, consumption at scale is probably in, like, the hundreds. whereas on PC, it's probably in the tens of thousands.”
Pim de Witte Dec 6, 2025 ▶ 29:57
Disclosure
Lead researchers from the GAIA-2 and DIAMOND papers joined General Intuition
“Anthony Tu, who led the research on Gaia II is, is also one of the engineers that joined our team. So it's all the Diamond the core contributors for Diamond, and then Anthony and we just had three more researchers joining this week.”
Pim de Witte Dec 6, 2025 ▶ 30:50
Prediction Open · timeframe Dec 2028
General Intuition's world models will initially be orchestrated by Vision-Language Models
“I think our models will initially be used as like you'll have like an orchestrator VLM of sorts. That's kind of like managing instances and instructing them.”
Pim de Witte Dec 6, 2025 ▶ 32:44
Insight
Vinod Khosla tests founders by rigorously reverse-engineering their five-year vision
“He asked you to, like, draw a twenty-thirty picture of your company, and I think he just picks And plus five years. Whatever. I don't know. I did the same to you. Yeah. He asks you to like walk that back from first principles all the way from today. And he ask…”
Pim de Witte Dec 6, 2025 ▶ 33:48
Insight
Proprietary data cannot be accurately valued without training models to evaluate capabilities
“I don't think you can value it unless you actually model it yourself and see what the capabilities are.”
Pim de Witte Dec 6, 2025 ▶ 35:44
Insight
Founders negotiating large AI data deals should demand equity, says de Witte
“I would recommend, if you're gonna do large data deals, like, just try to get like a large chunk of equity in the company that you're doing it with if you can.”
Pim de Witte Dec 6, 2025 ▶ 36:22
Insight
World models are significantly more complex than traditional generative video models
“What world models do is they actually have to understand the full range of possibilities and outcomes from the current states and based on the action that you take generates the next states, right? So the next frame. And so it is a much more sort of complex pr…”
Pim de Witte Dec 6, 2025 ▶ 41:14
Insight
Physics simulation compute complexity scales rapidly with agents, freedoms, and information
“The compute complexity of simulation goes up really, really rapidly with three variables. First, the numbers of agents in an environment. Second they're DOF. So they're individuals of freedom. Yeah. And then third, the information that each action reveals.”
Pim de Witte Dec 6, 2025 ▶ 42:30
Disclosure
General Intuition is betting maximally on video-to-action transfer over physics simulation
“And so I think for us, it is more so about making a maximal bet on video transfer and interacting with things that are difficult to simulate. And the steerability is also really interesting with text than it is on betting against simulation or something like t…”
Pim de Witte Dec 6, 2025 ▶ 43:22
Opinion
De Witte: Static 3D environments without interactivity are not true world models
“However, my understanding is they're currently not interactive, which in my opinion is like the whole point of world models, right? It's environments. They're great environments. And I think from a business perspective, I think they picked a really important p…”
Pim de Witte Dec 6, 2025 ▶ 44:27
Disclosure
Yann LeCun calling LLMs a dead end inspired General Intuition's founding
“Honestly, Jan's podcast that he did, I don't remember which one it was, but a long time ago, where he basically proclaimed LLMs to be a dead end it was one of the things that inspired me to do this.”
Pim de Witte Dec 6, 2025 ▶ 46:58
Insight
Text and language function fundamentally as compression methods for 3D perception
“The way I think about it is, at Siemens, right, we had sort of a three-dimensional world, then we invented text. As like a, in a way, a compression method, right? So you had, we invented text in order to communicate with each other in a common way, in a way th…”
Pim de Witte Dec 6, 2025 ▶ 47:41
Prediction Not checkable as stated
Text and speech generation will become actions emitted by world models
“I think text prediction is just one of the actions that is going to come out of these, you know, these policies and world models. I think speech and text generation will just be one of the actions that, that can be a part of that.”
Pim de Witte Dec 6, 2025 ▶ 49:27
Prediction Not checkable as stated
Language model labs and world model labs will ultimately converge on capabilities
“I think that there will just be labs coming at this problem from both sides. And everyone ends up in roughly the same place, and the same place will be whatever people think is cool.”
Pim de Witte Dec 6, 2025 ▶ 49:41
Insight
Player retention in scaled video games heavily depends on automated bot quality
“So if you're a game developer, how well you're actually retaining players, it's like if you have a game that's already at scale, it's like decently dependent on how good your bots are.”
Pim de Witte Dec 6, 2025 ▶ 53:00
Assertion Not checkable as stated
Medal has more concurrent sim-drivers than Waymo has autonomous cars
“We have more people at any given time on metal playing with steering wheels and like truck simulator and these types of games than Waymo has cars on the road. It's a ridiculous stat, but it's true.”
Pim de Witte Dec 6, 2025 ▶ 54:36
Disclosure
General Intuition's initial commercial business model will be an API
“Our business model is initially going to be an API, again, like the Anthropic API”
Pim de Witte Dec 6, 2025 ▶ 57:03
Insight
Making billions of gameplay clips playable bridges imitation learning to reinforcement learning
“Actually making every single clip on the platform playable at billions of clips scale is how we go from imitation learning to RL.”
Pim de Witte Dec 6, 2025 ▶ 1:00:30
Opinion
General Intuition publishes openly because its proprietary data moat prevents model replication
“Because we have such a large data mode, we don't have to be as concerned as the LLM companies about publishing, because we don't need the ones to be able to. Exactly. No one can replicate the models, right?”
Pim de Witte Dec 6, 2025 ▶ 1:01:10
Assertion Supported
General Intuition partnered with Eric Schmidt-backed Kyutai for open data research
“We just announced our partnership with Qtai in France, which is an open science lab in Paris, one of the best research labs in the world. Eric Schmidt, I believe funded in addition to some French people. They are essentially acting as the partner that's curren…”
Pim de Witte Dec 6, 2025 ▶ 1:01:29
Prediction Not checkable as stated
General Intuition aims to drive 80% of AI physical interactions by 2030
“In the atoms to atoms stage, I want, like, I want GI models to be responsible for 80% of all the atoms to atoms interactions driven by AI models and the reason for that is because we were able to unblock intelligence so quickly, and robotics, like, intelligenc…”
Pim de Witte Dec 6, 2025 ▶ 1:03:11
Prediction Not checkable as stated
Simulation will initially be a larger commercial AI market than physical deployment
“I think simulation will actually be the larger market initially. So I think in simulation because you have very little constraints also from a safety perspective, simulation is much easier. So I think a lot of the takeoff initially sits in simulation.”
Pim de Witte Dec 6, 2025 ▶ 1:03:57
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.