Jul 17, 2026 · 1h 14m · y-combinator

World Models, JEPA And The Path To Sample-Efficient RL · Y Combinator

0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of YC Decoded, Ankit Gupta and Francois Chaubard explore sample efficiency and world models in artificial intelligence, analyzing mathematical decision-making frameworks, robotics embodiment challenges, and latent predictive architectures like JEPA that aim to bridge the gap between human cognition and AI systems.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The partners as informed peer 6.4 Guest teaching 4.4 Guest disagreement 0.9 The partners pushing back 0.7
05100:0015:0030:0045:001:00:000:32–6:33 · The partners as informed peer 6/10 YC Decoded Title Sequence Ankit sets up the technical problem of sample efficiency and ARC-AGI benchmarks, demonstrating strong domain literacy. Francois builds upon the framing with cognitive science studies and neocortex evolution theories in a completely collaborative opening exchange.6:33–12:08 · The partners as informed peer 7/10 Optimal Control Theory and Differentiable Physics Francois walks through the mathematical formulation of optimal control, state vectors, and convex optimization via model predictive control. Ankit actively interjects to clarify the six-dimensional state vector and closed-form solutions.12:08–19:30 · The partners as informed peer 7/10 Non-Differentiable Systems and Reinforcement Learning Foundations The conversation shifts to non-differentiable stochastic systems with an adversarial multi-agent drone hypothetical. Ankit demonstrates solid technical comprehension of value functions, discounted rewards, and joint state-action modeling with video diffusion architectures.19:30–35:54 · The partners as informed peer 7/10 Evaluating Chess and AlphaGo Limits with MCTS Francois maps out state spaces and MCTS equations for chess and Go. Ankit pushes on combinatorial complexity, pointing out tree expansion dynamics and calculating sample limits if Go boards scale up.35:54–39:51 · The partners as informed peer 6/10 Self-Driving Cars and Non-Differentiable Multi-Agent Interactions Ankit contrasts discrete games with the continuous, infinite state space of autonomous driving. Francois explains non-differentiable multi-agent negotiation in driving environments like roundabouts.39:51–46:02 · The partners as informed peer 6/10 Y Combinator Application Announcement After the brief YC ad bumper, the pair examine action space discretization, tele-operation data scarcity, and the difference between model-free VLAs and model-based RL. Francois details the evolutionary significance of the neocortex.46:02–49:02 · The partners as informed peer 6/10 General Robotics Constraints and the Cross-Embodiment Gap The discussion covers high degrees of freedom in humanoid manipulators and the cross-embodiment gap across hardware variations. Ankit brings up Tesla fleet telemetry advantages while Francois discusses sharding data per vehicle model.49:02–57:00 · The partners as informed peer 7/10 Historical Evolution of World Models and Video Diffusion Francois traces the history of world models from Schmidhuber to Dreamer and recent video diffusion papers like DreamZero. Ankit connects these developments directly to generative flow matching and weather forecasting architectures.57:00–1:03:45 · The partners as informed peer 7/10 Latent Representation Learning and Joint Embedding Predictive Architectures Francois details the mathematics of Joint Embedding Predictive Architecture (JEPA), modal collapse mitigation via regularizers, and latent space optimization. Ankit provides parallel examples from graph neural network drug discovery pipelines.1:03:45–1:12:33 · The partners as informed peer 6/10 Open Challenges in Physics-Informed Networks and Sleep Consolidation Francois challenges the optimism around near-term robotics by pointing out fundamental failures in physics-informed neural networks and out-of-distribution interpolation. He highlights biological sleep consolidation and hippocampal sharp-wave ripples as missing components in modern AI.1:12:33–1:14:18 · The partners as informed peer 5/10 Tactile Sensing Gaps and Future Outlook for Robotics The episode wraps up on physical hardware sensing limitations, specifically tactile shear force and friction estimation in biological epidermis versus robotic sensors. Both hosts conclude collaboratively on the remaining research horizon.0:32–6:33 · Guest teaching 3/10 YC Decoded Title Sequence Ankit sets up the technical problem of sample efficiency and ARC-AGI benchmarks, demonstrating strong domain literacy. Francois builds upon the framing with cognitive science studies and neocortex evolution theories in a completely collaborative opening exchange.6:33–12:08 · Guest teaching 5/10 Optimal Control Theory and Differentiable Physics Francois walks through the mathematical formulation of optimal control, state vectors, and convex optimization via model predictive control. Ankit actively interjects to clarify the six-dimensional state vector and closed-form solutions.12:08–19:30 · Guest teaching 4/10 Non-Differentiable Systems and Reinforcement Learning Foundations The conversation shifts to non-differentiable stochastic systems with an adversarial multi-agent drone hypothetical. Ankit demonstrates solid technical comprehension of value functions, discounted rewards, and joint state-action modeling with video diffusion architectures.19:30–35:54 · Guest teaching 4/10 Evaluating Chess and AlphaGo Limits with MCTS Francois maps out state spaces and MCTS equations for chess and Go. Ankit pushes on combinatorial complexity, pointing out tree expansion dynamics and calculating sample limits if Go boards scale up.35:54–39:51 · Guest teaching 4/10 Self-Driving Cars and Non-Differentiable Multi-Agent Interactions Ankit contrasts discrete games with the continuous, infinite state space of autonomous driving. Francois explains non-differentiable multi-agent negotiation in driving environments like roundabouts.39:51–46:02 · Guest teaching 5/10 Y Combinator Application Announcement After the brief YC ad bumper, the pair examine action space discretization, tele-operation data scarcity, and the difference between model-free VLAs and model-based RL. Francois details the evolutionary significance of the neocortex.46:02–49:02 · Guest teaching 4/10 General Robotics Constraints and the Cross-Embodiment Gap The discussion covers high degrees of freedom in humanoid manipulators and the cross-embodiment gap across hardware variations. Ankit brings up Tesla fleet telemetry advantages while Francois discusses sharding data per vehicle model.49:02–57:00 · Guest teaching 4/10 Historical Evolution of World Models and Video Diffusion Francois traces the history of world models from Schmidhuber to Dreamer and recent video diffusion papers like DreamZero. Ankit connects these developments directly to generative flow matching and weather forecasting architectures.57:00–1:03:45 · Guest teaching 5/10 Latent Representation Learning and Joint Embedding Predictive Architectures Francois details the mathematics of Joint Embedding Predictive Architecture (JEPA), modal collapse mitigation via regularizers, and latent space optimization. Ankit provides parallel examples from graph neural network drug discovery pipelines.1:03:45–1:12:33 · Guest teaching 6/10 Open Challenges in Physics-Informed Networks and Sleep Consolidation Francois challenges the optimism around near-term robotics by pointing out fundamental failures in physics-informed neural networks and out-of-distribution interpolation. He highlights biological sleep consolidation and hippocampal sharp-wave ripples as missing components in modern AI.1:12:33–1:14:18 · Guest teaching 4/10 Tactile Sensing Gaps and Future Outlook for Robotics The episode wraps up on physical hardware sensing limitations, specifically tactile shear force and friction estimation in biological epidermis versus robotic sensors. Both hosts conclude collaboratively on the remaining research horizon.0:32–6:33 · Guest disagreement 1/10 YC Decoded Title Sequence Ankit sets up the technical problem of sample efficiency and ARC-AGI benchmarks, demonstrating strong domain literacy. Francois builds upon the framing with cognitive science studies and neocortex evolution theories in a completely collaborative opening exchange.6:33–12:08 · Guest disagreement 0/10 Optimal Control Theory and Differentiable Physics Francois walks through the mathematical formulation of optimal control, state vectors, and convex optimization via model predictive control. Ankit actively interjects to clarify the six-dimensional state vector and closed-form solutions.12:08–19:30 · Guest disagreement 1/10 Non-Differentiable Systems and Reinforcement Learning Foundations The conversation shifts to non-differentiable stochastic systems with an adversarial multi-agent drone hypothetical. Ankit demonstrates solid technical comprehension of value functions, discounted rewards, and joint state-action modeling with video diffusion architectures.19:30–35:54 · Guest disagreement 1/10 Evaluating Chess and AlphaGo Limits with MCTS Francois maps out state spaces and MCTS equations for chess and Go. Ankit pushes on combinatorial complexity, pointing out tree expansion dynamics and calculating sample limits if Go boards scale up.35:54–39:51 · Guest disagreement 1/10 Self-Driving Cars and Non-Differentiable Multi-Agent Interactions Ankit contrasts discrete games with the continuous, infinite state space of autonomous driving. Francois explains non-differentiable multi-agent negotiation in driving environments like roundabouts.39:51–46:02 · Guest disagreement 1/10 Y Combinator Application Announcement After the brief YC ad bumper, the pair examine action space discretization, tele-operation data scarcity, and the difference between model-free VLAs and model-based RL. Francois details the evolutionary significance of the neocortex.46:02–49:02 · Guest disagreement 1/10 General Robotics Constraints and the Cross-Embodiment Gap The discussion covers high degrees of freedom in humanoid manipulators and the cross-embodiment gap across hardware variations. Ankit brings up Tesla fleet telemetry advantages while Francois discusses sharding data per vehicle model.49:02–57:00 · Guest disagreement 0/10 Historical Evolution of World Models and Video Diffusion Francois traces the history of world models from Schmidhuber to Dreamer and recent video diffusion papers like DreamZero. Ankit connects these developments directly to generative flow matching and weather forecasting architectures.57:00–1:03:45 · Guest disagreement 1/10 Latent Representation Learning and Joint Embedding Predictive Architectures Francois details the mathematics of Joint Embedding Predictive Architecture (JEPA), modal collapse mitigation via regularizers, and latent space optimization. Ankit provides parallel examples from graph neural network drug discovery pipelines.1:03:45–1:12:33 · Guest disagreement 2/10 Open Challenges in Physics-Informed Networks and Sleep Consolidation Francois challenges the optimism around near-term robotics by pointing out fundamental failures in physics-informed neural networks and out-of-distribution interpolation. He highlights biological sleep consolidation and hippocampal sharp-wave ripples as missing components in modern AI.1:12:33–1:14:18 · Guest disagreement 1/10 Tactile Sensing Gaps and Future Outlook for Robotics The episode wraps up on physical hardware sensing limitations, specifically tactile shear force and friction estimation in biological epidermis versus robotic sensors. Both hosts conclude collaboratively on the remaining research horizon.0:32–6:33 · The partners pushing back 1/10 YC Decoded Title Sequence Ankit sets up the technical problem of sample efficiency and ARC-AGI benchmarks, demonstrating strong domain literacy. Francois builds upon the framing with cognitive science studies and neocortex evolution theories in a completely collaborative opening exchange.6:33–12:08 · The partners pushing back 0/10 Optimal Control Theory and Differentiable Physics Francois walks through the mathematical formulation of optimal control, state vectors, and convex optimization via model predictive control. Ankit actively interjects to clarify the six-dimensional state vector and closed-form solutions.12:08–19:30 · The partners pushing back 1/10 Non-Differentiable Systems and Reinforcement Learning Foundations The conversation shifts to non-differentiable stochastic systems with an adversarial multi-agent drone hypothetical. Ankit demonstrates solid technical comprehension of value functions, discounted rewards, and joint state-action modeling with video diffusion architectures.19:30–35:54 · The partners pushing back 2/10 Evaluating Chess and AlphaGo Limits with MCTS Francois maps out state spaces and MCTS equations for chess and Go. Ankit pushes on combinatorial complexity, pointing out tree expansion dynamics and calculating sample limits if Go boards scale up.35:54–39:51 · The partners pushing back 1/10 Self-Driving Cars and Non-Differentiable Multi-Agent Interactions Ankit contrasts discrete games with the continuous, infinite state space of autonomous driving. Francois explains non-differentiable multi-agent negotiation in driving environments like roundabouts.39:51–46:02 · The partners pushing back 1/10 Y Combinator Application Announcement After the brief YC ad bumper, the pair examine action space discretization, tele-operation data scarcity, and the difference between model-free VLAs and model-based RL. Francois details the evolutionary significance of the neocortex.46:02–49:02 · The partners pushing back 1/10 General Robotics Constraints and the Cross-Embodiment Gap The discussion covers high degrees of freedom in humanoid manipulators and the cross-embodiment gap across hardware variations. Ankit brings up Tesla fleet telemetry advantages while Francois discusses sharding data per vehicle model.49:02–57:00 · The partners pushing back 0/10 Historical Evolution of World Models and Video Diffusion Francois traces the history of world models from Schmidhuber to Dreamer and recent video diffusion papers like DreamZero. Ankit connects these developments directly to generative flow matching and weather forecasting architectures.57:00–1:03:45 · The partners pushing back 0/10 Latent Representation Learning and Joint Embedding Predictive Architectures Francois details the mathematics of Joint Embedding Predictive Architecture (JEPA), modal collapse mitigation via regularizers, and latent space optimization. Ankit provides parallel examples from graph neural network drug discovery pipelines.1:03:45–1:12:33 · The partners pushing back 1/10 Open Challenges in Physics-Informed Networks and Sleep Consolidation Francois challenges the optimism around near-term robotics by pointing out fundamental failures in physics-informed neural networks and out-of-distribution interpolation. He highlights biological sleep consolidation and hippocampal sharp-wave ripples as missing components in modern AI.1:12:33–1:14:18 · The partners pushing back 0/10 Tactile Sensing Gaps and Future Outlook for Robotics The episode wraps up on physical hardware sensing limitations, specifically tactile shear force and friction estimation in biological epidermis versus robotic sensors. Both hosts conclude collaboratively on the remaining research horizon.

speaking balance: gold is the partners, purple is the guest (3 minute bins)

0:00 · the partners 0% · guest 100%0:00 · the partners 0% · guest 100%3:00 · the partners 0% · guest 100%3:00 · the partners 0% · guest 100%6:00 · the partners 0% · guest 100%6:00 · the partners 0% · guest 100%9:00 · the partners 0% · guest 100%9:00 · the partners 0% · guest 100%12:00 · the partners 0% · guest 100%12:00 · the partners 0% · guest 100%15:00 · the partners 0% · guest 100%15:00 · the partners 0% · guest 100%18:00 · the partners 0% · guest 100%18:00 · the partners 0% · guest 100%21:00 · the partners 0% · guest 100%21:00 · the partners 0% · guest 100%24:00 · the partners 0% · guest 100%24:00 · the partners 0% · guest 100%27:00 · the partners 0% · guest 100%27:00 · the partners 0% · guest 100%30:00 · the partners 0% · guest 100%30:00 · the partners 0% · guest 100%33:00 · the partners 0% · guest 100%33:00 · the partners 0% · guest 100%36:00 · the partners 0% · guest 100%36:00 · the partners 0% · guest 100%39:00 · the partners 0% · guest 100%39:00 · the partners 0% · guest 100%42:00 · the partners 0% · guest 100%42:00 · the partners 0% · guest 100%45:00 · the partners 0% · guest 100%45:00 · the partners 0% · guest 100%48:00 · the partners 0% · guest 100%48:00 · the partners 0% · guest 100%51:00 · the partners 0% · guest 100%51:00 · the partners 0% · guest 100%54:00 · the partners 0% · guest 100%54:00 · the partners 0% · guest 100%57:00 · the partners 0% · guest 100%57:00 · the partners 0% · guest 100%1:00:00 · the partners 0% · guest 100%1:00:00 · the partners 0% · guest 100%1:03:00 · the partners 0% · guest 100%1:03:00 · the partners 0% · guest 100%1:06:00 · the partners 0% · guest 100%1:06:00 · the partners 0% · guest 100%1:09:00 · the partners 0% · guest 100%1:09:00 · the partners 0% · guest 100%1:12:00 · the partners 0% · guest 100%1:12:00 · the partners 0% · guest 100%
Sharpest disagreement ▶ 1:03:48 Dismissing near-term robot optimism

Francois bluntly rejects the optimistic notion that 2026 will bring general-purpose home humanoid robots, arguing that PINNs and current architectures fundamentally fail to capture precise physics.

Hardest push from the partners ▶ 22:15 Challenging Go state-space tractability assumptions

Ankit pushes back on the game tractability framing by illustrating how scaling up board dimensions causes Monte Carlo tree search test-time compute requirements to explode exponentially.

Biggest teaching moment ▶ 9:05 Deriving convex MPC trajectory optimization

Francois takes the host through the exact disciplinary convex programming formulation for real-time trajectory optimization, explaining how Lagrangian gradient descent solves deterministic control.

The partners hold their own ▶ 1:02:55 Synthesizing JEPA with graph neural network architectures

Ankit demonstrates deep firsthand expertise by mapping latent JEPA loss mechanics directly to his own background building graph convolutional networks for molecular generation.

the scores for every segment, with the reasoning behind each
ChapterTopicThe partners as informed peerGuest teachingGuest disagreementThe partners pushing backWhy
YC Decoded Title Sequence 6311 Ankit sets up the technical problem of sample efficiency and ARC-AGI benchmarks, demonstrating strong domain literacy. Francois builds upon the framing with cognitive science studies and neocortex evolution theories in a completely collaborative opening exchange.
Optimal Control Theory and Differentiable Physics 7500 Francois walks through the mathematical formulation of optimal control, state vectors, and convex optimization via model predictive control. Ankit actively interjects to clarify the six-dimensional state vector and closed-form solutions.
Non-Differentiable Systems and Reinforcement Learning Foundations 7411 The conversation shifts to non-differentiable stochastic systems with an adversarial multi-agent drone hypothetical. Ankit demonstrates solid technical comprehension of value functions, discounted rewards, and joint state-action modeling with video diffusion architectures.
Evaluating Chess and AlphaGo Limits with MCTS 7412 Francois maps out state spaces and MCTS equations for chess and Go. Ankit pushes on combinatorial complexity, pointing out tree expansion dynamics and calculating sample limits if Go boards scale up.
Self-Driving Cars and Non-Differentiable Multi-Agent Interactions 6411 Ankit contrasts discrete games with the continuous, infinite state space of autonomous driving. Francois explains non-differentiable multi-agent negotiation in driving environments like roundabouts.
Y Combinator Application Announcement 6511 After the brief YC ad bumper, the pair examine action space discretization, tele-operation data scarcity, and the difference between model-free VLAs and model-based RL. Francois details the evolutionary significance of the neocortex.
General Robotics Constraints and the Cross-Embodiment Gap 6411 The discussion covers high degrees of freedom in humanoid manipulators and the cross-embodiment gap across hardware variations. Ankit brings up Tesla fleet telemetry advantages while Francois discusses sharding data per vehicle model.
Historical Evolution of World Models and Video Diffusion 7400 Francois traces the history of world models from Schmidhuber to Dreamer and recent video diffusion papers like DreamZero. Ankit connects these developments directly to generative flow matching and weather forecasting architectures.
Latent Representation Learning and Joint Embedding Predictive Architectures 7510 Francois details the mathematics of Joint Embedding Predictive Architecture (JEPA), modal collapse mitigation via regularizers, and latent space optimization. Ankit provides parallel examples from graph neural network drug discovery pipelines.
Open Challenges in Physics-Informed Networks and Sleep Consolidation 6621 Francois challenges the optimism around near-term robotics by pointing out fundamental failures in physics-informed neural networks and out-of-distribution interpolation. He highlights biological sleep consolidation and hippocampal sharp-wave ripples as missing components in modern AI.
Tactile Sensing Gaps and Future Outlook for Robotics 5410 The episode wraps up on physical hardware sensing limitations, specifically tactile shear force and friction estimation in biological epidermis versus robotic sensors. Both hosts conclude collaboratively on the remaining research horizon.

Statements from this episode (0)

Nothing in this episode matches those filters. clear them

Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.