Jul 17, 2026 · 1h 14m · y-combinator
World Models, JEPA And The Path To Sample-Efficient RL · Y Combinator
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of YC Decoded, Ankit Gupta and Francois Chaubard explore sample efficiency and world models in artificial intelligence, analyzing mathematical decision-making frameworks, robotics embodiment challenges, and latent predictive architectures like JEPA that aim to bridge the gap between human cognition and AI systems.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the partners, purple is the guest (3 minute bins)
Francois bluntly rejects the optimistic notion that 2026 will bring general-purpose home humanoid robots, arguing that PINNs and current architectures fundamentally fail to capture precise physics.
Hardest push from the partners ▶ 22:15 Challenging Go state-space tractability assumptionsAnkit pushes back on the game tractability framing by illustrating how scaling up board dimensions causes Monte Carlo tree search test-time compute requirements to explode exponentially.
Biggest teaching moment ▶ 9:05 Deriving convex MPC trajectory optimizationFrancois takes the host through the exact disciplinary convex programming formulation for real-time trajectory optimization, explaining how Lagrangian gradient descent solves deterministic control.
The partners hold their own ▶ 1:02:55 Synthesizing JEPA with graph neural network architecturesAnkit demonstrates deep firsthand expertise by mapping latent JEPA loss mechanics directly to his own background building graph convolutional networks for molecular generation.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The partners as informed peer | Guest teaching | Guest disagreement | The partners pushing back | Why |
|---|---|---|---|---|---|---|
| YC Decoded Title Sequence | 6 | 3 | 1 | 1 | Ankit sets up the technical problem of sample efficiency and ARC-AGI benchmarks, demonstrating strong domain literacy. Francois builds upon the framing with cognitive science studies and neocortex evolution theories in a completely collaborative opening exchange. | |
| Optimal Control Theory and Differentiable Physics | 7 | 5 | 0 | 0 | Francois walks through the mathematical formulation of optimal control, state vectors, and convex optimization via model predictive control. Ankit actively interjects to clarify the six-dimensional state vector and closed-form solutions. | |
| Non-Differentiable Systems and Reinforcement Learning Foundations | 7 | 4 | 1 | 1 | The conversation shifts to non-differentiable stochastic systems with an adversarial multi-agent drone hypothetical. Ankit demonstrates solid technical comprehension of value functions, discounted rewards, and joint state-action modeling with video diffusion architectures. | |
| Evaluating Chess and AlphaGo Limits with MCTS | 7 | 4 | 1 | 2 | Francois maps out state spaces and MCTS equations for chess and Go. Ankit pushes on combinatorial complexity, pointing out tree expansion dynamics and calculating sample limits if Go boards scale up. | |
| Self-Driving Cars and Non-Differentiable Multi-Agent Interactions | 6 | 4 | 1 | 1 | Ankit contrasts discrete games with the continuous, infinite state space of autonomous driving. Francois explains non-differentiable multi-agent negotiation in driving environments like roundabouts. | |
| Y Combinator Application Announcement | 6 | 5 | 1 | 1 | After the brief YC ad bumper, the pair examine action space discretization, tele-operation data scarcity, and the difference between model-free VLAs and model-based RL. Francois details the evolutionary significance of the neocortex. | |
| General Robotics Constraints and the Cross-Embodiment Gap | 6 | 4 | 1 | 1 | The discussion covers high degrees of freedom in humanoid manipulators and the cross-embodiment gap across hardware variations. Ankit brings up Tesla fleet telemetry advantages while Francois discusses sharding data per vehicle model. | |
| Historical Evolution of World Models and Video Diffusion | 7 | 4 | 0 | 0 | Francois traces the history of world models from Schmidhuber to Dreamer and recent video diffusion papers like DreamZero. Ankit connects these developments directly to generative flow matching and weather forecasting architectures. | |
| Latent Representation Learning and Joint Embedding Predictive Architectures | 7 | 5 | 1 | 0 | Francois details the mathematics of Joint Embedding Predictive Architecture (JEPA), modal collapse mitigation via regularizers, and latent space optimization. Ankit provides parallel examples from graph neural network drug discovery pipelines. | |
| Open Challenges in Physics-Informed Networks and Sleep Consolidation | 6 | 6 | 2 | 1 | Francois challenges the optimism around near-term robotics by pointing out fundamental failures in physics-informed neural networks and out-of-distribution interpolation. He highlights biological sleep consolidation and hippocampal sharp-wave ripples as missing components in modern AI. | |
| Tactile Sensing Gaps and Future Outlook for Robotics | 5 | 4 | 1 | 0 | The episode wraps up on physical hardware sensing limitations, specifically tactile shear force and friction estimation in biological epidermis versus robotic sensors. Both hosts conclude collaboratively on the remaining research horizon. |
Statements from this episode (0)
Nothing in this episode matches those filters. clear them