Mar 20, 2025 · 35m · no-priors

No Priors Ep. 107 | With Physical Intelligence Co-Founder Chelsea Finn

Chelsea Finn · 24m spoken Elad Gil · 8m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of No Priors, Stanford professor and Physical Intelligence co-founder Chelsea Finn explores the breakthroughs, architectural designs, and data strategies powering general-purpose robotic foundation models. She discusses multi-embodiment learning, hardware pragmaticism, and the critical importance of real-world environmental diversity in achieving physical artificial intelligence.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 25% of the talking time here. How this is scored →

The hosts as informed peer 5.5 Guest teaching 4.2 Guest disagreement 0.9 The hosts pushing back 0.7
05100:0010:0020:0030:000:34–3:09 · The hosts as informed peer 4/10 Chelsea Finn's Background and Robotic Research Evolution Gil opens with a polite biographical prompt acknowledging Finn's storied academic and industry research career. Finn gives an expansive overview of pixel-to-torque motor control evolution and the persistent bottleneck of broad generalizability.3:09–7:37 · The hosts as informed peer 5/10 The Mission and Multi-Embodiment Strategy of Physical Intelligence Gil draws relevant analogies to LLM architecture and scaling trends. Finn explains Physical Intelligence's multi-embodiment data collection approach and how vision-language web pre-training enables zero-shot physical transfer.7:37–9:46 · The hosts as informed peer 6/10 Achieving True Generalizability Through Environmental Diversity and Reasoning Gil asks technical questions distinguishing compute scaling from reasoning modules and post-training. Finn clarifies that environmental and task diversity across multiple physical sites is the primary bottleneck rather than pure model reasoning.9:46–12:31 · The hosts as informed peer 5/10 Open-Source Strategy, Hardware Sharing, and Core Engineering Risks Gil probes the commercial trade-offs between open source, open core, and proprietary models. Finn explains why sharing IP and hardware accelerates the ecosystem, framing the existential risk not as competition but as robotics being intrinsically unforgiving of motor errors.12:31–16:37 · The hosts as informed peer 5/10 Commercial Viability, Error Tolerance, and Human-Robot Collaboration Gil asks about humanoid embodiment versus domain-specialized form factors. Finn takes a mildly contrarian stance, arguing humanoids are overrated because teleoperation data collection is far slower and more cumbersome on them than on simpler platforms.16:38–20:07 · The hosts as informed peer 6/10 The Deep Complexity and Underestimated Value of Embodied Intelligence Gil demonstrates deep familiarity with recent literature by citing Finn's ALOHA paper. Finn points out that AI hype often underappreciates evolutionary motor intelligence, listing SayCan, RT-2, and RT-X as seminal turning points.20:07–22:20 · The hosts as informed peer 5/10 HiRobot: Hierarchical Interactive Architectures for Long-Horizon Tasks Gil prompts Finn on the announcement of HiRobot. Finn explains the architecture's division of labor between high-level prompt reasoning and low-level reactive motor commands.22:20–25:20 · The hosts as informed peer 6/10 Sensor Modalities, Tactile Feedback, and the Primacy of Policy Memory Gil contrasts camera-only systems with multi-sensor AV stacks like Waymo's LiDAR. Finn educates on why current tactile sensors are fragile and low-resolution, emphasizing that temporal policy memory is far more urgent than new sensory modalities.25:20–28:36 · The hosts as informed peer 7/10 Startup Agility, Autonomous Vehicle Comparisons, and Corporate Constraints Gil provides a strong analytical breakdown of why autonomous vehicles consolidated around early incumbents like Tesla and Waymo. Finn agrees on the difficulty of physical AI but notes startup velocity and Google's internal security red tape as key differentiators.28:36–31:56 · The hosts as informed peer 5/10 Entrepreneurial Advice for Aspiring Robotics Founders Gil asks whether passive observational human video can substitute for robot data collection. Finn pushes back using an Olympic swimmer analogy, arguing passive observation cannot teach proprioceptive motor control without embodied experience.31:56–34:47 · The hosts as informed peer 7/10 The Future Hardware Landscape: Humanoids Versus a Cambrian Explosion Gil references Neal Stephenson's The Diamond Age and argues manufacturing supply chain economies favor a few standardized form factors. Finn counters that downstream intelligent robots could manufacture bespoke, specialized hardware directly.0:34–3:09 · Guest teaching 3/10 Chelsea Finn's Background and Robotic Research Evolution Gil opens with a polite biographical prompt acknowledging Finn's storied academic and industry research career. Finn gives an expansive overview of pixel-to-torque motor control evolution and the persistent bottleneck of broad generalizability.3:09–7:37 · Guest teaching 4/10 The Mission and Multi-Embodiment Strategy of Physical Intelligence Gil draws relevant analogies to LLM architecture and scaling trends. Finn explains Physical Intelligence's multi-embodiment data collection approach and how vision-language web pre-training enables zero-shot physical transfer.7:37–9:46 · Guest teaching 4/10 Achieving True Generalizability Through Environmental Diversity and Reasoning Gil asks technical questions distinguishing compute scaling from reasoning modules and post-training. Finn clarifies that environmental and task diversity across multiple physical sites is the primary bottleneck rather than pure model reasoning.9:46–12:31 · Guest teaching 5/10 Open-Source Strategy, Hardware Sharing, and Core Engineering Risks Gil probes the commercial trade-offs between open source, open core, and proprietary models. Finn explains why sharing IP and hardware accelerates the ecosystem, framing the existential risk not as competition but as robotics being intrinsically unforgiving of motor errors.12:31–16:37 · Guest teaching 5/10 Commercial Viability, Error Tolerance, and Human-Robot Collaboration Gil asks about humanoid embodiment versus domain-specialized form factors. Finn takes a mildly contrarian stance, arguing humanoids are overrated because teleoperation data collection is far slower and more cumbersome on them than on simpler platforms.16:38–20:07 · Guest teaching 4/10 The Deep Complexity and Underestimated Value of Embodied Intelligence Gil demonstrates deep familiarity with recent literature by citing Finn's ALOHA paper. Finn points out that AI hype often underappreciates evolutionary motor intelligence, listing SayCan, RT-2, and RT-X as seminal turning points.20:07–22:20 · Guest teaching 4/10 HiRobot: Hierarchical Interactive Architectures for Long-Horizon Tasks Gil prompts Finn on the announcement of HiRobot. Finn explains the architecture's division of labor between high-level prompt reasoning and low-level reactive motor commands.22:20–25:20 · Guest teaching 5/10 Sensor Modalities, Tactile Feedback, and the Primacy of Policy Memory Gil contrasts camera-only systems with multi-sensor AV stacks like Waymo's LiDAR. Finn educates on why current tactile sensors are fragile and low-resolution, emphasizing that temporal policy memory is far more urgent than new sensory modalities.25:20–28:36 · Guest teaching 3/10 Startup Agility, Autonomous Vehicle Comparisons, and Corporate Constraints Gil provides a strong analytical breakdown of why autonomous vehicles consolidated around early incumbents like Tesla and Waymo. Finn agrees on the difficulty of physical AI but notes startup velocity and Google's internal security red tape as key differentiators.28:36–31:56 · Guest teaching 5/10 Entrepreneurial Advice for Aspiring Robotics Founders Gil asks whether passive observational human video can substitute for robot data collection. Finn pushes back using an Olympic swimmer analogy, arguing passive observation cannot teach proprioceptive motor control without embodied experience.31:56–34:47 · Guest teaching 4/10 The Future Hardware Landscape: Humanoids Versus a Cambrian Explosion Gil references Neal Stephenson's The Diamond Age and argues manufacturing supply chain economies favor a few standardized form factors. Finn counters that downstream intelligent robots could manufacture bespoke, specialized hardware directly.0:34–3:09 · Guest disagreement 0/10 Chelsea Finn's Background and Robotic Research Evolution Gil opens with a polite biographical prompt acknowledging Finn's storied academic and industry research career. Finn gives an expansive overview of pixel-to-torque motor control evolution and the persistent bottleneck of broad generalizability.3:09–7:37 · Guest disagreement 0/10 The Mission and Multi-Embodiment Strategy of Physical Intelligence Gil draws relevant analogies to LLM architecture and scaling trends. Finn explains Physical Intelligence's multi-embodiment data collection approach and how vision-language web pre-training enables zero-shot physical transfer.7:37–9:46 · Guest disagreement 0/10 Achieving True Generalizability Through Environmental Diversity and Reasoning Gil asks technical questions distinguishing compute scaling from reasoning modules and post-training. Finn clarifies that environmental and task diversity across multiple physical sites is the primary bottleneck rather than pure model reasoning.9:46–12:31 · Guest disagreement 1/10 Open-Source Strategy, Hardware Sharing, and Core Engineering Risks Gil probes the commercial trade-offs between open source, open core, and proprietary models. Finn explains why sharing IP and hardware accelerates the ecosystem, framing the existential risk not as competition but as robotics being intrinsically unforgiving of motor errors.12:31–16:37 · Guest disagreement 2/10 Commercial Viability, Error Tolerance, and Human-Robot Collaboration Gil asks about humanoid embodiment versus domain-specialized form factors. Finn takes a mildly contrarian stance, arguing humanoids are overrated because teleoperation data collection is far slower and more cumbersome on them than on simpler platforms.16:38–20:07 · Guest disagreement 1/10 The Deep Complexity and Underestimated Value of Embodied Intelligence Gil demonstrates deep familiarity with recent literature by citing Finn's ALOHA paper. Finn points out that AI hype often underappreciates evolutionary motor intelligence, listing SayCan, RT-2, and RT-X as seminal turning points.20:07–22:20 · Guest disagreement 0/10 HiRobot: Hierarchical Interactive Architectures for Long-Horizon Tasks Gil prompts Finn on the announcement of HiRobot. Finn explains the architecture's division of labor between high-level prompt reasoning and low-level reactive motor commands.22:20–25:20 · Guest disagreement 1/10 Sensor Modalities, Tactile Feedback, and the Primacy of Policy Memory Gil contrasts camera-only systems with multi-sensor AV stacks like Waymo's LiDAR. Finn educates on why current tactile sensors are fragile and low-resolution, emphasizing that temporal policy memory is far more urgent than new sensory modalities.25:20–28:36 · Guest disagreement 1/10 Startup Agility, Autonomous Vehicle Comparisons, and Corporate Constraints Gil provides a strong analytical breakdown of why autonomous vehicles consolidated around early incumbents like Tesla and Waymo. Finn agrees on the difficulty of physical AI but notes startup velocity and Google's internal security red tape as key differentiators.28:36–31:56 · Guest disagreement 2/10 Entrepreneurial Advice for Aspiring Robotics Founders Gil asks whether passive observational human video can substitute for robot data collection. Finn pushes back using an Olympic swimmer analogy, arguing passive observation cannot teach proprioceptive motor control without embodied experience.31:56–34:47 · Guest disagreement 2/10 The Future Hardware Landscape: Humanoids Versus a Cambrian Explosion Gil references Neal Stephenson's The Diamond Age and argues manufacturing supply chain economies favor a few standardized form factors. Finn counters that downstream intelligent robots could manufacture bespoke, specialized hardware directly.0:34–3:09 · The hosts pushing back 0/10 Chelsea Finn's Background and Robotic Research Evolution Gil opens with a polite biographical prompt acknowledging Finn's storied academic and industry research career. Finn gives an expansive overview of pixel-to-torque motor control evolution and the persistent bottleneck of broad generalizability.3:09–7:37 · The hosts pushing back 0/10 The Mission and Multi-Embodiment Strategy of Physical Intelligence Gil draws relevant analogies to LLM architecture and scaling trends. Finn explains Physical Intelligence's multi-embodiment data collection approach and how vision-language web pre-training enables zero-shot physical transfer.7:37–9:46 · The hosts pushing back 0/10 Achieving True Generalizability Through Environmental Diversity and Reasoning Gil asks technical questions distinguishing compute scaling from reasoning modules and post-training. Finn clarifies that environmental and task diversity across multiple physical sites is the primary bottleneck rather than pure model reasoning.9:46–12:31 · The hosts pushing back 1/10 Open-Source Strategy, Hardware Sharing, and Core Engineering Risks Gil probes the commercial trade-offs between open source, open core, and proprietary models. Finn explains why sharing IP and hardware accelerates the ecosystem, framing the existential risk not as competition but as robotics being intrinsically unforgiving of motor errors.12:31–16:37 · The hosts pushing back 1/10 Commercial Viability, Error Tolerance, and Human-Robot Collaboration Gil asks about humanoid embodiment versus domain-specialized form factors. Finn takes a mildly contrarian stance, arguing humanoids are overrated because teleoperation data collection is far slower and more cumbersome on them than on simpler platforms.16:38–20:07 · The hosts pushing back 0/10 The Deep Complexity and Underestimated Value of Embodied Intelligence Gil demonstrates deep familiarity with recent literature by citing Finn's ALOHA paper. Finn points out that AI hype often underappreciates evolutionary motor intelligence, listing SayCan, RT-2, and RT-X as seminal turning points.20:07–22:20 · The hosts pushing back 0/10 HiRobot: Hierarchical Interactive Architectures for Long-Horizon Tasks Gil prompts Finn on the announcement of HiRobot. Finn explains the architecture's division of labor between high-level prompt reasoning and low-level reactive motor commands.22:20–25:20 · The hosts pushing back 1/10 Sensor Modalities, Tactile Feedback, and the Primacy of Policy Memory Gil contrasts camera-only systems with multi-sensor AV stacks like Waymo's LiDAR. Finn educates on why current tactile sensors are fragile and low-resolution, emphasizing that temporal policy memory is far more urgent than new sensory modalities.25:20–28:36 · The hosts pushing back 1/10 Startup Agility, Autonomous Vehicle Comparisons, and Corporate Constraints Gil provides a strong analytical breakdown of why autonomous vehicles consolidated around early incumbents like Tesla and Waymo. Finn agrees on the difficulty of physical AI but notes startup velocity and Google's internal security red tape as key differentiators.28:36–31:56 · The hosts pushing back 0/10 Entrepreneurial Advice for Aspiring Robotics Founders Gil asks whether passive observational human video can substitute for robot data collection. Finn pushes back using an Olympic swimmer analogy, arguing passive observation cannot teach proprioceptive motor control without embodied experience.31:56–34:47 · The hosts pushing back 4/10 The Future Hardware Landscape: Humanoids Versus a Cambrian Explosion Gil references Neal Stephenson's The Diamond Age and argues manufacturing supply chain economies favor a few standardized form factors. Finn counters that downstream intelligent robots could manufacture bespoke, specialized hardware directly.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 23.9% · guest 76.1%0:00 · the hosts 23.9% · guest 76.1%3:00 · the hosts 20.4% · guest 79.6%3:00 · the hosts 20.4% · guest 79.6%6:00 · the hosts 15.1% · guest 84.9%6:00 · the hosts 15.1% · guest 84.9%9:00 · the hosts 18.9% · guest 81.1%9:00 · the hosts 18.9% · guest 81.1%12:00 · the hosts 41.9% · guest 58.1%12:00 · the hosts 41.9% · guest 58.1%15:00 · the hosts 30% · guest 70%15:00 · the hosts 30% · guest 70%18:00 · the hosts 7.6% · guest 92.4%18:00 · the hosts 7.6% · guest 92.4%21:00 · the hosts 27.9% · guest 72.1%21:00 · the hosts 27.9% · guest 72.1%24:00 · the hosts 26.5% · guest 73.5%24:00 · the hosts 26.5% · guest 73.5%27:00 · the hosts 22.2% · guest 77.8%27:00 · the hosts 22.2% · guest 77.8%30:00 · the hosts 21.8% · guest 78.2%30:00 · the hosts 21.8% · guest 78.2%33:00 · the hosts 52.3% · guest 47.7%33:00 · the hosts 52.3% · guest 47.7%
Sharpest disagreement ▶ 29:36 Refuting passive video data via the Olympic swimmer analogy

Finn directly dismisses the premise that web video or human observation is sufficient for robotics, pointing out that watching an Olympic swimmer never provides the requisite motor coordination.

Hardest push from the hosts ▶ 34:12 Gil pushes back on hardware variety using supply chain constraints

Gil directly challenges Finn's hypothesis of a Cambrian explosion of robot forms by citing the overwhelming cost and manufacturing efficiencies of standardized supply chains.

Biggest teaching moment ▶ 24:50 Finn explains policy memory deficiency over tactile sensing

Finn educates the host on current hardware limits in tactile sensing, demonstrating that existing policies cannot even retain half-second history and that temporal memory is a far higher technical priority.

The host holds their own ▶ 26:21 Gil analyzes autonomous driving market dynamics and incumbent advantage

Gil displays deep industry knowledge by outlining how dozens of autonomous vehicle startups over fifteen years collapsed into two capital-intensive incumbents.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Chelsea Finn's Background and Robotic Research Evolution 4300 Gil opens with a polite biographical prompt acknowledging Finn's storied academic and industry research career. Finn gives an expansive overview of pixel-to-torque motor control evolution and the persistent bottleneck of broad generalizability.
The Mission and Multi-Embodiment Strategy of Physical Intelligence 5400 Gil draws relevant analogies to LLM architecture and scaling trends. Finn explains Physical Intelligence's multi-embodiment data collection approach and how vision-language web pre-training enables zero-shot physical transfer.
Achieving True Generalizability Through Environmental Diversity and Reasoning 6400 Gil asks technical questions distinguishing compute scaling from reasoning modules and post-training. Finn clarifies that environmental and task diversity across multiple physical sites is the primary bottleneck rather than pure model reasoning.
Open-Source Strategy, Hardware Sharing, and Core Engineering Risks 5511 Gil probes the commercial trade-offs between open source, open core, and proprietary models. Finn explains why sharing IP and hardware accelerates the ecosystem, framing the existential risk not as competition but as robotics being intrinsically unforgiving of motor errors.
Commercial Viability, Error Tolerance, and Human-Robot Collaboration 5521 Gil asks about humanoid embodiment versus domain-specialized form factors. Finn takes a mildly contrarian stance, arguing humanoids are overrated because teleoperation data collection is far slower and more cumbersome on them than on simpler platforms.
The Deep Complexity and Underestimated Value of Embodied Intelligence 6410 Gil demonstrates deep familiarity with recent literature by citing Finn's ALOHA paper. Finn points out that AI hype often underappreciates evolutionary motor intelligence, listing SayCan, RT-2, and RT-X as seminal turning points.
HiRobot: Hierarchical Interactive Architectures for Long-Horizon Tasks 5400 Gil prompts Finn on the announcement of HiRobot. Finn explains the architecture's division of labor between high-level prompt reasoning and low-level reactive motor commands.
Sensor Modalities, Tactile Feedback, and the Primacy of Policy Memory 6511 Gil contrasts camera-only systems with multi-sensor AV stacks like Waymo's LiDAR. Finn educates on why current tactile sensors are fragile and low-resolution, emphasizing that temporal policy memory is far more urgent than new sensory modalities.
Startup Agility, Autonomous Vehicle Comparisons, and Corporate Constraints 7311 Gil provides a strong analytical breakdown of why autonomous vehicles consolidated around early incumbents like Tesla and Waymo. Finn agrees on the difficulty of physical AI but notes startup velocity and Google's internal security red tape as key differentiators.
Entrepreneurial Advice for Aspiring Robotics Founders 5520 Gil asks whether passive observational human video can substitute for robot data collection. Finn pushes back using an Olympic swimmer analogy, arguing passive observation cannot teach proprioceptive motor control without embodied experience.
The Future Hardware Landscape: Humanoids Versus a Cambrian Explosion 7424 Gil references Neal Stephenson's The Diamond Age and argues manufacturing supply chain economies favor a few standardized form factors. Finn counters that downstream intelligent robots could manufacture bespoke, specialized hardware directly.

Statements from this episode (23)

Assertion Not checkable as stated
Finn: End-to-end pixel-to-torque robot control shifted from unpopular to mainstream
“I started working more seriously in robotics more than 10 years ago at this point at the start of my PhD at Berkeley, and we were working on neural network control Trying to train neural networks that map from image pixels to directly actually to motor torques…”
Chelsea Finn Mar 20, 2025 ▶ 1:17
Disclosure
Physical Intelligence aims to build a single model for any robot
“So we're trying to build a big neural network model that could ultimately control any robot to do anything in any scenario.”
Chelsea Finn Mar 20, 2025 ▶ 3:22
Assertion Supported
Finn: Robot training data transfers across different physical hardware embodiments
“We've seen a lot of evidence that you could actually transfer a lot of rich information across these different embodiments and allows you to use data. And also if you iterate on your robot platform, you don't have to throw all your data away.”
Chelsea Finn Mar 20, 2025 ▶ 4:21
Disclosure
Physical Intelligence uses transformers and pre-trained VLMs for robot models
“And in terms of the architecture, we're using transformers and we are using pre-trained models, pre-trained vision language models, and that allows you to leverage all of the rich information on the internet.”
Chelsea Finn Mar 20, 2025 ▶ 6:51
Insight
Finn: Pre-trained VLMs let robots perform tasks with unseen internet concepts
“We had a research result a couple years ago where we showed that if you leverage vision language models, then you could actually get the robot to do tasks that require concepts that were never in the robot's training data, but were in the internet.”
Chelsea Finn Mar 20, 2025 ▶ 7:05
Disclosure
Finn: Physical Intelligence's October release trained on data from three buildings
“So for that release that we had in late October last year, we collected data in three buildings, technically.”
Chelsea Finn Mar 20, 2025 ▶ 8:11
Disclosure
Finn: Physical Intelligence gives robot designs directly to hardware companies
“We've not only have we open-sourced some of the weights and released details in, in technical papers, we've actually also been working with hardware companies and giving designs of robots to hardware companies.”
Chelsea Finn Mar 20, 2025 ▶ 10:21
Opinion
Finn: Biggest risk in generalist robotics is technical failure, not competition
“The last thing that I'll mention is that I think the biggest risk with this bet is that it won't work. Like, I'm not really worried about competitors. I'm more worried that no one will solve the problem.”
Chelsea Finn Mar 20, 2025 ▶ 11:45
Insight
Finn: Robotics is harder than digital ML because humans cannot verify real-time outputs
“Typically in machine learning, a lot of the successful applications of like recommender systems, language models like image detection, a lot of the consumers of that Of the model outputs are actually humans who could actually check it, and the humans are good …”
Chelsea Finn Mar 20, 2025 ▶ 13:31
Opinion
Finn: Humanoids are overrated because collecting teleoperation data is too difficult
“On the other hand, I think that they're a little overrated, and one way it kind of to practically look at it is, I think that we're generally fairly bottlenecked on data right now, and some people argue that with humanoids, you can maybe collect data more easi…”
Chelsea Finn Mar 20, 2025 ▶ 15:05
Disclosure
Physical Intelligence prioritizes cheap robots for faster teleoperation data collection
“That's one of the things we're kind of optimizing for, and so we're using cheap robots. We're using robots that we can very easily develop teleoperation interfaces for, in which you can do teleoperation very quickly and collect diverse data, collect lots of da…”
Chelsea Finn Mar 20, 2025 ▶ 15:54
Insight
Finn: Embodied AI and motor control are underrated compared to language models
“I feel like actually people underestimate how much intelligence goes into motor control. Many, many years of evolution is what led to us being able to use our hands the way that we do. And there are many animals that they can't do it even though they had so mu…”
Chelsea Finn Mar 20, 2025 ▶ 16:51
Assertion Supported
Finn: General multi-robot models outperformed research labs' custom single-robot policies
“We actually found that we could take a checkpoint, send that model checkpoint to another lab halfway across the country, and the grad student at that lab could run the checkpoint on the robot, and it would actually More often than not do better than the model …”
Chelsea Finn Mar 20, 2025 ▶ 18:29
Assertion Supported
Finn: ALOHA proved teleoperation trains complex dexterous robotic manipulation
“I think the Aloha work on, and later the mobile Aloha work was work that showed that you can teleoperate and get models to train pretty complicated dexterous manipulation tasks.”
Chelsea Finn Mar 20, 2025 ▶ 19:15
Insight
Finn: Direct action policies fail on long-horizon multi-minute robotics tasks
“If you need to do like a longer horizon task, meaning a task that might take minutes to do, then if you just train a single policy to like output actions based on images like if you're trying to make a sandwich and you train a policy that's just outputting the…”
Chelsea Finn Mar 20, 2025 ▶ 20:15
Assertion Not checkable as stated
Finn: Current robotic tactile sensors are fragile, expensive, or low-resolution
“Unfortunately, a lot of the tactile sensors that are out there Are either far less robust than skin, far more expensive or very, very low resolution.”
Chelsea Finn Mar 20, 2025 ▶ 23:18
Insight
Finn: Wrist-mounted RGB cameras capture much of what tactile sensors provide
“And we found that actually that mounting RGB cameras to the wrists ends up being very, very helpful and probably giving you a lot of the same information that tactile sensors can give you.”
Chelsea Finn Mar 20, 2025 ▶ 23:31
Disclosure
Finn: Physical Intelligence's current state-of-the-art robot policies operate entirely without memory
“The other thing that I'll mention is actually right now we're most like, our policies right now do not have any memory. They only look at the current image frame. They can't remember even half a second prior.”
Chelsea Finn Mar 20, 2025 ▶ 24:56
Insight
Finn: Robotics avoids self-driving's edge-case distribution problem by finding commercial niches
“With driving, I feel like you kind of need to solve the entire distribution to be have anything that's viable. You have to be able to handle an intersection at any time of day or with any kind of possible pedestrian scenario or other cars and all that. Whereas…”
Chelsea Finn Mar 20, 2025 ▶ 25:44
Disclosure
Finn: Google's code security rules made real-world robot data collection nearly impossible
“As one example, taking a robot off campus was like almost a non-starter just for code security reasons. And if you want to collect diverse data, taking robots off campus is, is valuable.”
Chelsea Finn Mar 20, 2025 ▶ 28:04
Insight
Chelsea Finn: Observational video alone cannot train robot foundation models
“I think that data can have a lot of value, but I think that by itself, it won't get you very far and I think that there's actually some really nice analogies you can make where for example, if you watch, like, an Olympic swimmer, swimmer race even if you had t…”
Chelsea Finn Mar 20, 2025 ▶ 29:37
Prediction Not checkable as stated
Chelsea Finn: Autonomous RL experience will play a huge role in robotics
“And then I also think that autonomous experience will play a huge role, just like we've seen in language models. After you get an initial language model, if you can use reinforcement learning to have the robot, the language model bootstrap on its own experienc…”
Chelsea Finn Mar 20, 2025 ▶ 30:54
Prediction Not checkable as stated
Finn predicts a 'Cambrian explosion' of diverse robot hardware platforms
“I don't know exactly, but I think that my bet would be on something where there's actually a A really wide range of different robot platforms. I think Sergei my co-founder likes to call it a Cambrian explosion of different robot hardware types and so forth. O…”
Chelsea Finn Mar 20, 2025 ▶ 32:19
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.