Vision Language Action Models
topic on 3 shows · 4 statements across 4 episodes
the Y Combinator Startup Podcast
Invest Like the Best
TBPN
4 statements about Vision Language Action Models, every show
Patel: VLA models will probably not scale for robotics due to data inefficiency
“Robots, currently the robot models VLAs Vision Language Action Models, which is very popular right now, is probably Not going to be the thing that ultimately scales beyond. They are inefficient in data. And we can't scale the data for them fast enough.”
Levine: Chain-of-thought reasoning allows robots to handle edge cases
“So the way you get common sense is by essentially using chain of thought. So the robot enters a scene and instead of directly starting to move, it thinks about what it was asked to do. So if it was told to clean up the kitchen, looks at the scene and says, lik…”
Finn: World models hallucinate success when evaluating suboptimal actions
“You might train it on demonstration data of successful data of completing the task, and then evaluate it on to try to actually use it to evaluate actions that are not optimally completing the task, and then the world model will hallucinate a video of completin…”