The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Fei-Fei Li no published score: only 12 usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.0/5 from 12 raw tape exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
12exchanges match
12on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q and managing builds and scale of manufacturing, you're going to have many fewer form factors. And the other argument is the economic value of specialization is very high. And therefore there'll be, you know, thousands and thousands of different form factors as we move to sort of a robot driven future. Do you have a point of view on sort of where we're likely to land between those two viewpoints?

A I think we're going to gradient descending to optimization of productivity and efficiency. My hypothesis is that the requirements of different tasks are so vast that having very few form or, or sticking with one form is. Energy, energy inefficient. And a lot of tasks can be done and should be done by much more energy efficient form factors. Just an extreme and, and trivial example. If we put robots underwater, they should not be in the shape of humans. They better be in the shape of fish, right? Just think about energy efficiency. And the same with flying. I don't think human form is, uh, our airplanes are becoming more and more robots. Um, so I, I, I do think there's gonna be diversity.

AI assessment note: “I do think there's gonna be diversity.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q to date? And obviously there's still a lot of career to come, but I'm just sort of curious. I mean, obviously there's a lot of things that you did in terms of Sort of, uh, image and visual recognition-related systems and all sorts, but I'm just sort of curious, like, when you think, think of the last 20 years, what stands out the most, just given everything that you've done?

A Oh, thank you for asking that question. Of course, ImageNet is one of those, uh, ImageNet consists of multiple moments from the early struggles and being told I will not get tenure to, um, To actually realizing Amazon Mechanical Turk comes to rescue to the moment of Alex Nett winning, and also to a couple of years ago, I was at an event in Toronto with Jeff Hinton, and he said publicly, like, how that was so defining, and he, he was almost a little bit, um, apologetic, that image that was not As recognized as, uh, as neural networks. So that journey is very validating. And for scientists, the validation is not about recognition or awards. It's that you made a difference. Like that conjecture that no one believed in, that hypothesis that no one believed in, we were able to make it happen. So that's one thread.

AI assessment note: “Of course, ImageNet is one of those, uh, ImageNet consists of multiple moments”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q to date? And obviously there's still a lot of career to come, but I'm just sort of curious. I mean, obviously there's a lot of things that you did in terms of Sort of, uh, image and visual recognition-related systems and all sorts, but I'm just sort of curious, like, when you think, think of the last 20 years, what stands out the most, just given everything that you've done?

A Oh, thank you for asking that question. Of course, ImageNet is one of those, uh, ImageNet consists of multiple moments from the early struggles and being told I will not get tenure to, um, To actually realizing Amazon Mechanical Turk comes to rescue to the moment of Alex Nett winning, and also to a couple of years ago, I was at an event in Toronto with Jeff Hinton, and he said publicly, like, how that was so defining, and he, he was almost a little bit, um, apologetic, that image that was not As recognized as, uh, as neural networks. So that journey is very validating. And for scientists, the validation is not about recognition or awards. It's that you made a difference. Like that conjecture that no one believed in, that hypothesis that no one believed in, we were able to make it happen. So that's one thread.

AI assessment note: “Of course, ImageNet is one of those, uh, ImageNet consists of multiple moments”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Does that imply you have a particular, um, point of view, um, from a, like, a neuroscience perspective of, like, you know, how fundamental visual, you've, I mean, you've always been a leader in, um, uh, computer vision, right? But in how important visual intelligence is versus, let's say, like, Large language models and textual intelligence.

A I actually do. I think from a neural and cognitive science point of view that spatial intelligence is a really hard problem that evolution has to solve for animals. And what's really interesting is I think animals have solved it to an extent, but not fully solved it. It's one of the hardest problem because, um, what is the problem animal has to solve? Animals Have to evolve the capability of collecting lights in something which we call eyes mostly. And then with that collection of eyes, it has to reconstruct a three D world in their mind somehow so that they can navigate and they can do things. And of course they can interact for humans where the most Capable animal in terms of manipulation. We can do a lot of things, and all this is spatial intelligence. To me, that's, um, that's just rooted in, in our intelligence. What is interesting is, it's not a fully solved problem, even in animals. We, uh, for example, uh, for humans, right, um, if I ask you to close your eyes right now and draw out or Or, or build a three D model of the environment around you. It's not that easy. We don't have that much capability to generate extremely complicated three D model till we get trained. You know, there are some of us, whether they're architects or, or designers, or just people with a lot of training and a lot of talent. And that's, that's, uh, That's a hard thing to do. And imagine you do i…

AI assessment note: “I actually do. I think from a neural and cognitive science point of view”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q and managing builds and scale of manufacturing, you're going to have many fewer form factors. And the other argument is the economic value of specialization is very high. And therefore there'll be, you know, thousands and thousands of different form factors as we move to sort of a robot driven future. Do you have a point of view on sort of where we're likely to land between those two viewpoints?

A I think we're going to gradient descending to optimization of productivity and efficiency. My hypothesis is that the requirements of different tasks are so vast that having very few form or, or sticking with one form is. Energy, energy inefficient. And a lot of tasks can be done and should be done by much more energy efficient form factors. Just an extreme and, and trivial example. If we put robots underwater, they should not be in the shape of humans. They better be in the shape of fish, right? Just think about energy efficiency. And the same with flying. I don't think human form is, uh, our airplanes are becoming more and more robots. Um, so I, I, I do think there's gonna be diversity.

AI assessment note: “so I, I, I do think there's gonna be diversity.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q What are the biggest challenges in, um, I guess, trying to go down this path of, you know, designing and training world models? I imagine one is, like, you worked on images, you worked on video, but we, we have images, and we have video, and we don't have lots of, you know, three-D worlds, like, in, in a format I assume you're building.

A Yeah, data is absolutely a challenge. You're totally right about that. Um, you know, to create, uh, world models, three-D foundation models, uh, we, we require More and more sophisticated data engineering, data acquisition, data processing, and data, uh, synthesis. So, um, uh, I am envious of my, uh, NLP LLM colleagues that the data is so abundant on the internet, and we don't necessarily have that luxury. So that's definitely one, um, one challenge. Another one challenge is that, um, three D is, this is kind of, um, ironic, right? Every one of us use three D every day, like in so many settings. Basically you open your eye and, and, and the whole life that you experience is three D. Okay. Even when we type on the computer or stare at a screen all the time, yet it's still not as easy a form factor to deliver in the hands of people compared to language. The language is just so easy. And, uh, it's also a very active form of, it's not a passive consumption of viewing. Nobody wakes up and say, I'm just gonna sit here and watch three D, you know. So, um, that, uh, creates challenges for, for, for productization and how to do it in the right way.

AI assessment note: “data is absolutely a challenge. You're totally right about that.”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q What are the biggest challenges in, um, I guess, trying to go down this path of, you know, designing and training world models? I imagine one is, like, you worked on images, you worked on video, but we, we have images, and we have video, and we don't have lots of, you know, three-D worlds, like, in, in a format I assume you're building.

A Yeah, data is absolutely a challenge. You're totally right about that. Um, you know, to create, uh, world models, three-D foundation models, uh, we, we require More and more sophisticated data engineering, data acquisition, data processing, and data, uh, synthesis. So, um, uh, I am envious of my, uh, NLP LLM colleagues that the data is so abundant on the internet, and we don't necessarily have that luxury. So that's definitely one, um, one challenge. Another one challenge is that, um, three D is, this is kind of, um, ironic, right? Every one of us use three D every day, like in so many settings. Basically you open your eye and, and, and the whole life that you experience is three D. Okay. Even when we type on the computer or stare at a screen all the time, yet it's still not as easy a form factor to deliver in the hands of people compared to language. The language is just so easy. And, uh, it's also a very active form of, it's not a passive consumption of viewing. Nobody wakes up and say, I'm just gonna sit here and watch three D, you know. So, um, that, uh, creates challenges for, for, for productization and how to do it in the right way.

AI assessment note: “Yeah, data is absolutely a challenge. You're totally right about that.”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q can get from that today. Perhaps people are, do not see the future of, like, the quality and the physics that are going to be available to us. Um, and then there's, you know, close to embodied, like, different forms of tele-op and then, like, embodied data collection. Is that the hierarchy you have in your mind, or do you think people underestimate simulation and world models for the future?

A Yeah, great question. First of all, I, like you said, I do work in robotics, especially in my lab at Stanford. I have no doubt that humanity will move into an age where we cohabit with robots. And also the world, the world robot is not humanoid per se. Robots taking All kind of forms and shapes. Actually, a few years ago, my lab wrote a really fun paper about morphological intelligence is where the, the morphology of a, a, an agent actually can change by optimizing the tasks they're trying to achieve. So, so we should be a little more imaginative than just human, humanoids. Having said that, how to train robot Uh, you mentioned this whole data. Some people call it data pyramids or data cakes or whatever. I agree. I think it's gonna be a hybrid of, uh, many different forms of data. I also think, uh, simulation is underrated. It's, um, actually, it's not underrated by a lot of experts and people in the field. If you look at a lot of robotics companies, they are working on simulated, uh, simulation and synthetic data. I also think we have to be, um, also aware that unlike language models or even unlike, um, Spatial Intelligence Foundation models, robotics is a highly multimodal, um, um, uh, system that I think what is truly underappreciated, in my opinion, is haptics. Is there so much, especially if we want to do manipulation, not just Navigation. I think haptics data and the abil…

AI assessment note: “I agree. I think it's gonna be a hybrid of, uh, many different forms”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q Does that imply you have a particular, um, point of view, um, from a, like, a neuroscience perspective of, like, you know, how fundamental visual, you've, I mean, you've always been a leader in, um, uh, computer vision, right? But in how important visual intelligence is versus, let's say, like, Large language models and textual intelligence.

A I actually do. I think from a neural and cognitive science point of view that spatial intelligence is a really hard problem that evolution has to solve for animals. And what's really interesting is I think animals have solved it to an extent, but not fully solved it. It's one of the hardest problem because, um, what is the problem animal has to solve? Animals Have to evolve the capability of collecting lights in something which we call eyes mostly. And then with that collection of eyes, it has to reconstruct a three D world in their mind somehow so that they can navigate and they can do things. And of course they can interact for humans where the most Capable animal in terms of manipulation. We can do a lot of things, and all this is spatial intelligence. To me, that's, um, that's just rooted in, in our intelligence. What is interesting is, it's not a fully solved problem, even in animals. We, uh, for example, uh, for humans, right, um, if I ask you to close your eyes right now and draw out or Or, or build a three D model of the environment around you. It's not that easy. We don't have that much capability to generate extremely complicated three D model till we get trained. You know, there are some of us, whether they're architects or, or designers, or just people with a lot of training and a lot of talent. And that's, that's, uh, That's a hard thing to do. And imagine you do i…

AI assessment note: “I actually do. I think from a neural and cognitive science point of view”

Answered raw tape D 4 · C 4 · P 3 · Cm 3 3.60

Q How do you do that? Like, how do you identify if somebody has fearlessness in their background or in their thinking processes?

A It's in their background. You talk to them. You can sense someone is fearless. You know, you can sense what drives them. You know, you can sense the questions they ask. If they are If they start to asking you a lot of things about I don't know how to get this done. I mean, of course you have to ask those questions, because you want to get it done, but if, if you sense that it comes from the, the, the, the, the point of view of, um, being scared of solving that, then that's not fearlessness. But, um, but those fearless people, they are creative, they're ambitious, they, they, they, they can, they're not afraid of, um, The uncertainty or the unknown. And I really love that.

AI assessment note: “You talk to them. You can sense someone is fearless... the questions they ask.”

Answered raw tape D 3 · C 4 · P 4 · Cm 3 3.55

Q can get from that today. Perhaps people are, do not see the future of, like, the quality and the physics that are going to be available to us. Um, and then there's, you know, close to embodied, like, different forms of tele-op and then, like, embodied data collection. Is that the hierarchy you have in your mind, or do you think people underestimate simulation and world models for the future?

A Yeah, great question. First of all, I, like you said, I do work in robotics, especially in my lab at Stanford. I have no doubt that humanity will move into an age where we cohabit with robots. And also the world, the world robot is not humanoid per se. Robots taking All kind of forms and shapes. Actually, a few years ago, my lab wrote a really fun paper about morphological intelligence is where the, the morphology of a, a, an agent actually can change by optimizing the tasks they're trying to achieve. So, so we should be a little more imaginative than just human, humanoids. Having said that, how to train robot Uh, you mentioned this whole data. Some people call it data pyramids or data cakes or whatever. I agree. I think it's gonna be a hybrid of, uh, many different forms of data. I also think, uh, simulation is underrated. It's, um, actually, it's not underrated by a lot of experts and people in the field. If you look at a lot of robotics companies, they are working on simulated, uh, simulation and synthetic data. I also think we have to be, um, also aware that unlike language models or even unlike, um, Spatial Intelligence Foundation models, robotics is a highly multimodal, um, um, uh, system that I think what is truly underappreciated, in my opinion, is haptics. Is there so much, especially if we want to do manipulation, not just Navigation. I think haptics data and the abil…

AI assessment note: “I also think, uh, simulation is underrated. It's, um, actually, it's not underrated”

Answered raw tape D 4 · C 4 · P 2 · Cm 3 3.35

Q How do you do that? Like, how do you identify if somebody has fearlessness in their background or in their thinking processes?

A It's in their background. You talk to them. You can sense someone is fearless. You know, you can sense what drives them. You know, you can sense the questions they ask. If they are If they start to asking you a lot of things about I don't know how to get this done. I mean, of course you have to ask those questions, because you want to get it done, but if, if you sense that it comes from the, the, the, the, the point of view of, um, being scared of solving that, then that's not fearlessness. But, um, but those fearless people, they are creative, they're ambitious, they, they, they, they can, they're not afraid of, um, The uncertainty or the unknown. And I really love that.

AI assessment note: “You talk to them. You can sense someone is fearless. You know, you can sense the questions they ask.”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.