Spatial Intelligence
topic on 5 shows · 18 statements across 8 episodes
the Y Combinator Startup Podcast
Latent Space
Lenny's Podcast
No Priors
the a16z Podcast
18 statements about Spatial Intelligence, every show
Fei-Fei Li: Spatial intelligence requires generation, reasoning, and interaction
“Spatial intelligence eventually must enable us to both generate what the space is, reason within it, and being able to edit and interact within it.”
Fei-Fei Li: Atlas grounds generative visual models in spatial context
“So, so in the, on the path to spatial intelligence, generating pixels is definitely a early step, which we have seen with What you call it, gazillions of models. But generating pixels that are truly spatially contextualized and grounded is absolutely another m…”
Li: World Labs founded without a scaling law for spatial intelligence
“There was no scaling law of spatial intelligence.”
Li: Spatial intelligence requires closing the loop between seeing and interacting
“Intelligence is not sitting there stuck and just seeing something or interpreting something when it comes to space and physical space, right? It's really this closing the loop between seeing and experiencing an interaction.”
Manning: Video Models Lack Genuine 3D Spatial Understanding and Causality
“The reality is that although the visuals do look fantastic, those visuals actually aren't accompanied by an understanding of the three-d world, understanding how objects can move, what the consequences of different actions are, and that's what's really needed …”
Li: World Labs was founded to build spatial intelligence beyond LLMs
“I think around, you know, more than two years ago, for sure, I think both independently, both of us have been looking at the development of the large models and thinking about What's beyond language models and this idea of building world models, spatial intell…”
Li: Evolution spent 540M years on spatial intelligence versus 0.5M on language
“And in nature, you know, it took five hundred, forty million years to optimize perception and spatial intelligence and language in the most Generous estimation of language development is probably half a million years.”
Li: World models and spatial intelligence are essential for embodied AI
“I think world modeling and spatial intelligence Is a key missing piece of embodied AI.”
Li: Spatial intelligence and world models are as important as LLMs
“We believe that spatial intelligence and world modeling is As important, if not more, to language models and complementary to language models.”
Fei-Fei Li: AGI will not be complete without spatial intelligence
“To me, AGI will not be complete. Without spatial intelligence.”
Fei-Fei Li: 3D spatial intelligence is combinatorially harder than language models
“The real world is three D, and if you add time, it's four D, but just, let's just confine ourselves within space. It's fundamentally three D. So that by itself is a much more combinatorially harder problem.”
Fei-Fei Li calls solving spatial intelligence a problem 'bordering delusional'
“My entire career is going after problems that are just so hard, bordering delusional, and I think this is this is this is the delusional problem.”
Fei-Fei Li: Hardware and software convergence will eventually enable the metaverse
“I'm actually really, really excited by metaverse. I know so many people are kind of still like it's still not working. I know it's still not working. That's why I'm excited because I think the convergence of hardware and software will be coming. So that's also…”
Dr. Fei-Fei Li: AI is not complete without spatial intelligence
“AI is not complete without spatial intelligence because The humans interact in in three D worlds and in the digital world, we need all kinds of interaction”
Li: Spatial intelligence is a critical component of intelligence
“Space, the three D space, the space out there, the space in your mind's eye, the spatial intelligence is a part of a critical part of intelligence.”
Fei-Fei Li: All embodied robotics require 3D spatial intelligence
“Robotics to me is any body machines. It's not just humanoids or cars. There's so much in between, but all of them have to somehow figured out, ah, the three D space it lives in, have to be trained, ah, to understand the three D space and have to do things, som…”
Justin Johnson: Multimodal LLMs shoehorn visual data into 1D token sequences
“And now the multimodal LLMs that we're seeing now, you kind of end up shoehorning the other modalities into this underlying representation of a one D sequence of tokens. Now when we move to spatial intelligence, it's kind of going the other way. Where we're sa…”