World Labs co-founder Justin Johnson explains the underlying architecture of Atlas, the company's 3D world model, on The a16z Podcast.
Prediction Not checkable as stated
Justin Johnson: Native 3D AI representations will outperform 2D video generation
“Modeling the two D projections of a dynamic three D world is, is a function that probably can be modeled, but by putting a three D representation into the heart of a model, there's just going to be a better fit between the kind of representation that the model…”
Prediction Not checkable as stated
Justin Johnson: Seamless mixed reality will deprecate phones, TVs, and monitors
“If you've got the ability to seamlessly blend virtual content with the physical world, it kind of deprecates the need for all of those.”
Opinion
Johnson: Scaling spatial world models is primarily bottlenecked by training compute
“I think we're basically at the beginning, and we're basically limited by compute at this point, right? Like, data is very important, as Fei-Fei likes to point out, but, like, everything has a bottleneck, and I think the main bottleneck on continuing to scale t…”
Opinion
Johnson: Generative new view prediction in Atlas is AI-complete
“But I think that's something we're kind of realizing, and Ben was talking about this earlier today, is, like, new view prediction, this primitive that we have in Atlas, especially generative new view prediction, this is also AI complete.”
Insight
Justin Johnson: Multimodal LLMs shoehorn visual data into 1D token sequences
“And now the multimodal LLMs that we're seeing now, you kind of end up shoehorning the other modalities into this underlying representation of a one D sequence of tokens. Now when we move to spatial intelligence, it's kind of going the other way. Where we're sa…”
Assertion Partly supported
Johnson: World Labs' Atlas model generates, reconstructs, and simulates the world
“Atlas is our new next generation world model. It has three basic things. It can generate, reconstruct, and simulate the world.”