Justin Johnson, co-founder of World Labs, discusses the structural limitations of using 1D language model sequence architectures for spatial and visual AI reasoning.
Prediction Not checkable as stated
Justin Johnson: Native 3D AI representations will outperform 2D video generation
“Modeling the two D projections of a dynamic three D world is, is a function that probably can be modeled, but by putting a three D representation into the heart of a model, there's just going to be a better fit between the kind of representation that the model…”
Prediction Not checkable as stated
Justin Johnson: Seamless mixed reality will deprecate phones, TVs, and monitors
“If you've got the ability to seamlessly blend virtual content with the physical world, it kind of deprecates the need for all of those.”
Assertion Not yet assessed · timeframe Sep 2026
Johnson: No base model has used new view prediction before Atlas
“And this is a really fundamental primitive that we think is super exciting, a super new primitive For base models that no one's ever done before, right?”
Opinion
Johnson: Scaling spatial world models is primarily bottlenecked by training compute
“I think we're basically at the beginning, and we're basically limited by compute at this point, right? Like, data is very important, as Fei-Fei likes to point out, but, like, everything has a bottleneck, and I think the main bottleneck on continuing to scale t…”
Opinion
Johnson: Generative new view prediction in Atlas is AI-complete
“But I think that's something we're kind of realizing, and Ben was talking about this earlier today, is, like, new view prediction, this primitive that we have in Atlas, especially generative new view prediction, this is also AI complete.”
Assertion Partly supported
Johnson: World Labs' Atlas model generates, reconstructs, and simulates the world
“Atlas is our new next generation world model. It has three basic things. It can generate, reconstruct, and simulate the world.”