Everything Justin Johnson said on any show that made the record, most notable first. Each card names its show and opens the statement there.
Johnson: No base model has used new view prediction before Atlas
“And this is a really fundamental primitive that we think is super exciting, a super new primitive For base models that no one's ever done before, right?”
Johnson: Scaling spatial world models is primarily bottlenecked by training compute
“I think we're basically at the beginning, and we're basically limited by compute at this point, right? Like, data is very important, as Fei-Fei likes to point out, but, like, everything has a bottleneck, and I think the main bottleneck on continuing to scale t…”
Johnson: Generative new view prediction in Atlas is AI-complete
“But I think that's something we're kind of realizing, and Ben was talking about this earlier today, is, like, new view prediction, this primitive that we have in Atlas, especially generative new view prediction, this is also AI complete.”
Johnson: Nvidia Blackwell offers roughly same performance per watt as Hopper
“Like, if you look at the numbers, like, even going from Hopper to Blackwell, like, the performance per watt is about the same. They mostly make the number of transistors go up, and they make the chip size go up, and they make the power usage go up. But even fr…”
Johnson: Pixels offer a more lossless world representation than tokenized text
“And then like you actually lose something if you translate to this like purely tokenized representations that we use in LLMs, right? Like you lose the font, you lose the line breaks, you lose sort of the two D arrangement on the page. And for a lot of cases, f…”
Justin Johnson: Native 3D AI representations will outperform 2D video generation
“Modeling the two D projections of a dynamic three D world is, is a function that probably can be modeled, but by putting a three D representation into the heart of a model, there's just going to be a better fit between the kind of representation that the model…”
Justin Johnson: Seamless mixed reality will deprecate phones, TVs, and monitors
“If you've got the ability to seamlessly blend virtual content with the physical world, it kind of deprecates the need for all of those.”
Johnson: World Labs' Atlas model generates, reconstructs, and simulates the world
“Atlas is our new next generation world model. It has three basic things. It can generate, reconstruct, and simulate the world.”
Johnson: Atlas achieves bullet-time effects using only three iPhones
“But now with Atlas, we can do this with just as few, through three cameras. So like no studio capture, no green screen, no expensive calibration. We can literally stick like three cameras on tri, three iPhones on tripods use these to take sort of a video of so…”
Justin Johnson: Multimodal LLMs shoehorn visual data into 1D token sequences
“And now the multimodal LLMs that we're seeing now, you kind of end up shoehorning the other modalities into this underlying representation of a one D sequence of tokens. Now when we move to spatial intelligence, it's kind of going the other way. Where we're sa…”
Johnson: Atlas reconstructs 3D spaces from up to 100 frames
“You can input one or multiple up to a hundred frames that are views of the real world, and use those to reconstruct the real world, and that reconstruction can take the case either of a novel, a video flying through the space, or an explicit three-D reconstruc…”
Johnson: Feeding camera poses natively during pre-training is unprecedented
“It also works on camera poses as a native input to the model, which I don't think anyone's ever done at the pre-training phase before.”
Johnson: World Labs' models improved significantly with each increase in compute
“Each time we made the model bigger and each time we trained it for longer, each time we put it on more chips, like it got significantly better.”
Johnson: Academic labs can no longer train state-of-the-art AI on few GPUs
“Like five or 10 years ago, you really could train state-of-the-art models in the lab even with just a couple of GPUs. But, you know, because that technology was so successful and scaled up so much, then you can't train state-of-the-art models with a couple of …”
Johnson: Gaussian splats render in real time on nearly any client device
“Gaussian splats are really cool because you can render them in real time really efficiently. So you can render on your iPhone, render, render everything. And that's how we get that sort of precise camera control because The splats can be rendered real time on …”
Johnson: Physical theory-building stems from interactive falsification, not model modality
“Because we're constantly interacting with the world, we're constantly having to build theories about what's happening in the world around us, and then falsify or add evidence to those theories. And I think that that kind of process writ large and scaled up is …”
Johnson: Transformers are natively models of sets, not sequences
“Transformers are actually not a model of sequences. A transformer is natively a model of sets.”
Justin Johnson: AI is shifting from analyzing web data to sensor data
“The previous decade had mostly been about understanding data that already exists. But the next decade was going to be about understanding new data.”
Justin Johnson: Ben Mildenhall's 2020 NeRF paper ignited 3D computer vision
“In 2020, you asked about breakthrough moments. There was a really big breakthrough moment from our co-founder Ben Mildenhall at the time with his paper, NERF neural radiance fields. And that was a very simple, very clear way of backing out three D structure fr…”
Justin Johnson: VR headsets like Vision Pro lack mass market readiness
“But I think the reality is it's just not there yet as a platform for mass market appeal.”
Johnson: AI compute per model has scaled one million-fold since 2012
“And if you think about, you know, AlexNet required this jump from CPUs to GPUs, but even from AlexNet to today, we're getting about a thousand times more performance per card than we had in AlexNet days. And now it's common to train models, not just on one GPU…”
Justin Johnson: AlexNet's 6-day training takes under 5 minutes on GB200
“So I ran the numbers last night, like that two week training run, that of six days on two GTX five eighties, if you scale, it comes out to just under five minutes on a single GB 200.”
Justin Johnson: Creating interactive 3D worlds currently costs hundreds of millions
“Because we already have the ability to create virtual interactive worlds but it costs hundreds and hundreds of millions of dollars and a ton of development time.”
Justin Johnson: AlexNet trained for six days on two consumer GPUs
“That AlexNet was a sixty million parameter deep neural network and it was trained for six days on two GTX five eighties, which was the top consumer card at the time, which came out in 2010.”