World Labs co-founder Justin Johnson clarifies architectural misconceptions about transformers versus recurrent neural networks.
Assertion Contradicted
Johnson: Nvidia Blackwell offers roughly same performance per watt as Hopper
“Like, if you look at the numbers, like, even going from Hopper to Blackwell, like, the performance per watt is about the same. They mostly make the number of transistors go up, and they make the chip size go up, and they make the power usage go up. But even fr…”
Insight
Johnson: Pixels offer a more lossless world representation than tokenized text
“And then like you actually lose something if you translate to this like purely tokenized representations that we use in LLMs, right? Like you lose the font, you lose the line breaks, you lose sort of the two D arrangement on the page. And for a lot of cases, f…”
Assertion Not checkable as stated
Johnson: Academic labs can no longer train state-of-the-art AI on few GPUs
“Like five or 10 years ago, you really could train state-of-the-art models in the lab even with just a couple of GPUs. But, you know, because that technology was so successful and scaled up so much, then you can't train state-of-the-art models with a couple of …”
Assertion Supported
Johnson: Gaussian splats render in real time on nearly any client device
“Gaussian splats are really cool because you can render them in real time really efficiently. So you can render on your iPhone, render, render everything. And that's how we get that sort of precise camera control because The splats can be rendered real time on …”
Insight
Johnson: Physical theory-building stems from interactive falsification, not model modality
“Because we're constantly interacting with the world, we're constantly having to build theories about what's happening in the world around us, and then falsify or add evidence to those theories. And I think that that kind of process writ large and scaled up is …”
Assertion Supported
Johnson: AI compute per model has scaled one million-fold since 2012
“And if you think about, you know, AlexNet required this jump from CPUs to GPUs, but even from AlexNet to today, we're getting about a thousand times more performance per card than we had in AlexNet days. And now it's common to train models, not just on one GPU…”