The Ledger, every show
Every statement that passed quotation and attribution checks, across all 44 shows. Pick shows below, then mix any filter with any other.
shows 




every show 44 of 44
Luan: Base AI model performance requires doubling compute for consistent gains
“So put another way for just scaling up a base language model you need to double the amount of compute for that language model for it to be predictably consistently smarter.”
Multimodal models will completely supplant text-only large language models
“I actually think like it's really clear today. Multimodal models are the default foundation model, right? It's just going to supplant LLMs. Like why did you just train a giant multimodal model?”
Luan: ChatGPT was just GPT-3 with instruction tuning released a year later
“ChatGPT was really just GPT-III with instruction. It was basically like more chat tuning, but GPT-III API came out, I think, over a year before ChatGPT did, but only developers could play with it.”
Luan: Scaling GPT-2 without architectural changes unlocked three-digit arithmetic
“When we were training GPT-II, we trained GPT-II in various different sizes. And at the smallest size, the model was just, like, unable to do three-digit arithmetic. But as the models got bigger and bigger and bigger, we didn't change anything else. We just had…”
Luan: Every major cloud and LLM provider is developing custom chips
“Every one of the major clouds is working on And every major LLM provider is working on a strategy to have their in-house chips because that way they have better margins.”
Dextro was the first to offer automated video analysis as a service
“We were the first company to figure out how to get this level of analysis of what's happening in videos as a service.”
AlexNet halved state-of-the-art visual recognition error rates in a single year
“Krzyzewski's work on ILS VRC, which is the ImageNet Large Scale Visual Recognition Challenge, where in one year, essentially they blew away the previous state of the art with a deep convolutional neural network by about, like, half of the final error rate on t…”