Training
topic on 11 shows · 22 statements across 20 episodes
BG2 Pod
Cheeky Pint
Latent Space
No Priors
Catalyst
the MAD Podcast
the a16z Podcast
Big Technology
All-In
TBPN
20VC
22 statements about Training, every show
Optimal inference parallelism cannot be mathematically calculated; it must be auto-tuned
“And with training, it's more of like a math, like you can run the math and see the flops and maximize it. With inference, it's more of like an auto-tuning, like GPU kernel auto-tuning... You shadow the same traffic, like real traffic, and you just see which co…”
Katti: Modern AI model training consists heavily of inference workloads
“We don't like to make a distinction between Training and infants, because a lot of training is now infants. So when we train a new model, we are generating synthetic data, for example. That's inference. When we train a new model, we are doing post-train, and t…”
Pope: MatX will sell AI inference chips first due to lower risk
“Our product is both training and inference, but I think the first sales will be an inference. That's mostly just a market effect where It's easier to buy, like, it's not as big of a risk to go to buy an inference cluster than as a training cluster.”
Bishop: China manufactures capable inference chips but cannot produce training chips
“China has enough, like they can actually make, looks like decent inference chips. They just can't make the chips they need for training, right?”
Altman: Society will deem AI training fair use but create IP licensing models
“So like, you'll see this continue to move, but forced guests from the position we're in today, I would say that society decides training is fair use, but There's a new model for generating content in the style of or with the IPF or something else.”
Feldman: Vastly more people do AI inference than AI training
“To move people off GPUs in inference, and the number of people doing inference is vastly higher than the number of people doing training.”
Kupor: OPM employees granted two hours monthly for AI training
“Everybody, you know, you're entitled to two hours a month of training, using all these free resources. You can take time out of work, just clear it with your manager.”
Guthrie: Training-only datacenters cannot serve inference without global networking
“If you are, for example, building one large data center that only does training and it's not connected to a wide area network around the world, that's close to the users, it's hard to use that same infrastructure For inferencing because you can't go faster tha…”
Ross: AI Training and Inference Form a Virtuous Hardware Demand Cycle
“The more inference you have, as mentioned before, the more you need to train the model to optimize for the inference. And the more training you have the more inference you want to deploy to optimize for the cost of that training, to amortize the cost of the tr…”
Kohli: AlphaEvolve has successfully made AI model training more computationally efficient
“What Alpha Evolve has been able to do is basically make training more efficient.”
Long: Bad security training signals that company security is unimportant
“When the training is kind of a joke at the company, I think it also sends a signal to all of the employees that the company's security posture is also not very important.”
Kimber: AI inference will ultimately draw far more power than training
“I think now what we're seeing is that inference in the aggregate is actually probably a much larger draw than the training in the long run.”
Monday Afternoon Is Optimal for Pipeline Generation Training
“Monday afternoons is a great time to do training, especially training that's related to PG skills.”
Bernhardsson: Modal Is Expanding into Bursty Experimental AI Training
“Traditionally, most of modal has always been inference. Like that's been our main use case, but we're really interested also in training. So in particular, like probably focused more on these like shorter, like very bursty sort of experimental training runs, n…”
AI inference matters more than training because it scales with global population
“Our prediction is for those kind of applications, the inference is much more important than training. Because inference scale is proportional to the upliminal world population. And training. Training scale is proportional to the number of researchers.”
Huang: NVIDIA's moat in inference will be greater than in training
“And I'm sure I said it would be greater.”
Chamath predicts AI inference market will be 100x larger than training
“AI is really two markets, training and inference is going to be a hundred times bigger than training.”
Chamath asserts Nvidia hardware is miscast for AI inference
“NVIDIA is really good at training and very miscast at inference.”
Srivastava: Inter-rack networking matters less for AI inference than training
“Even the GPU clusters themselves, like, you know, the full training networking is a very, very important Piece to have networking on the racks themselves with inference and matters a little less because you're doing a little bit more on individual GPUs and les…”
Srivastava: AI inference demands strict uptime, while training tolerates node terminations
“Resiliency and reliability matters a lot more. You know, downtime is unacceptable from an input perspective. Nodes get terminated all the time from a training perspective.”
Polosukhin: Inference demands vastly more aggregate compute than AI model training
“I think an inference is really interesting because we do need so much more compute for inference than we need for training, right? Like it's a very interesting like economy of scale. You train once, like Lama trained once and then everybody runs it everywhere.”