DeepSpeed, every mention

8 scenes across 2 shows · ← back to DeepSpeed

tap a year for its mentions
003263202320242025episodesmentions
023202320242025episodes it came up in
0011.523202320242025episodesmentions per episode

Latent Space 10the MAD Podcast 1

every year every show Latent Space 10 the MAD Podcast 1

Verbatim, from the transcripts: passages where DeepSpeed comes up on Latent Space, the MAD Podcast

loading…

How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony Nov 3, 2025 · 1 mention

  • ▶ 29:54 unnamed speaker Lead development on a model training framework called GPT NeoX, which is like a model, a Megatron deep speed style framework for pre-training models on HPC systems.

Building Jamba 3B: the tiny Hybrid Transformer State Space Reasoning Model - Barak Lenz, CTO of AI21 Oct 11, 2025 · 1 mention

  • ▶ 20:33 unnamed speaker Like the DeepSpeed team from Microsoft is like, uh, is doing good stuff.

Information Theory for Language Models: Jack Morris Jul 2, 2025 · 2 mentions

  • ▶ 9:49 Jack Morris And it's not like they're learning how to do like multi-node distributed FSTP training, like with whatever deep speed, you have to learn that from the internet and from other people. 2 times in the scene

[Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras) Dec 7, 2024 · 1 mention

  • ▶ 42:07 Sarah Chieng Um, so I actually didn't dive too deep into these techniques, um, so I just kind of gave a couple of, uh, techniques and what different companies are doing, um, just so you're aware of it and feel free to deep dive further, and so when…

[LLM Paper Club] Llama 3.1 Paper: The Llama Family of Models Jul 29, 2024 · 2 mentions

  • ▶ 16:33 Eugene Cheah The reason why people try to avoid pipeline parallelism at all costs, and they use like DeepSpeedTree, for example, where the weights are shuddered around all the other GPUs, is that 2 times in the scene

State of the Art: Training 70B LLMs on 10,000 H100 clusters Jun 25, 2024 · 3 mentions

  • ▶ 32:23 unnamed speaker I noticed also in the piece that you mentioned FSTP with zero three, um, actually when we met, um, I went to iClear and, uh, Guanhua from the DSP team was, was there presenting zero plus plus.
  • ▶ 32:57 Josh Albrecht We use stuff from probably deep speed. 2 times in the scene

The Race to Build the Ultimate AI Programmer | Poolside CTO Eiso Kant Dec 20, 2023 · 1 mention

  • ▶ 35:53 Eiso Kant The other thing that surprises a lot of people about us, uh, technically, and again, multiple paths to, to the same goal is that we didn't take Megatron or DeepSpeed or kind of the big open source frameworks for, for training large…
Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.