DeepSpeed, every mention

7 scenes · ← back to DeepSpeed

tap a year for its mentions
00326320242025episodesmentions
02320242025episodes it came up in
0011.52320242025episodesmentions per episode

every year anyone Josh Albrecht 2Jack Morris 2Eugene Cheah 2Sarah Chieng 1

Verbatim, from the transcripts: the passages where DeepSpeed comes up

loading…

How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony Nov 3, 2025 · 1 mention

  • ▶ 29:54 unnamed speaker Lead development on a model training framework called GPT NeoX, which is like a model, a Megatron deep speed style framework for pre-training models on HPC systems.

Building Jamba 3B: the tiny Hybrid Transformer State Space Reasoning Model - Barak Lenz, CTO of AI21 Oct 11, 2025 · 1 mention

  • ▶ 20:33 unnamed speaker Like the DeepSpeed team from Microsoft is like, uh, is doing good stuff.

Information Theory for Language Models: Jack Morris Jul 2, 2025 · 2 mentions

  • ▶ 9:49 Jack Morris And it's not like they're learning how to do like multi-node distributed FSTP training, like with whatever deep speed, you have to learn that from the internet and from other people. 2 times in the scene

[Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras) Dec 7, 2024 · 1 mention

  • ▶ 42:07 Sarah Chieng Um, so I actually didn't dive too deep into these techniques, um, so I just kind of gave a couple of, uh, techniques and what different companies are doing, um, just so you're aware of it and feel free to deep dive further, and so when…

[LLM Paper Club] Llama 3.1 Paper: The Llama Family of Models Jul 29, 2024 · 2 mentions

  • ▶ 16:33 Eugene Cheah The reason why people try to avoid pipeline parallelism at all costs, and they use like DeepSpeedTree, for example, where the weights are shuddered around all the other GPUs, is that 2 times in the scene

State of the Art: Training 70B LLMs on 10,000 H100 clusters Jun 25, 2024 · 3 mentions

  • ▶ 32:23 unnamed speaker I noticed also in the piece that you mentioned FSTP with zero three, um, actually when we met, um, I went to iClear and, uh, Guanhua from the DSP team was, was there presenting zero plus plus.
  • ▶ 32:57 Josh Albrecht We use stuff from probably deep speed. 2 times in the scene
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.