CUTLASS, every mention

5 scenes · ← back to CUTLASS

tap a year for its mentions
002243202320242025episodesmentions
023202320242025episodes it came up in
001.51.533202320242025episodesmentions per episode

every year anyone Tri Dao 3Yining Zhang 1Ethan He 1Chris Lattner 1

Verbatim, from the transcripts: the passages where CUTLASS comes up

loading…

How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony Nov 3, 2025 · 2 mentions

  • ▶ 10:27 unnamed speaker and then it's like cut less. 2 times in the scene

The Shape of Compute (Chris Lattner of Modular) Jun 13, 2025 · 1 mention

DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing) Jan 19, 2025 · 1 mention

  • ▶ 11:57 Yining Zhang To implement the kernel, or you should use something like a catalyst to implement that kernel.

[Paper Club] Upcycling Large Language Models into Mixture of Experts Oct 29, 2024 · 1 mention

  • ▶ 13:05 Ethan He Uh, we provide a interface to Catalyst Group Jam, where the, the Catalyst Group Jam group all of the, like, looping over experts and calculate Jam into a single operation.

FlashAttention-2: Making Transformers 800% faster AND exact Aug 3, 2023 · 3 mentions

  • ▶ 31:50 Tri Dao The, the NVIDIA Cutlass team, um, they release a new version of their, their library, which contains all these primitives to allow you to do, like, you know, matrix multiply or memory loading on, on GPU, uh, efficiently. 3 times in the scene
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.