Parallel Thread Execution, every mention
8 scenes across 5 shows · ← back to Parallel Thread Execution
Latent Space 5
the Y Combinator Startup Podcast 2
the a16z Podcast 1
All-In 1
TBPN 1
every year every show
Latent Space 5
the Y Combinator Startup Podcast 2
the a16z Podcast 1
All-In 1
TBPN 1
Verbatim, from the transcripts: passages where Parallel Thread Execution comes up on Latent Space, the Y Combinator Startup Podcast, the a16z Podcast, All-In, TBPN
Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club · Y Combinator
- ▶ 10:48 Stuart (Stu) They produce kernels that are sometimes slower than non-overlapped baselines, or to use low-level primitives like OS inter-process communication calls or, uh, PTX assembly, and they often require reverse-engineered understanding of the…
- ▶ 33:43 unnamed speaker And sort of, like, at the sort of, like, lowest end, you have, like, the, what the cracked people do, which is, like, write PTX or inline SAS.
How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
- ▶ 10:44 Quentin Anthony So from the very, from going from bottom up, you have, um, PTX, so basically GPU assembly on video world and AMD world that's AMD GCN. 3 times in the scene
- ▶ 13:46 Quentin Anthony So when I try and get a model to do something in CUDA or PTX or something, as long, if it's not dead basic, it's bad really fast.
Dylan Patel on the AI Chip Race - NVIDIA, Intel & the US Government vs. China
- ▶ 1:04:17 Dylan Patel You know, when you're running inference, you're either, um, you know, using Cutlass or stamping out your own PTX,
Information Theory for Language Models: Jack Morris
- ▶ 13:10 Shawn Wang And like, we actually are faster than like the native, like, uh, sometimes the PTX implementation.
AI Czar David Sacks Explains the DeepSeek Freak Out
- ▶ 10:40 Chamath Palihapitiya And these guys worked totally around CUDA, and they did something called PTX, which goes right to the bare metal, and it's controllable, and it's effectively like writing assembly.
What Open AI Got WRONG
- ▶ 16:35 unnamed speaker There is an opportunity cost to research your time and there, and if they're spending it, micro optimizing PTX to make best