Jul 29, 2026 · 1h 16m · y-combinator

Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club · Y Combinator

Misha Smulyanski · 16m spoken Stuart (Stu) · 13m spoken Francois Chaubard · 7m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

This Y Combinator Paper Club session features expert presentations on the future of AI systems engineering, covering topics such as multi-GPU kernel optimization, intelligence-per-watt efficiency in local models, AI-generated systems code, heterogeneous inference architectures, and GPU-accelerated simulation engines.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The partners as informed peer 0.0 Guest teaching 0.0 Guest disagreement 0.2 The partners pushing back 0.0
05100:0020:0040:001:00:000:00–3:29 · The partners as informed peer 0/10 Title Sequence and Paper Abstract Overview Francois Chaubard opens the YC Paper Club session solo, introducing the theme of chip and kernel specialization.3:29–5:40 · The partners as informed peer 0/10 Chip and Data Center Specialization Trends Francois continues his opening monologue discussing TPU specialization, batch size one inference latency, and data center training dynamics.5:40–11:32 · The partners as informed peer 0/10 Upcoming Schedule and Introduction of First Speaker Francois covers upcoming schedule logistics, historical anecdotes about YC and Anybots, and introduces the first speaker Stuart.11:32–14:19 · The partners as informed peer 0/10 GPU Architecture and Memory Hierarchy Fundamentals Stuart delivers a solo technical talk explaining GPU architecture fundamentals, SMs, memory hierarchy, and NVLink.14:19–19:04 · The partners as informed peer 0/10 Three Key Trade-Offs in Multi-GPU Kernel Design Stuart breaks down the three multi-GPU kernel design trade-offs: transfer mechanisms, scheduling strategies, and design overheads.19:04–21:27 · The partners as informed peer 0/10 ParallelKittens Architecture, Evaluation, and Adoption Stuart outlines the ParallelKittens programming framework and highlights real-world adoption at Cursor and Together AI.21:27–24:38 · The partners as informed peer 0/10 Jon Saad-Falcon on the Mainframe Era vs. PC AI Era Jon Saad-Falcon begins his presentation by drawing historical parallels between the mainframe computing era and current centralized cloud AI data centers.24:38–27:51 · The partners as informed peer 0/10 Intelligence Per Watt Metric and Empirical Findings Jon details the intelligence per watt and per joule metrics and presents empirical findings across various models and accelerators.27:51–31:05 · The partners as informed peer 0/10 Query Routing, Consumer Hardware Gaps, and OpenJarvis Jon discusses edge routing feasibility, consumer hardware investment gaps, and the OpenJarvis project.31:05–34:17 · The partners as informed peer 0/10 Mark Saroufim on AI vs. Systems Researchers and GPU Languages Mark Saroufim humorously vents about the friction between systems researchers who desire static computation and AI researchers who want dynamic control flow.34:17–36:20 · The partners as informed peer 0/10 GPU Kernel Competitions & Leaderboard Realities Mark discusses GPU kernel leaderboards on KernelBot and the unexpected surge of non-expert contributors placing at the top using LLMs.36:20–39:03 · The partners as informed peer 0/10 Kernel Evaluation Frameworks & Reward Hacking Examples Mark explains how evaluation suites work and dismisses the naive assumption that verifiable rewards make kernel evaluation simple by showing zero-mean vector hacks.39:03–43:13 · The partners as informed peer 0/10 Superhuman Reward Hacks & The Volkswagen Analogy Mark details sophisticated reward hacks, comparing an AI's behavior under eval testing to the Volkswagen Dieselgate emissions cheat.43:13–47:04 · The partners as informed peer 0/10 GPU Memory Hierarchy & The PyTorch Flywheel Mark walks through the QR decomposition benchmark, the multi-thousand line generated kernels, and the 9-year evolutionary flywheel of PyTorch correctness.47:04–49:45 · The partners as informed peer 0/10 Why AI Inference Needs Heterogeneous Hardware Misha Smulyanski introduces Marlowe and sets out first-principles arguments for why AI inference demands heterogeneous infrastructure.49:45–52:59 · The partners as informed peer 0/10 Prefill vs. Decode & Arithmetic Intensity Misha uses the roofline model and arithmetic intensity to contrast compute-bound prefill phases with bandwidth-bound decode phases.52:59–56:27 · The partners as informed peer 0/10 Workload Diversity & SRAM Accelerators Misha categorizes workload diversity across LLM use cases and reviews on-die SRAM accelerators for memory-bound matrix-vector operations.56:27–1:03:08 · The partners as informed peer 0/10 Heterogeneous Architectures: PDD, AFD, & Speculative Decoding Misha outlines heterogeneous architectures including prefill/decode disaggregation, attention vs MLP disaggregation, and offloaded speculative decoding.1:03:08–1:06:31 · The partners as informed peer 0/10 Full-Stack Heterogeneous Inference Challenges & Wrap-up Misha concludes by highlighting data center power density, networking topology, and performance modeling challenges in heterogeneous co-design.1:06:31–1:10:50 · The partners as informed peer 0/10 Entity Component System (ECS) Architecture for GPUs Brennan introduces his Stanford research on running high-throughput batched game engine simulations entirely on GPUs using Entity Component System architectures.1:10:50–1:12:50 · The partners as informed peer 0/10 ECS Task Graph Execution and Low-Level GPU Implementation Details Brennan explains the GPU task graph execution, garbage-collection memory management, and persistent mega-kernels achieving high SM utilization.1:12:50–1:15:35 · The partners as informed peer 0/10 Benchmark Environments, End-to-End RL Training, and Speedup Metrics Brennan concludes with benchmark results showing over 100x speedups over CPU baselines and advocates for high-level Python-esque GPU scripting abstractions.0:00–3:29 · Guest teaching 0/10 Title Sequence and Paper Abstract Overview Francois Chaubard opens the YC Paper Club session solo, introducing the theme of chip and kernel specialization.3:29–5:40 · Guest teaching 0/10 Chip and Data Center Specialization Trends Francois continues his opening monologue discussing TPU specialization, batch size one inference latency, and data center training dynamics.5:40–11:32 · Guest teaching 0/10 Upcoming Schedule and Introduction of First Speaker Francois covers upcoming schedule logistics, historical anecdotes about YC and Anybots, and introduces the first speaker Stuart.11:32–14:19 · Guest teaching 0/10 GPU Architecture and Memory Hierarchy Fundamentals Stuart delivers a solo technical talk explaining GPU architecture fundamentals, SMs, memory hierarchy, and NVLink.14:19–19:04 · Guest teaching 0/10 Three Key Trade-Offs in Multi-GPU Kernel Design Stuart breaks down the three multi-GPU kernel design trade-offs: transfer mechanisms, scheduling strategies, and design overheads.19:04–21:27 · Guest teaching 0/10 ParallelKittens Architecture, Evaluation, and Adoption Stuart outlines the ParallelKittens programming framework and highlights real-world adoption at Cursor and Together AI.21:27–24:38 · Guest teaching 0/10 Jon Saad-Falcon on the Mainframe Era vs. PC AI Era Jon Saad-Falcon begins his presentation by drawing historical parallels between the mainframe computing era and current centralized cloud AI data centers.24:38–27:51 · Guest teaching 0/10 Intelligence Per Watt Metric and Empirical Findings Jon details the intelligence per watt and per joule metrics and presents empirical findings across various models and accelerators.27:51–31:05 · Guest teaching 0/10 Query Routing, Consumer Hardware Gaps, and OpenJarvis Jon discusses edge routing feasibility, consumer hardware investment gaps, and the OpenJarvis project.31:05–34:17 · Guest teaching 0/10 Mark Saroufim on AI vs. Systems Researchers and GPU Languages Mark Saroufim humorously vents about the friction between systems researchers who desire static computation and AI researchers who want dynamic control flow.34:17–36:20 · Guest teaching 0/10 GPU Kernel Competitions & Leaderboard Realities Mark discusses GPU kernel leaderboards on KernelBot and the unexpected surge of non-expert contributors placing at the top using LLMs.36:20–39:03 · Guest teaching 0/10 Kernel Evaluation Frameworks & Reward Hacking Examples Mark explains how evaluation suites work and dismisses the naive assumption that verifiable rewards make kernel evaluation simple by showing zero-mean vector hacks.39:03–43:13 · Guest teaching 0/10 Superhuman Reward Hacks & The Volkswagen Analogy Mark details sophisticated reward hacks, comparing an AI's behavior under eval testing to the Volkswagen Dieselgate emissions cheat.43:13–47:04 · Guest teaching 0/10 GPU Memory Hierarchy & The PyTorch Flywheel Mark walks through the QR decomposition benchmark, the multi-thousand line generated kernels, and the 9-year evolutionary flywheel of PyTorch correctness.47:04–49:45 · Guest teaching 0/10 Why AI Inference Needs Heterogeneous Hardware Misha Smulyanski introduces Marlowe and sets out first-principles arguments for why AI inference demands heterogeneous infrastructure.49:45–52:59 · Guest teaching 0/10 Prefill vs. Decode & Arithmetic Intensity Misha uses the roofline model and arithmetic intensity to contrast compute-bound prefill phases with bandwidth-bound decode phases.52:59–56:27 · Guest teaching 0/10 Workload Diversity & SRAM Accelerators Misha categorizes workload diversity across LLM use cases and reviews on-die SRAM accelerators for memory-bound matrix-vector operations.56:27–1:03:08 · Guest teaching 0/10 Heterogeneous Architectures: PDD, AFD, & Speculative Decoding Misha outlines heterogeneous architectures including prefill/decode disaggregation, attention vs MLP disaggregation, and offloaded speculative decoding.1:03:08–1:06:31 · Guest teaching 0/10 Full-Stack Heterogeneous Inference Challenges & Wrap-up Misha concludes by highlighting data center power density, networking topology, and performance modeling challenges in heterogeneous co-design.1:06:31–1:10:50 · Guest teaching 0/10 Entity Component System (ECS) Architecture for GPUs Brennan introduces his Stanford research on running high-throughput batched game engine simulations entirely on GPUs using Entity Component System architectures.1:10:50–1:12:50 · Guest teaching 0/10 ECS Task Graph Execution and Low-Level GPU Implementation Details Brennan explains the GPU task graph execution, garbage-collection memory management, and persistent mega-kernels achieving high SM utilization.1:12:50–1:15:35 · Guest teaching 0/10 Benchmark Environments, End-to-End RL Training, and Speedup Metrics Brennan concludes with benchmark results showing over 100x speedups over CPU baselines and advocates for high-level Python-esque GPU scripting abstractions.0:00–3:29 · Guest disagreement 0/10 Title Sequence and Paper Abstract Overview Francois Chaubard opens the YC Paper Club session solo, introducing the theme of chip and kernel specialization.3:29–5:40 · Guest disagreement 0/10 Chip and Data Center Specialization Trends Francois continues his opening monologue discussing TPU specialization, batch size one inference latency, and data center training dynamics.5:40–11:32 · Guest disagreement 0/10 Upcoming Schedule and Introduction of First Speaker Francois covers upcoming schedule logistics, historical anecdotes about YC and Anybots, and introduces the first speaker Stuart.11:32–14:19 · Guest disagreement 0/10 GPU Architecture and Memory Hierarchy Fundamentals Stuart delivers a solo technical talk explaining GPU architecture fundamentals, SMs, memory hierarchy, and NVLink.14:19–19:04 · Guest disagreement 0/10 Three Key Trade-Offs in Multi-GPU Kernel Design Stuart breaks down the three multi-GPU kernel design trade-offs: transfer mechanisms, scheduling strategies, and design overheads.19:04–21:27 · Guest disagreement 0/10 ParallelKittens Architecture, Evaluation, and Adoption Stuart outlines the ParallelKittens programming framework and highlights real-world adoption at Cursor and Together AI.21:27–24:38 · Guest disagreement 0/10 Jon Saad-Falcon on the Mainframe Era vs. PC AI Era Jon Saad-Falcon begins his presentation by drawing historical parallels between the mainframe computing era and current centralized cloud AI data centers.24:38–27:51 · Guest disagreement 0/10 Intelligence Per Watt Metric and Empirical Findings Jon details the intelligence per watt and per joule metrics and presents empirical findings across various models and accelerators.27:51–31:05 · Guest disagreement 0/10 Query Routing, Consumer Hardware Gaps, and OpenJarvis Jon discusses edge routing feasibility, consumer hardware investment gaps, and the OpenJarvis project.31:05–34:17 · Guest disagreement 2/10 Mark Saroufim on AI vs. Systems Researchers and GPU Languages Mark Saroufim humorously vents about the friction between systems researchers who desire static computation and AI researchers who want dynamic control flow.34:17–36:20 · Guest disagreement 0/10 GPU Kernel Competitions & Leaderboard Realities Mark discusses GPU kernel leaderboards on KernelBot and the unexpected surge of non-expert contributors placing at the top using LLMs.36:20–39:03 · Guest disagreement 2/10 Kernel Evaluation Frameworks & Reward Hacking Examples Mark explains how evaluation suites work and dismisses the naive assumption that verifiable rewards make kernel evaluation simple by showing zero-mean vector hacks.39:03–43:13 · Guest disagreement 1/10 Superhuman Reward Hacks & The Volkswagen Analogy Mark details sophisticated reward hacks, comparing an AI's behavior under eval testing to the Volkswagen Dieselgate emissions cheat.43:13–47:04 · Guest disagreement 0/10 GPU Memory Hierarchy & The PyTorch Flywheel Mark walks through the QR decomposition benchmark, the multi-thousand line generated kernels, and the 9-year evolutionary flywheel of PyTorch correctness.47:04–49:45 · Guest disagreement 0/10 Why AI Inference Needs Heterogeneous Hardware Misha Smulyanski introduces Marlowe and sets out first-principles arguments for why AI inference demands heterogeneous infrastructure.49:45–52:59 · Guest disagreement 0/10 Prefill vs. Decode & Arithmetic Intensity Misha uses the roofline model and arithmetic intensity to contrast compute-bound prefill phases with bandwidth-bound decode phases.52:59–56:27 · Guest disagreement 0/10 Workload Diversity & SRAM Accelerators Misha categorizes workload diversity across LLM use cases and reviews on-die SRAM accelerators for memory-bound matrix-vector operations.56:27–1:03:08 · Guest disagreement 0/10 Heterogeneous Architectures: PDD, AFD, & Speculative Decoding Misha outlines heterogeneous architectures including prefill/decode disaggregation, attention vs MLP disaggregation, and offloaded speculative decoding.1:03:08–1:06:31 · Guest disagreement 0/10 Full-Stack Heterogeneous Inference Challenges & Wrap-up Misha concludes by highlighting data center power density, networking topology, and performance modeling challenges in heterogeneous co-design.1:06:31–1:10:50 · Guest disagreement 0/10 Entity Component System (ECS) Architecture for GPUs Brennan introduces his Stanford research on running high-throughput batched game engine simulations entirely on GPUs using Entity Component System architectures.1:10:50–1:12:50 · Guest disagreement 0/10 ECS Task Graph Execution and Low-Level GPU Implementation Details Brennan explains the GPU task graph execution, garbage-collection memory management, and persistent mega-kernels achieving high SM utilization.1:12:50–1:15:35 · Guest disagreement 0/10 Benchmark Environments, End-to-End RL Training, and Speedup Metrics Brennan concludes with benchmark results showing over 100x speedups over CPU baselines and advocates for high-level Python-esque GPU scripting abstractions.0:00–3:29 · The partners pushing back 0/10 Title Sequence and Paper Abstract Overview Francois Chaubard opens the YC Paper Club session solo, introducing the theme of chip and kernel specialization.3:29–5:40 · The partners pushing back 0/10 Chip and Data Center Specialization Trends Francois continues his opening monologue discussing TPU specialization, batch size one inference latency, and data center training dynamics.5:40–11:32 · The partners pushing back 0/10 Upcoming Schedule and Introduction of First Speaker Francois covers upcoming schedule logistics, historical anecdotes about YC and Anybots, and introduces the first speaker Stuart.11:32–14:19 · The partners pushing back 0/10 GPU Architecture and Memory Hierarchy Fundamentals Stuart delivers a solo technical talk explaining GPU architecture fundamentals, SMs, memory hierarchy, and NVLink.14:19–19:04 · The partners pushing back 0/10 Three Key Trade-Offs in Multi-GPU Kernel Design Stuart breaks down the three multi-GPU kernel design trade-offs: transfer mechanisms, scheduling strategies, and design overheads.19:04–21:27 · The partners pushing back 0/10 ParallelKittens Architecture, Evaluation, and Adoption Stuart outlines the ParallelKittens programming framework and highlights real-world adoption at Cursor and Together AI.21:27–24:38 · The partners pushing back 0/10 Jon Saad-Falcon on the Mainframe Era vs. PC AI Era Jon Saad-Falcon begins his presentation by drawing historical parallels between the mainframe computing era and current centralized cloud AI data centers.24:38–27:51 · The partners pushing back 0/10 Intelligence Per Watt Metric and Empirical Findings Jon details the intelligence per watt and per joule metrics and presents empirical findings across various models and accelerators.27:51–31:05 · The partners pushing back 0/10 Query Routing, Consumer Hardware Gaps, and OpenJarvis Jon discusses edge routing feasibility, consumer hardware investment gaps, and the OpenJarvis project.31:05–34:17 · The partners pushing back 0/10 Mark Saroufim on AI vs. Systems Researchers and GPU Languages Mark Saroufim humorously vents about the friction between systems researchers who desire static computation and AI researchers who want dynamic control flow.34:17–36:20 · The partners pushing back 0/10 GPU Kernel Competitions & Leaderboard Realities Mark discusses GPU kernel leaderboards on KernelBot and the unexpected surge of non-expert contributors placing at the top using LLMs.36:20–39:03 · The partners pushing back 0/10 Kernel Evaluation Frameworks & Reward Hacking Examples Mark explains how evaluation suites work and dismisses the naive assumption that verifiable rewards make kernel evaluation simple by showing zero-mean vector hacks.39:03–43:13 · The partners pushing back 0/10 Superhuman Reward Hacks & The Volkswagen Analogy Mark details sophisticated reward hacks, comparing an AI's behavior under eval testing to the Volkswagen Dieselgate emissions cheat.43:13–47:04 · The partners pushing back 0/10 GPU Memory Hierarchy & The PyTorch Flywheel Mark walks through the QR decomposition benchmark, the multi-thousand line generated kernels, and the 9-year evolutionary flywheel of PyTorch correctness.47:04–49:45 · The partners pushing back 0/10 Why AI Inference Needs Heterogeneous Hardware Misha Smulyanski introduces Marlowe and sets out first-principles arguments for why AI inference demands heterogeneous infrastructure.49:45–52:59 · The partners pushing back 0/10 Prefill vs. Decode & Arithmetic Intensity Misha uses the roofline model and arithmetic intensity to contrast compute-bound prefill phases with bandwidth-bound decode phases.52:59–56:27 · The partners pushing back 0/10 Workload Diversity & SRAM Accelerators Misha categorizes workload diversity across LLM use cases and reviews on-die SRAM accelerators for memory-bound matrix-vector operations.56:27–1:03:08 · The partners pushing back 0/10 Heterogeneous Architectures: PDD, AFD, & Speculative Decoding Misha outlines heterogeneous architectures including prefill/decode disaggregation, attention vs MLP disaggregation, and offloaded speculative decoding.1:03:08–1:06:31 · The partners pushing back 0/10 Full-Stack Heterogeneous Inference Challenges & Wrap-up Misha concludes by highlighting data center power density, networking topology, and performance modeling challenges in heterogeneous co-design.1:06:31–1:10:50 · The partners pushing back 0/10 Entity Component System (ECS) Architecture for GPUs Brennan introduces his Stanford research on running high-throughput batched game engine simulations entirely on GPUs using Entity Component System architectures.1:10:50–1:12:50 · The partners pushing back 0/10 ECS Task Graph Execution and Low-Level GPU Implementation Details Brennan explains the GPU task graph execution, garbage-collection memory management, and persistent mega-kernels achieving high SM utilization.1:12:50–1:15:35 · The partners pushing back 0/10 Benchmark Environments, End-to-End RL Training, and Speedup Metrics Brennan concludes with benchmark results showing over 100x speedups over CPU baselines and advocates for high-level Python-esque GPU scripting abstractions.

speaking balance: gold is the partners, purple is the guest (3 minute bins)

0:00 · the partners 0% · guest 100%0:00 · the partners 0% · guest 100%3:00 · the partners 0% · guest 100%3:00 · the partners 0% · guest 100%6:00 · the partners 0% · guest 100%6:00 · the partners 0% · guest 100%9:00 · the partners 0% · guest 100%9:00 · the partners 0% · guest 100%12:00 · the partners 0% · guest 100%12:00 · the partners 0% · guest 100%15:00 · the partners 0% · guest 100%15:00 · the partners 0% · guest 100%18:00 · the partners 0% · guest 100%18:00 · the partners 0% · guest 100%21:00 · the partners 0% · guest 100%21:00 · the partners 0% · guest 100%24:00 · the partners 0% · guest 100%24:00 · the partners 0% · guest 100%27:00 · the partners 0% · guest 100%27:00 · the partners 0% · guest 100%30:00 · the partners 0% · guest 100%30:00 · the partners 0% · guest 100%33:00 · the partners 0% · guest 100%33:00 · the partners 0% · guest 100%36:00 · the partners 0% · guest 100%36:00 · the partners 0% · guest 100%39:00 · the partners 0% · guest 100%39:00 · the partners 0% · guest 100%42:00 · the partners 0% · guest 100%42:00 · the partners 0% · guest 100%45:00 · the partners 0% · guest 100%45:00 · the partners 0% · guest 100%48:00 · the partners 0% · guest 100%48:00 · the partners 0% · guest 100%51:00 · the partners 0% · guest 100%51:00 · the partners 0% · guest 100%54:00 · the partners 0% · guest 100%54:00 · the partners 0% · guest 100%57:00 · the partners 0% · guest 100%57:00 · the partners 0% · guest 100%1:00:00 · the partners 0% · guest 100%1:00:00 · the partners 0% · guest 100%1:03:00 · the partners 0% · guest 100%1:03:00 · the partners 0% · guest 100%1:06:00 · the partners 0% · guest 100%1:06:00 · the partners 0% · guest 100%1:09:00 · the partners 0% · guest 100%1:09:00 · the partners 0% · guest 100%1:12:00 · the partners 0% · guest 100%1:12:00 · the partners 0% · guest 100%1:15:00 · the partners 0% · guest 100%1:15:00 · the partners 0% · guest 100%
Sharpest disagreement ▶ 36:20 Challenging verifiable reward assumptions

Mark pushes back against the widely held belief that kernel evaluation is straightforward with verifiable rewards, demonstrating how LLMs exploit test harness quirks.

Hardest push from the partners ▶ 3:55 Critiquing throughput priority over latency

Francois rejects prioritizing pure throughput in conversational voice AI setups because high latency degrades user interaction.

Biggest teaching moment ▶ 14:20 Explaining transfer mechanism trade-offs

Stuart breaks down why copy engines fall short for fine-grained multi-GPU communication and why TMAs or register instructions are required.

The partners hold their own ▶ 4:20 Differentiating training and inference interconnects

Francois demonstrates systems domain knowledge by explaining how training data centers require all-to-all gradient passing whereas inference only requires all-to-all for MOE routing.

the scores for every segment, with the reasoning behind each
ChapterTopicThe partners as informed peerGuest teachingGuest disagreementThe partners pushing backWhy
Title Sequence and Paper Abstract Overview 0000 Francois Chaubard opens the YC Paper Club session solo, introducing the theme of chip and kernel specialization.
Chip and Data Center Specialization Trends 0000 Francois continues his opening monologue discussing TPU specialization, batch size one inference latency, and data center training dynamics.
Upcoming Schedule and Introduction of First Speaker 0000 Francois covers upcoming schedule logistics, historical anecdotes about YC and Anybots, and introduces the first speaker Stuart.
GPU Architecture and Memory Hierarchy Fundamentals 0000 Stuart delivers a solo technical talk explaining GPU architecture fundamentals, SMs, memory hierarchy, and NVLink.
Three Key Trade-Offs in Multi-GPU Kernel Design 0000 Stuart breaks down the three multi-GPU kernel design trade-offs: transfer mechanisms, scheduling strategies, and design overheads.
ParallelKittens Architecture, Evaluation, and Adoption 0000 Stuart outlines the ParallelKittens programming framework and highlights real-world adoption at Cursor and Together AI.
Jon Saad-Falcon on the Mainframe Era vs. PC AI Era 0000 Jon Saad-Falcon begins his presentation by drawing historical parallels between the mainframe computing era and current centralized cloud AI data centers.
Intelligence Per Watt Metric and Empirical Findings 0000 Jon details the intelligence per watt and per joule metrics and presents empirical findings across various models and accelerators.
Query Routing, Consumer Hardware Gaps, and OpenJarvis 0000 Jon discusses edge routing feasibility, consumer hardware investment gaps, and the OpenJarvis project.
Mark Saroufim on AI vs. Systems Researchers and GPU Languages 0020 Mark Saroufim humorously vents about the friction between systems researchers who desire static computation and AI researchers who want dynamic control flow.
GPU Kernel Competitions & Leaderboard Realities 0000 Mark discusses GPU kernel leaderboards on KernelBot and the unexpected surge of non-expert contributors placing at the top using LLMs.
Kernel Evaluation Frameworks & Reward Hacking Examples 0020 Mark explains how evaluation suites work and dismisses the naive assumption that verifiable rewards make kernel evaluation simple by showing zero-mean vector hacks.
Superhuman Reward Hacks & The Volkswagen Analogy 0010 Mark details sophisticated reward hacks, comparing an AI's behavior under eval testing to the Volkswagen Dieselgate emissions cheat.
GPU Memory Hierarchy & The PyTorch Flywheel 0000 Mark walks through the QR decomposition benchmark, the multi-thousand line generated kernels, and the 9-year evolutionary flywheel of PyTorch correctness.
Why AI Inference Needs Heterogeneous Hardware 0000 Misha Smulyanski introduces Marlowe and sets out first-principles arguments for why AI inference demands heterogeneous infrastructure.
Prefill vs. Decode & Arithmetic Intensity 0000 Misha uses the roofline model and arithmetic intensity to contrast compute-bound prefill phases with bandwidth-bound decode phases.
Workload Diversity & SRAM Accelerators 0000 Misha categorizes workload diversity across LLM use cases and reviews on-die SRAM accelerators for memory-bound matrix-vector operations.
Heterogeneous Architectures: PDD, AFD, & Speculative Decoding 0000 Misha outlines heterogeneous architectures including prefill/decode disaggregation, attention vs MLP disaggregation, and offloaded speculative decoding.
Full-Stack Heterogeneous Inference Challenges & Wrap-up 0000 Misha concludes by highlighting data center power density, networking topology, and performance modeling challenges in heterogeneous co-design.
Entity Component System (ECS) Architecture for GPUs 0000 Brennan introduces his Stanford research on running high-throughput batched game engine simulations entirely on GPUs using Entity Component System architectures.
ECS Task Graph Execution and Low-Level GPU Implementation Details 0000 Brennan explains the GPU task graph execution, garbage-collection memory management, and persistent mega-kernels achieving high SM utilization.
Benchmark Environments, End-to-End RL Training, and Speedup Metrics 0000 Brennan concludes with benchmark results showing over 100x speedups over CPU baselines and advocates for high-level Python-esque GPU scripting abstractions.

Statements from this episode (17)

Prediction Not checkable as stated
Chaubard: Massive chip specialization will split training and inference data center specs
“We just see that at the chip level, there's going to be massive specialization. The specs that you would need for a training data center are going to be very different than the specs for a inference data center.”
Francois Chaubard Jul 29, 2026 ▶ 0:38
Assertion Not checkable as stated
Chaubard: CPU-based environment simulation bottlenecks on-policy reinforcement learning rollouts
“The, it's amazing how much of simulators, when you call environment.step, is still run on the CPU, and so that's usually the bottleneck for a lot of your on-policy rollouts”
Francois Chaubard Jul 29, 2026 ▶ 2:50
Prediction Not checkable as stated
Chaubard: Chip specialization and ASICs will proliferate across AI workloads
“From my vantage point, we're gonna see this proliferation, and where we have sufficient demand now, because there's so much demand for tokens, Where it makes sense to specialize at the chip level that will pay out and allow you to go through the full new produ…”
Francois Chaubard Jul 29, 2026 ▶ 3:32
Assertion Supported
Chaubard: Inference stacks split between Nvidia for pre-fill and Cerebras for decode
“We already do see it a little bit on where a lot of the stack peep for inference. People will go to Nvidia for pre-fill and then they'll go to cerebris for decode engine.”
Francois Chaubard Jul 29, 2026 ▶ 4:00
Opinion
Chaubard: AI models are recursively distilling into one another, Claude into Kimi
“What I think is kind of happening right now, but like clawed kind of mother birds into, Kimmy too, and then now Kimmy too is post-training thinking machines, and so it just keeps going and going.”
Francois Chaubard Jul 29, 2026 ▶ 6:51
Assertion Supported
Stuart: GPU networking can consume up to 50% of LLaMA pre-fill runtime
“For example, networking can still consume up to 50% of total runtime for workloads like Lama's MDB pre-fill.”
Stuart (Stu) Jul 29, 2026 ▶ 8:30
Prediction Held up
Stuart: Single NVLink domains will soon scale to hundreds of GPUs
“And you also have a scale of architectures like NBL-L-Seventy-Two, which packs 72 GPUs inside a single NVLink domain, and soon this is going to extend to hundreds of GPUs.”
Stuart (Stu) Jul 29, 2026 ▶ 9:43
Assertion Supported
Stuart: Multi-GPU compilers often produce kernels slower than baseline
“Two is to use compiler-based approaches, but we find these compilers, compilers to be quite suboptimal. They produce kernels that are sometimes slower than non-overlapped baselines”
Stuart (Stu) Jul 29, 2026 ▶ 10:43
Insight
Stuart: Multi-GPU optimization requires overlapping compute with remote memory prefetching
“For multi-GPU kernels a simple, a similar idea applies, except that you're overlapping computation with communication with other GPUs, such that when the current computation is done, the data for next computation is ready and fetched from remote GPU HPMs.”
Stuart (Stu) Jul 29, 2026 ▶ 14:02
Assertion Supported
Stuart: TMA saturates NVLink on Blackwell using roughly 15 SMs
“So on Blackwell, we find that with roughly 15 SMs out of one 48 TMA is able to saturate the NVLink.”
Stuart (Stu) Jul 29, 2026 ▶ 15:27
Assertion Supported
Stuart: Bypassing NCCL intermediate buffers speeds up all-reduce by 80%
“For example, Nickel's default mode forces intermediate buffers which adds extra data movement between the sender and the receiver. And for fine-grained communication, this overhead really accumulates, and by stripping it out, you can speed up an operation as s…”
Stuart (Stu) Jul 29, 2026 ▶ 18:36
Assertion Supported
Stuart: ParallelKittens matches hand-optimized kernels in 50 to 100 lines
“And what we find is that with roughly 50 to 100 lines of device code, PK is able to surpass or match hand-optimized kernels that are often hundreds to thousands lines of code.”
Stuart (Stu) Jul 29, 2026 ▶ 20:38
Assertion Supported
Stuart: Cursor uses ParallelKittens on tens of thousands of Blackwell GPUs
“For example, Cursor is using it to train Composer on tens of thousands of Blackwell GPUs.”
Stuart (Stu) Jul 29, 2026 ▶ 20:57
Insight
Smulyanski: AI inference demands heterogeneous hardware co-designed for different phases
“Inference Is a very heterogeneous workload, right? Different phases of inference exercise, compute, network, storage, memory bandwidths differently, and so when we look at it makes sense to actually co-design the systems that will opt to, you know, use differe…”
Misha Smulyanski Jul 29, 2026 ▶ 47:35
Assertion Supported
Smulyanski: LLM decode workloads remain bandwidth-bound even across large batch sizes
“Pre-fill is generally very compute bound. Right because you basically do, ah, like, attention, ah, you do, ah, ah, a lot of, ah work for, you know, for, you know, all the tokens that you are fetching, right? You can, ah, you fetch the weights while I'm once an…”
Misha Smulyanski Jul 29, 2026 ▶ 51:45
Insight
Smulyanski: On-die SRAM accelerators excel at LLM decode due to high bandwidth
“The SRA machine basically keeps the entire weight matrix in SRA memory on DAI, so the, you got a lot more bandwidth, right, because it's on chip, right, so you can access, you know bytes over cycles, right, the chip interconnect is also fast, you can go, like …”
Misha Smulyanski Jul 29, 2026 ▶ 55:29
Insight
Smulyanski: GPU throughput drops sharply at low concurrency due to kernel overheads
“The moment that you start basically going to lower concurrency because you want better interactivity and better latency, Right? The performance the throughput drops. And it drops very sharply because all of a sudden you have a lot of, like, smaller kernels, yo…”
Misha Smulyanski Jul 29, 2026 ▶ 59:40
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.