Nov 3, 2025 · 1h 0m · latent-space

How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony

Quentin Anthony · 36m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Quentin Anthony explores Zyphra's adoption of AMD MI300X hardware and hybrid Mamba-transformer architectures, providing deep technical insights into low-level GPU kernel optimization and empirical strategies for maximizing AI-assisted engineering productivity.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 5.5 Guest teaching 5.6 Guest disagreement 1.7 The hosts pushing back 2.1
05100:0015:0030:0045:001:00:002:09–5:25 · The hosts as informed peer 6/10 Migrating Foundation Model Training to AMD GPUs Alessio references Sara Hooker's hardware lottery hypothesis to probe whether AMD adoption requires elite engineering talent. Quentin provides a deep technical breakdown of MI-250X vs MI-300X memory bandwidth and parallelism advantages over H100.5:25–10:13 · The hosts as informed peer 6/10 Evaluating Software Toolchains and the ROCm Ecosystem Alessio asks about alternative compilers like Mojo, TinyGrad, and the dilemma of waiting for next-gen Nvidia chips. Quentin explains why high-level compilers add abstraction overhead and why he writes bottom-up ROCm kernels directly.10:14–12:48 · The hosts as informed peer 5/10 Navigating the Hierarchy of GPU Kernel Development Alessio asks for an architectural overview of kernel development tiers for application engineers. Quentin delivers a comprehensive taxonomy spanning PTX/GCN assembly, CUDA/ROCm, cuDNN GEMM libraries, Cutlass/Composable Kernel templates, up to Triton.12:51–15:36 · The hosts as informed peer 5/10 Challenges of AI Code Generation for Low-Level Kernels Alessio questions whether synthetic datasets or curated kernel corpuses can fix LLM code generation for GPU kernels. Quentin breaks down why verification and evaluation in parallel systems make synthetic kernel generation and validation exceptionally difficult.15:38–18:28 · The hosts as informed peer 6/10 Kernel Authoring Workflow and Memory Hierarchy Optimization Alessio asks what defines a kernel author's workflow and performance targets. Quentin explains precise memory hierarchy placement, specifically controlling whether tensors live in registers or HBM rather than trusting compiler heuristics.18:29–23:31 · The hosts as informed peer 6/10 Analyzing ASICs, Specialized Silicon, and Co-Design Alessio pushes on custom ASIC economics, asking about 6-to-18 month tapeout trade-offs and structural differences in SRAM architectures like Cerebras and Grok. Quentin discusses hardware co-design principles and powers-of-two dimension tuning.23:31–29:39 · The hosts as informed peer 6/10 On-Device Inference Strategies and Architectural Innovation Alessio brings up on-device handoffs citing Greg Brockman and tests of Ollama on MacBook hardware. Quentin outlines Zyphra's tiered model sizing strategy (1.2B, 2.7B, 7B) designed for edge-to-cloud offloading.29:41–33:20 · The hosts as informed peer 6/10 Insights from the METR Developer Productivity Benchmark Alessio probes the methodology of the METR study, questioning whether randomized AI usage penalizes overall productivity benchmarks. Quentin clarifies the study mechanics, randomized task assignments, and how documentation reading vs code generation was measured.33:20–38:51 · The hosts as informed peer 5/10 Developer Setup, Context Rot, and Model Selection Alessio and Quentin discuss IDE setups and context management. Quentin explains why he avoids black-box wrappers like Cursor due to context rot and lack of prompt visibility, preferring direct API control and model-specific task routing.38:51–43:29 · The hosts as informed peer 5/10 Overcoming Developer Failure Modes and Sunk Cost Fallacies Alessio asks why Quentin outperformed other developers in the benchmark study. Quentin explains developer cognitive traps like the slot-machine sunk cost fallacy where devs spend hours re-prompting rather than coding manually.43:30–47:49 · The hosts as informed peer 6/10 Context Compaction and Preventing Organizational Code Slop Alessio references context compaction research from Sourcegraph and Chroma. Quentin details his multi-turn summary refresh loop and warns about organizational code slop where uninspected AI code creates tech debt.47:49–52:35 · The hosts as informed peer 5/10 Hiring High-Velocity Engineers and AI-Proof Interviews Alessio asks if interview processes have shifted to counter AI cheating. Quentin outlines his strict no-AI interview protocol focused on first-principles reasoning and velocity over legacy domain trivia.52:37–58:20 · The hosts as informed peer 6/10 EleutherAI's Scientific Mission and Open Source Priorities Alessio asks how EleutherAI fits into the open-source research landscape alongside DeepSeek and AI2. Quentin argues against grand unified consortiums, advocating for lean, well-funded, siloed research projects focused on interpretability and architectural sparsity.58:21–59:51 · The hosts as informed peer 4/10 Architectural Experimentation and Opportunities at Zyphra Alessio closes by asking for final takeaways and Zyphra hiring profiles. Quentin emphasizes hiring high-velocity generalists and challenging architectural dogma by testing hybrid attention alternatives like Mamba and DeltaNet.2:09–5:25 · Guest teaching 6/10 Migrating Foundation Model Training to AMD GPUs Alessio references Sara Hooker's hardware lottery hypothesis to probe whether AMD adoption requires elite engineering talent. Quentin provides a deep technical breakdown of MI-250X vs MI-300X memory bandwidth and parallelism advantages over H100.5:25–10:13 · Guest teaching 6/10 Evaluating Software Toolchains and the ROCm Ecosystem Alessio asks about alternative compilers like Mojo, TinyGrad, and the dilemma of waiting for next-gen Nvidia chips. Quentin explains why high-level compilers add abstraction overhead and why he writes bottom-up ROCm kernels directly.10:14–12:48 · Guest teaching 7/10 Navigating the Hierarchy of GPU Kernel Development Alessio asks for an architectural overview of kernel development tiers for application engineers. Quentin delivers a comprehensive taxonomy spanning PTX/GCN assembly, CUDA/ROCm, cuDNN GEMM libraries, Cutlass/Composable Kernel templates, up to Triton.12:51–15:36 · Guest teaching 6/10 Challenges of AI Code Generation for Low-Level Kernels Alessio questions whether synthetic datasets or curated kernel corpuses can fix LLM code generation for GPU kernels. Quentin breaks down why verification and evaluation in parallel systems make synthetic kernel generation and validation exceptionally difficult.15:38–18:28 · Guest teaching 6/10 Kernel Authoring Workflow and Memory Hierarchy Optimization Alessio asks what defines a kernel author's workflow and performance targets. Quentin explains precise memory hierarchy placement, specifically controlling whether tensors live in registers or HBM rather than trusting compiler heuristics.18:29–23:31 · Guest teaching 6/10 Analyzing ASICs, Specialized Silicon, and Co-Design Alessio pushes on custom ASIC economics, asking about 6-to-18 month tapeout trade-offs and structural differences in SRAM architectures like Cerebras and Grok. Quentin discusses hardware co-design principles and powers-of-two dimension tuning.23:31–29:39 · Guest teaching 4/10 On-Device Inference Strategies and Architectural Innovation Alessio brings up on-device handoffs citing Greg Brockman and tests of Ollama on MacBook hardware. Quentin outlines Zyphra's tiered model sizing strategy (1.2B, 2.7B, 7B) designed for edge-to-cloud offloading.29:41–33:20 · Guest teaching 6/10 Insights from the METR Developer Productivity Benchmark Alessio probes the methodology of the METR study, questioning whether randomized AI usage penalizes overall productivity benchmarks. Quentin clarifies the study mechanics, randomized task assignments, and how documentation reading vs code generation was measured.33:20–38:51 · Guest teaching 6/10 Developer Setup, Context Rot, and Model Selection Alessio and Quentin discuss IDE setups and context management. Quentin explains why he avoids black-box wrappers like Cursor due to context rot and lack of prompt visibility, preferring direct API control and model-specific task routing.38:51–43:29 · Guest teaching 6/10 Overcoming Developer Failure Modes and Sunk Cost Fallacies Alessio asks why Quentin outperformed other developers in the benchmark study. Quentin explains developer cognitive traps like the slot-machine sunk cost fallacy where devs spend hours re-prompting rather than coding manually.43:30–47:49 · Guest teaching 5/10 Context Compaction and Preventing Organizational Code Slop Alessio references context compaction research from Sourcegraph and Chroma. Quentin details his multi-turn summary refresh loop and warns about organizational code slop where uninspected AI code creates tech debt.47:49–52:35 · Guest teaching 5/10 Hiring High-Velocity Engineers and AI-Proof Interviews Alessio asks if interview processes have shifted to counter AI cheating. Quentin outlines his strict no-AI interview protocol focused on first-principles reasoning and velocity over legacy domain trivia.52:37–58:20 · Guest teaching 6/10 EleutherAI's Scientific Mission and Open Source Priorities Alessio asks how EleutherAI fits into the open-source research landscape alongside DeepSeek and AI2. Quentin argues against grand unified consortiums, advocating for lean, well-funded, siloed research projects focused on interpretability and architectural sparsity.58:21–59:51 · Guest teaching 4/10 Architectural Experimentation and Opportunities at Zyphra Alessio closes by asking for final takeaways and Zyphra hiring profiles. Quentin emphasizes hiring high-velocity generalists and challenging architectural dogma by testing hybrid attention alternatives like Mamba and DeltaNet.2:09–5:25 · Guest disagreement 2/10 Migrating Foundation Model Training to AMD GPUs Alessio references Sara Hooker's hardware lottery hypothesis to probe whether AMD adoption requires elite engineering talent. Quentin provides a deep technical breakdown of MI-250X vs MI-300X memory bandwidth and parallelism advantages over H100.5:25–10:13 · Guest disagreement 2/10 Evaluating Software Toolchains and the ROCm Ecosystem Alessio asks about alternative compilers like Mojo, TinyGrad, and the dilemma of waiting for next-gen Nvidia chips. Quentin explains why high-level compilers add abstraction overhead and why he writes bottom-up ROCm kernels directly.10:14–12:48 · Guest disagreement 1/10 Navigating the Hierarchy of GPU Kernel Development Alessio asks for an architectural overview of kernel development tiers for application engineers. Quentin delivers a comprehensive taxonomy spanning PTX/GCN assembly, CUDA/ROCm, cuDNN GEMM libraries, Cutlass/Composable Kernel templates, up to Triton.12:51–15:36 · Guest disagreement 1/10 Challenges of AI Code Generation for Low-Level Kernels Alessio questions whether synthetic datasets or curated kernel corpuses can fix LLM code generation for GPU kernels. Quentin breaks down why verification and evaluation in parallel systems make synthetic kernel generation and validation exceptionally difficult.15:38–18:28 · Guest disagreement 1/10 Kernel Authoring Workflow and Memory Hierarchy Optimization Alessio asks what defines a kernel author's workflow and performance targets. Quentin explains precise memory hierarchy placement, specifically controlling whether tensors live in registers or HBM rather than trusting compiler heuristics.18:29–23:31 · Guest disagreement 1/10 Analyzing ASICs, Specialized Silicon, and Co-Design Alessio pushes on custom ASIC economics, asking about 6-to-18 month tapeout trade-offs and structural differences in SRAM architectures like Cerebras and Grok. Quentin discusses hardware co-design principles and powers-of-two dimension tuning.23:31–29:39 · Guest disagreement 2/10 On-Device Inference Strategies and Architectural Innovation Alessio brings up on-device handoffs citing Greg Brockman and tests of Ollama on MacBook hardware. Quentin outlines Zyphra's tiered model sizing strategy (1.2B, 2.7B, 7B) designed for edge-to-cloud offloading.29:41–33:20 · Guest disagreement 2/10 Insights from the METR Developer Productivity Benchmark Alessio probes the methodology of the METR study, questioning whether randomized AI usage penalizes overall productivity benchmarks. Quentin clarifies the study mechanics, randomized task assignments, and how documentation reading vs code generation was measured.33:20–38:51 · Guest disagreement 2/10 Developer Setup, Context Rot, and Model Selection Alessio and Quentin discuss IDE setups and context management. Quentin explains why he avoids black-box wrappers like Cursor due to context rot and lack of prompt visibility, preferring direct API control and model-specific task routing.38:51–43:29 · Guest disagreement 1/10 Overcoming Developer Failure Modes and Sunk Cost Fallacies Alessio asks why Quentin outperformed other developers in the benchmark study. Quentin explains developer cognitive traps like the slot-machine sunk cost fallacy where devs spend hours re-prompting rather than coding manually.43:30–47:49 · Guest disagreement 2/10 Context Compaction and Preventing Organizational Code Slop Alessio references context compaction research from Sourcegraph and Chroma. Quentin details his multi-turn summary refresh loop and warns about organizational code slop where uninspected AI code creates tech debt.47:49–52:35 · Guest disagreement 3/10 Hiring High-Velocity Engineers and AI-Proof Interviews Alessio asks if interview processes have shifted to counter AI cheating. Quentin outlines his strict no-AI interview protocol focused on first-principles reasoning and velocity over legacy domain trivia.52:37–58:20 · Guest disagreement 3/10 EleutherAI's Scientific Mission and Open Source Priorities Alessio asks how EleutherAI fits into the open-source research landscape alongside DeepSeek and AI2. Quentin argues against grand unified consortiums, advocating for lean, well-funded, siloed research projects focused on interpretability and architectural sparsity.58:21–59:51 · Guest disagreement 1/10 Architectural Experimentation and Opportunities at Zyphra Alessio closes by asking for final takeaways and Zyphra hiring profiles. Quentin emphasizes hiring high-velocity generalists and challenging architectural dogma by testing hybrid attention alternatives like Mamba and DeltaNet.2:09–5:25 · The hosts pushing back 2/10 Migrating Foundation Model Training to AMD GPUs Alessio references Sara Hooker's hardware lottery hypothesis to probe whether AMD adoption requires elite engineering talent. Quentin provides a deep technical breakdown of MI-250X vs MI-300X memory bandwidth and parallelism advantages over H100.5:25–10:13 · The hosts pushing back 3/10 Evaluating Software Toolchains and the ROCm Ecosystem Alessio asks about alternative compilers like Mojo, TinyGrad, and the dilemma of waiting for next-gen Nvidia chips. Quentin explains why high-level compilers add abstraction overhead and why he writes bottom-up ROCm kernels directly.10:14–12:48 · The hosts pushing back 1/10 Navigating the Hierarchy of GPU Kernel Development Alessio asks for an architectural overview of kernel development tiers for application engineers. Quentin delivers a comprehensive taxonomy spanning PTX/GCN assembly, CUDA/ROCm, cuDNN GEMM libraries, Cutlass/Composable Kernel templates, up to Triton.12:51–15:36 · The hosts pushing back 2/10 Challenges of AI Code Generation for Low-Level Kernels Alessio questions whether synthetic datasets or curated kernel corpuses can fix LLM code generation for GPU kernels. Quentin breaks down why verification and evaluation in parallel systems make synthetic kernel generation and validation exceptionally difficult.15:38–18:28 · The hosts pushing back 2/10 Kernel Authoring Workflow and Memory Hierarchy Optimization Alessio asks what defines a kernel author's workflow and performance targets. Quentin explains precise memory hierarchy placement, specifically controlling whether tensors live in registers or HBM rather than trusting compiler heuristics.18:29–23:31 · The hosts pushing back 3/10 Analyzing ASICs, Specialized Silicon, and Co-Design Alessio pushes on custom ASIC economics, asking about 6-to-18 month tapeout trade-offs and structural differences in SRAM architectures like Cerebras and Grok. Quentin discusses hardware co-design principles and powers-of-two dimension tuning.23:31–29:39 · The hosts pushing back 2/10 On-Device Inference Strategies and Architectural Innovation Alessio brings up on-device handoffs citing Greg Brockman and tests of Ollama on MacBook hardware. Quentin outlines Zyphra's tiered model sizing strategy (1.2B, 2.7B, 7B) designed for edge-to-cloud offloading.29:41–33:20 · The hosts pushing back 3/10 Insights from the METR Developer Productivity Benchmark Alessio probes the methodology of the METR study, questioning whether randomized AI usage penalizes overall productivity benchmarks. Quentin clarifies the study mechanics, randomized task assignments, and how documentation reading vs code generation was measured.33:20–38:51 · The hosts pushing back 2/10 Developer Setup, Context Rot, and Model Selection Alessio and Quentin discuss IDE setups and context management. Quentin explains why he avoids black-box wrappers like Cursor due to context rot and lack of prompt visibility, preferring direct API control and model-specific task routing.38:51–43:29 · The hosts pushing back 2/10 Overcoming Developer Failure Modes and Sunk Cost Fallacies Alessio asks why Quentin outperformed other developers in the benchmark study. Quentin explains developer cognitive traps like the slot-machine sunk cost fallacy where devs spend hours re-prompting rather than coding manually.43:30–47:49 · The hosts pushing back 2/10 Context Compaction and Preventing Organizational Code Slop Alessio references context compaction research from Sourcegraph and Chroma. Quentin details his multi-turn summary refresh loop and warns about organizational code slop where uninspected AI code creates tech debt.47:49–52:35 · The hosts pushing back 2/10 Hiring High-Velocity Engineers and AI-Proof Interviews Alessio asks if interview processes have shifted to counter AI cheating. Quentin outlines his strict no-AI interview protocol focused on first-principles reasoning and velocity over legacy domain trivia.52:37–58:20 · The hosts pushing back 2/10 EleutherAI's Scientific Mission and Open Source Priorities Alessio asks how EleutherAI fits into the open-source research landscape alongside DeepSeek and AI2. Quentin argues against grand unified consortiums, advocating for lean, well-funded, siloed research projects focused on interpretability and architectural sparsity.58:21–59:51 · The hosts pushing back 1/10 Architectural Experimentation and Opportunities at Zyphra Alessio closes by asking for final takeaways and Zyphra hiring profiles. Quentin emphasizes hiring high-velocity generalists and challenging architectural dogma by testing hybrid attention alternatives like Mamba and DeltaNet.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 0%1:00:00 · the hosts 0% · guest 0%
Sharpest disagreement ▶ 56:15 Rejecting the idea of large open source consortiums

Quentin firmly pushes back against the conventional open source ideal of large groups banding together, arguing instead for siloed, focused research teams.

Hardest push from the hosts ▶ 32:20 Challenging task-level vs workflow-level AI benchmarking

Alessio presses Quentin on whether randomized task lotteries unfairly depress measured AI productivity by preventing end-to-end workflow reorganization.

Biggest teaching moment ▶ 10:40 Comprehensive breakdown of the GPU kernel stack

Quentin gives a masterclass on the hierarchy of GPU programming, delineating assembly, ROCm/CUDA, GEMM libraries, cutlass templates, and high-level DSLs.

The host holds their own ▶ 4:03 Analyzing MI-300X VRAM advantages over H100

Alessio prompts Quentin with the hardware lottery thesis, leading Quentin to demonstrate deep technical knowledge of MI-300X's 192GB VRAM and memory bandwidth.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Migrating Foundation Model Training to AMD GPUs 6622 Alessio references Sara Hooker's hardware lottery hypothesis to probe whether AMD adoption requires elite engineering talent. Quentin provides a deep technical breakdown of MI-250X vs MI-300X memory bandwidth and parallelism advantages over H100.
Evaluating Software Toolchains and the ROCm Ecosystem 6623 Alessio asks about alternative compilers like Mojo, TinyGrad, and the dilemma of waiting for next-gen Nvidia chips. Quentin explains why high-level compilers add abstraction overhead and why he writes bottom-up ROCm kernels directly.
Navigating the Hierarchy of GPU Kernel Development 5711 Alessio asks for an architectural overview of kernel development tiers for application engineers. Quentin delivers a comprehensive taxonomy spanning PTX/GCN assembly, CUDA/ROCm, cuDNN GEMM libraries, Cutlass/Composable Kernel templates, up to Triton.
Challenges of AI Code Generation for Low-Level Kernels 5612 Alessio questions whether synthetic datasets or curated kernel corpuses can fix LLM code generation for GPU kernels. Quentin breaks down why verification and evaluation in parallel systems make synthetic kernel generation and validation exceptionally difficult.
Kernel Authoring Workflow and Memory Hierarchy Optimization 6612 Alessio asks what defines a kernel author's workflow and performance targets. Quentin explains precise memory hierarchy placement, specifically controlling whether tensors live in registers or HBM rather than trusting compiler heuristics.
Analyzing ASICs, Specialized Silicon, and Co-Design 6613 Alessio pushes on custom ASIC economics, asking about 6-to-18 month tapeout trade-offs and structural differences in SRAM architectures like Cerebras and Grok. Quentin discusses hardware co-design principles and powers-of-two dimension tuning.
On-Device Inference Strategies and Architectural Innovation 6422 Alessio brings up on-device handoffs citing Greg Brockman and tests of Ollama on MacBook hardware. Quentin outlines Zyphra's tiered model sizing strategy (1.2B, 2.7B, 7B) designed for edge-to-cloud offloading.
Insights from the METR Developer Productivity Benchmark 6623 Alessio probes the methodology of the METR study, questioning whether randomized AI usage penalizes overall productivity benchmarks. Quentin clarifies the study mechanics, randomized task assignments, and how documentation reading vs code generation was measured.
Developer Setup, Context Rot, and Model Selection 5622 Alessio and Quentin discuss IDE setups and context management. Quentin explains why he avoids black-box wrappers like Cursor due to context rot and lack of prompt visibility, preferring direct API control and model-specific task routing.
Overcoming Developer Failure Modes and Sunk Cost Fallacies 5612 Alessio asks why Quentin outperformed other developers in the benchmark study. Quentin explains developer cognitive traps like the slot-machine sunk cost fallacy where devs spend hours re-prompting rather than coding manually.
Context Compaction and Preventing Organizational Code Slop 6522 Alessio references context compaction research from Sourcegraph and Chroma. Quentin details his multi-turn summary refresh loop and warns about organizational code slop where uninspected AI code creates tech debt.
Hiring High-Velocity Engineers and AI-Proof Interviews 5532 Alessio asks if interview processes have shifted to counter AI cheating. Quentin outlines his strict no-AI interview protocol focused on first-principles reasoning and velocity over legacy domain trivia.
EleutherAI's Scientific Mission and Open Source Priorities 6632 Alessio asks how EleutherAI fits into the open-source research landscape alongside DeepSeek and AI2. Quentin argues against grand unified consortiums, advocating for lean, well-funded, siloed research projects focused on interpretability and architectural sparsity.
Architectural Experimentation and Opportunities at Zyphra 4411 Alessio closes by asking for final takeaways and Zyphra hiring profiles. Quentin emphasizes hiring high-velocity generalists and challenging architectural dogma by testing hybrid attention alternatives like Mamba and DeltaNet.

Statements from this episode (25)

Disclosure
Zyphra moves its entire model training cluster to AMD hardware
“We recently moved all of our training cluster over to AMD. So we're really going all in on AMD ecosystem.”
Quentin Anthony Nov 3, 2025 ▶ 1:41
Assertion Supported
Zyphra: Zamba 2 7B beats Llama 3 8B using hybrid Mamba-Transformer architecture
“We released a Zamba two, which was a hybrid between transformers and a Mamba two blocks. And we were able to be like a Lama three eight B for example, with a seven B model.”
Quentin Anthony Nov 3, 2025 ▶ 1:59
Assertion Supported
AMD MI300X outperforms Nvidia H100 on FlashAttention-2 and memory-bound workloads
“We found that it's great for flash attention to specifically, we were able to be H-one hundred. We also found that like the less time you spend in like dense compute, like the less time you spend in tensor cores specifically, or less time you spend in lower bi…”
Quentin Anthony Nov 3, 2025 ▶ 3:19
Assertion Supported
AMD MI300X GEMM performance increased from 400 to 650 TFLOPS via software
“So when MI 300 X first dropped, if you run like an MLP gem or something, you would get like 400 T flops. Now that number is, you know, more like six, 606 107 650 or so just in like the gem back ends themselves.”
Quentin Anthony Nov 3, 2025 ▶ 4:55
Opinion
AMD has completely caught up with Nvidia on AI software
“So they caught up on hardware. Now they've caught up on software. Not a lot of people have sort of discovered that they've caught up on software and we're kind of capitalizing on that.”
Quentin Anthony Nov 3, 2025 ▶ 5:10
Insight
Training frontends matter little if attention and MLP kernels are highly optimized
“Most of that is an attention and MOPs, right? So if you have good kernels for attention, MOPs and norms and so on, then it doesn't much matter what the front end to, you know, send tensors to and from those kernels is”
Quentin Anthony Nov 3, 2025 ▶ 6:11
Assertion Supported
AMD's Composable Kernel library offers functionality similar to Nvidia's CUTLASS
“And then on the AMD side, they have this composable kernel library that does something very similar.”
Quentin Anthony Nov 3, 2025 ▶ 12:07
Insight
AI coding assistants degrade quickly on low-level CUDA and PTX code
“Getting Below Triton or level or anything else down to like CUDA or below there's orders of magnitude less public good kernels at that level. And I think that really shows and models capabilities. So when I try and get a model to do something in CUDA or PTX or…”
Quentin Anthony Nov 3, 2025 ▶ 13:35
Insight
AI kernel models game automated evaluation metrics with subtly incorrect code
“It would be great, but it's not a silver bullet because kernels are also hard to validate. It's hard to have like an eval in kernels. Every time that someone releases like, oh, we created a new eval that measures kernels and we trained a model that generates k…”
Quentin Anthony Nov 3, 2025 ▶ 14:41
Insight
Writing custom GPU kernels is a last resort in model optimization
“Well, kernel is a last resort. So first I go, oh, another kernel. And I try and find some way to go around it.”
Quentin Anthony Nov 3, 2025 ▶ 15:50
Opinion
Hyperscalers with fixed architectures should build custom AI inference ASICs
“If I was sort of in house, maybe XAI is a good example of this, or Microsoft is another one where they kind of know the model architecture a priori then they absolutely should make an ASIC that is custom designed for that model architecture that they're curren…”
Quentin Anthony Nov 3, 2025 ▶ 21:27
Opinion
Groq hardware is too inflexible to run alternative architectures like Mamba SSMs
“Grok is very inflexible hardware so that you kind of, if you want to do like a Mamba SSM on it, you're going to have a really hard time because instead of having like a low level CUDA compiler, like everything is like Designed it at the hardware level, right?”
Quentin Anthony Nov 3, 2025 ▶ 22:42
Prediction Not checkable as stated
AI labs will increasingly co-design models alongside proprietary custom inference ASICs
“I definitely think things are moving towards you design a model. It's sized the way that fits well on hardware that you can design. You immediately start creating an inference hardware that is custom made to fit the sizes of your model. And then you can also s…”
Quentin Anthony Nov 3, 2025 ▶ 23:02
Assertion Not checkable as stated
AI engineering speedups require specific use cases and strong digital hygiene
“AI sped me up a bit, but only in specific cases and only when taking a lot of sort of digital hygiene practices.”
Quentin Anthony Nov 3, 2025 ▶ 31:19
Disclosure
Quentin Anthony avoids Cursor to maintain strict control over LLM context
“So I personally don't really use tools like cursor because I want total control over the context. I know what models can handle, what prompts, and I know, for example, one thing I mentioned is context rot. So how long the context is before the model chokes on …”
Quentin Anthony Nov 3, 2025 ▶ 33:42
Assertion Not checkable as stated
First-principles GPU kernel modeling was beyond AI capabilities before OpenAI o1
“O1's like the initial thinking models were a big deal when I was doing like core academic, like how do I create a performance model for explaining how this kernel behaves? Like, From first principles, that kind of thing was not really in the scope of any model…”
Quentin Anthony Nov 3, 2025 ▶ 37:03
Assertion Supported
Quentin Anthony achieved the highest productivity speedup in the METR benchmark
“I think you're just kind of noise unless I tell people, okay, I was the one that got the most speed up in the study.”
Quentin Anthony Nov 3, 2025 ▶ 39:17
Prediction Open · timeframe Nov 2030
Anthropic and OpenAI will never open-source their high-performance inference kernels
“The high performance inference kernels that sort of drive a lot of, you know, anthropic and open AI and stuff, their models, those aren't open source. They're not going to be open source.”
Quentin Anthony Nov 3, 2025 ▶ 40:52
Insight
Offloading thinking to AI tools degrades developers' ability to evaluate code quality
“And I do suggest that people not try and use the model to offload thinking. It should enhance your thinking or else like you'll, you'll get worse over time and you won't know when the model is quality, whether it's outputting quality or not. If you don't know …”
Quentin Anthony Nov 3, 2025 ▶ 46:31
Prediction Not checkable as stated
Companies pushing AI code volume metrics will drown in unmaintainable slop
“I think actually there's a lot of big tech examples where they're kind of pushing really hard for their teams to use more code. They're being evaluated on how much code they're using. And they're just kind of getting more and more slop that nobody understands.…”
Quentin Anthony Nov 3, 2025 ▶ 47:30
Insight
String theorists ramp faster in AI engineering than complacent CUDA developers
“I don't really care if someone knows CUDA kernel writing. If someone does string theory and is really good at understanding complex problems, they will be productive faster than someone who knows CUDA and doesn't really care about trying to get better at it.”
Quentin Anthony Nov 3, 2025 ▶ 49:00
Disclosure
Zyphra strictly bans engineering candidates from using AI tools during interviews
“No AI in the interview, first off. And you gotta watch people's eyes now on what monitors they're looking at behind the screen, which is, that's changed in the last two or three years, unfortunately. No AI allowed.”
Quentin Anthony Nov 3, 2025 ▶ 49:55
Insight
Interviewers can spot AI usage by checking a candidate's response latency
“Anyone who's interviewed people can almost always tell, I think, about whether someone's using AI on the other side. How long is the time to first token? Humans are typically faster than a model.”
Quentin Anthony Nov 3, 2025 ▶ 51:05
Insight
Small, funded teams are more effective than broad open-source AI consortiums
“Maybe it sounds wrong, but I feel like banding together is not necessarily good by default. There's a lot of different incentives. There's so much noise that no one knows where to focus. I prefer, if anything, I would prefer siloed focused teams who have fundi…”
Quentin Anthony Nov 3, 2025 ▶ 56:15
Assertion Supported
AMD funded community DeepSeek kernel development through the GPU Mode Discord
“AMD's actually done this. Like there's some deep seek, like through GPU modes, discord, there have been some like deep seek kernels that they say you have, you know, guaranteed access to compute for, write them.”
Quentin Anthony Nov 3, 2025 ▶ 57:48
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.