Quentin Anthony, Head of Model Training at Zyphra, discusses the limitations of AI coding assistants when generating low-level GPU kernels versus high-level Triton abstractions.
“Getting Below Triton or level or anything else down to like CUDA or below there's orders of magnitude less public good kernels at that level. And I think that really shows and models capabilities. So when I try and get a model to do something in CUDA or PTX or something, as long, if it's not dead basic, it's bad really fast.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Quentin Anthony
Opinion
AMD has completely caught up with Nvidia on AI software
“So they caught up on hardware. Now they've caught up on software. Not a lot of people have sort of discovered that they've caught up on software and we're kind of capitalizing on that.”
Quentin AnthonyNov 3, 2025▶ 5:10How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
AssertionSupported
AMD MI300X outperforms Nvidia H100 on FlashAttention-2 and memory-bound workloads
“We found that it's great for flash attention to specifically, we were able to be H-one hundred. We also found that like the less time you spend in like dense compute, like the less time you spend in tensor cores specifically, or less time you spend in lower bi…”
Quentin AnthonyNov 3, 2025▶ 3:19How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
Insight
AI kernel models game automated evaluation metrics with subtly incorrect code
“It would be great, but it's not a silver bullet because kernels are also hard to validate. It's hard to have like an eval in kernels. Every time that someone releases like, oh, we created a new eval that measures kernels and we trained a model that generates k…”
Quentin AnthonyNov 3, 2025▶ 14:41How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
Opinion
Groq hardware is too inflexible to run alternative architectures like Mamba SSMs
“Grok is very inflexible hardware so that you kind of, if you want to do like a Mamba SSM on it, you're going to have a really hard time because instead of having like a low level CUDA compiler, like everything is like Designed it at the hardware level, right?”
Quentin AnthonyNov 3, 2025▶ 22:42How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
PredictionNot checkable as stated
AI labs will increasingly co-design models alongside proprietary custom inference ASICs
“I definitely think things are moving towards you design a model. It's sized the way that fits well on hardware that you can design. You immediately start creating an inference hardware that is custom made to fit the sizes of your model. And then you can also s…”
Quentin AnthonyNov 3, 2025▶ 23:02How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
Insight
Offloading thinking to AI tools degrades developers' ability to evaluate code quality
“And I do suggest that people not try and use the model to offload thinking. It should enhance your thinking or else like you'll, you'll get worse over time and you won't know when the model is quality, whether it's outputting quality or not. If you don't know …”
Quentin AnthonyNov 3, 2025▶ 46:31How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 200 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.