The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 8 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Insight
AI kernel models game automated evaluation metrics with subtly incorrect code
“It would be great, but it's not a silver bullet because kernels are also hard to validate. It's hard to have like an eval in kernels. Every time that someone releases like, oh, we created a new eval that measures kernels and we trained a model that generates k…”
Quentin Anthony Nov 3, 2025 ▶ 14:41 How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
Assertion Supported
AMD MI300X outperforms Nvidia H100 on FlashAttention-2 and memory-bound workloads
“We found that it's great for flash attention to specifically, we were able to be H-one hundred. We also found that like the less time you spend in like dense compute, like the less time you spend in tensor cores specifically, or less time you spend in lower bi…”
Quentin Anthony Nov 3, 2025 ▶ 3:19 How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
Insight
String theorists ramp faster in AI engineering than complacent CUDA developers
“I don't really care if someone knows CUDA kernel writing. If someone does string theory and is really good at understanding complex problems, they will be productive faster than someone who knows CUDA and doesn't really care about trying to get better at it.”
Quentin Anthony Nov 3, 2025 ▶ 49:00 How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
Disclosure
Zyphra moves its entire model training cluster to AMD hardware
“We recently moved all of our training cluster over to AMD. So we're really going all in on AMD ecosystem.”
Quentin Anthony Nov 3, 2025 ▶ 1:41 How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
Disclosure
Zyphra strictly bans engineering candidates from using AI tools during interviews
“No AI in the interview, first off. And you gotta watch people's eyes now on what monitors they're looking at behind the screen, which is, that's changed in the last two or three years, unfortunately. No AI allowed.”
Quentin Anthony Nov 3, 2025 ▶ 49:55 How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
Insight
Writing custom GPU kernels is a last resort in model optimization
“Well, kernel is a last resort. So first I go, oh, another kernel. And I try and find some way to go around it.”
Quentin Anthony Nov 3, 2025 ▶ 15:50 How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
Insight
AI coding assistants degrade quickly on low-level CUDA and PTX code
“Getting Below Triton or level or anything else down to like CUDA or below there's orders of magnitude less public good kernels at that level. And I think that really shows and models capabilities. So when I try and get a model to do something in CUDA or PTX or…”
Quentin Anthony Nov 3, 2025 ▶ 13:35 How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
Assertion Supported
Zyphra: Zamba 2 7B beats Llama 3 8B using hybrid Mamba-Transformer architecture
“We released a Zamba two, which was a hybrid between transformers and a Mamba two blocks. And we were able to be like a Lama three eight B for example, with a seven B model.”
Quentin Anthony Nov 3, 2025 ▶ 1:59 How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.