Disclosure certainty 4/5 debate potential 3/5

Quentin Anthony avoids Cursor to maintain strict control over LLM context

Quentin Anthony · How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony · Nov 3, 2025 · at 33:42

Quentin Anthony, Head of Model Training at Zyphra, explains his engineering workflow and why he avoids automated AI coding editors.

0:00 / 0:27exact quote · 27.6s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“So I personally don't really use tools like cursor because I want total control over the context. I know what models can handle, what prompts, and I know, for example, one thing I mentioned is context rot. So how long the context is before the model chokes on it. And I can't see those details of what cursor is passing to the model when I tell it You know, write me a test for this other function in my code base. I don't know what it's feeding the model. Maybe it's too much and Claude chokes on it. So I don't use tools like that.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Quentin Anthony

Opinion
AMD has completely caught up with Nvidia on AI software
“So they caught up on hardware. Now they've caught up on software. Not a lot of people have sort of discovered that they've caught up on software and we're kind of capitalizing on that.”
Quentin Anthony Nov 3, 2025 ▶ 5:10 How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
Assertion Supported
AMD MI300X outperforms Nvidia H100 on FlashAttention-2 and memory-bound workloads
“We found that it's great for flash attention to specifically, we were able to be H-one hundred. We also found that like the less time you spend in like dense compute, like the less time you spend in tensor cores specifically, or less time you spend in lower bi…”
Quentin Anthony Nov 3, 2025 ▶ 3:19 How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
Insight
AI kernel models game automated evaluation metrics with subtly incorrect code
“It would be great, but it's not a silver bullet because kernels are also hard to validate. It's hard to have like an eval in kernels. Every time that someone releases like, oh, we created a new eval that measures kernels and we trained a model that generates k…”
Quentin Anthony Nov 3, 2025 ▶ 14:41 How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
Opinion
Groq hardware is too inflexible to run alternative architectures like Mamba SSMs
“Grok is very inflexible hardware so that you kind of, if you want to do like a Mamba SSM on it, you're going to have a really hard time because instead of having like a low level CUDA compiler, like everything is like Designed it at the hardware level, right?”
Quentin Anthony Nov 3, 2025 ▶ 22:42 How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
Prediction Not checkable as stated
AI labs will increasingly co-design models alongside proprietary custom inference ASICs
“I definitely think things are moving towards you design a model. It's sized the way that fits well on hardware that you can design. You immediately start creating an inference hardware that is custom made to fit the sizes of your model. And then you can also s…”
Quentin Anthony Nov 3, 2025 ▶ 23:02 How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
Insight
Offloading thinking to AI tools degrades developers' ability to evaluate code quality
“And I do suggest that people not try and use the model to offload thinking. It should enhance your thinking or else like you'll, you'll get worse over time and you won't know when the model is quality, whether it's outputting quality or not. If you don't know …”
Quentin Anthony Nov 3, 2025 ▶ 46:31 How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.