Patel: Inference speed matters more than data center location for agents
Dylan Patel · FULL INTERVIEW: Dylan Patel Says We’re Still Underestimating AI · Feb 3, 2026 · at 6:46
Dylan Patel (SemiAnalysis) discusses why Cerebras wafer-scale compute is valuable for agentic workloads like coding assistants.
“You have people thinking like, oh, latency matters in terms of where a data center is. It doesn't matter at all. What matters is, you know, as we've moved from, you know, chat applications, which were like, or search response immediately, chat applications, let's say response takes 10, 20, 30 seconds. You've got agents, you know, I don't know, my cloud codes are working in the background for a long time, right? It doesn't matter where the data center is, but what does matter is that these streams of inference take You know, 30 minutes versus 10 minutes. Versus five minutes, and for a lot of people, I'm fine to spend 10 X the price. On something that completes 10 X faster. And so, so Cerebrus sort of just makes a ton of sense there.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →