Groq

includes Groq LPU

7 statements across 5 episodes · 2 bullish · 5 bearish · 5 people on the record · first statement Oct 18, 2024 by Drew Houston · said 41 times in 15 episodes since 2024 · across every show →

Mentions by year, the whole family

brought up most by Sean Lie (8), Shawn Wang (5), Matthew Berman (3), Dylan Patel (3), Doug O'Laughlin (2), Mitesh Agrawal (1), Kyle Kranen (1), Jesse Han (1)

tap a year for its mentions
00134258202420252026episodesmentions
048202420252026episodes it came up in
001.5438202420252026episodesmentions per episode
2026 23 mentions in 8 episodes 3 per episode
2025 11 mentions in 4 episodes 3 per episode
2024 7 mentions in 3 episodes 2 per episode

every mention, scene by scene, with the transcript →

Everything said about Groq, oldest first

Oct 18, 2024 positive
Assertion Supported
Houston: Groq and Cerebras outperform Nvidia on latency
“There's also, like, non-NVIDIA stacks, like the Grok, or Cerebris, or some of these custom silicon companies that are super interesting, and all, and outperformed the NVIDIA stack in terms of latency and things like that.”
Drew Houston Oct 18, 2024 ▶ 52:59 Building the Silicon Brain - Drew Houston of Dropbox
Aug 18, 2025 negative
Opinion
Agrawal: Cerebras and Groq offer cloud APIs because their software struggles
“You know, if you look at Cerebrus and Grok and others, they've really tried to do this, their cloud kind of portal. And for us, you know, it shows two things. One is the difficulty of software for third party to implement that, that that's why they're kind of …”
Mitesh Agrawal Aug 18, 2025 ▶ 39:13 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Nov 3, 2025 negative
Opinion
Groq hardware is too inflexible to run alternative architectures like Mamba SSMs
“Grok is very inflexible hardware so that you kind of, if you want to do like a Mamba SSM on it, you're going to have a really hard time because instead of having like a low level CUDA compiler, like everything is like Designed it at the hardware level, right?”
Quentin Anthony Nov 3, 2025 ▶ 22:42 How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
Feb 24, 2026 bearish
Opinion
O'Laughlin: Prior to Cerebras and Groq, AI accelerator startups were failures
“The reason why my hit rate for every AI accelerator trip is, like, very, like, I just don't believe in them is because, like, where are they? Until Cerebrus and Grok, honestly, they were all considered failures, and even then, we're like, what are they gonna d…”
Doug O'Laughlin Feb 24, 2026 ▶ 1:53:36 Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
Sep 2, 2026 bearish
Prediction Not checkable as stated
Lie: Groq will be forced to focus on significantly smaller models
“I think what's, what's going to end up happening is they're going to end up focusing on significantly smaller models. You know, if you have that limitation in your architecture, then I think that's what ends up happening.”
Sean Lie Sep 2, 2026 ▶ 23:42 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Sep 2, 2026 negative
Assertion Supported
Lie: Trillion-parameter models require thousands of Groq LPUs for weights
“To run a frontier level model, like, let's say, a few trillion parameters, you need thousands and thousands of Grok LPUs just to hold the weights, right?”
Sean Lie Sep 2, 2026 ▶ 23:17 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Sep 2, 2026 bullish
Assertion Supported
Lie: Cerebras chips have 100x more memory than Groq LPUs
“One of our chips has, You know, order a hundred times more memory than one of their chips, right? So you got two orders of magnitude difference in scale kind of for free, right?”
Sean Lie Sep 2, 2026 ▶ 23:58 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.