GPU

also referred to as: gpus

56 statements across 39 episodes · 20 bullish · 10 bearish · 41 people on the record · first statement Jun 20, 2023 by George Hotz · across every show →

Everything said about GPU, oldest first

Jun 20, 2023 positive
Assertion Supported
Hotz: Halving GPU power yields 80% of peak performance
“Now, you can limit power on GPUs and still get, you can use like half the power and get 80% of the performance. This is a known fact about GPUs”
George Hotz Jun 20, 2023 ▶ 40:06 Ep 18: Petaflops to the People — with George Hotz of tinycorp
Aug 31, 2023 bullish
Assertion Supported
RWKV Trains in Parallel Across GPUs, Unlike Traditional RNNs
“And in practice, once you start cascading there, you just saturate the GPU, and that's how it starts being paralysable trained. You no longer need to train in slices like traditional RNNs.”
Eugene Cheah Aug 31, 2023 ▶ 58:45 RWKV: Reinventing RNNs for the Transformer Era
Nov 3, 2023 positive
Insight
Royzen: NVIDIA Remains Cloud-Agnostic Because It Wins Regardless
“At NVIDIA, They know that they're going to win regardless. So they don't care where you get the GPUs from. They're like, they're truly neutral, unlike various sales reps that you might encounter at various like clouds and, you know, hardware companies, et cete…”
Michael Royzen Nov 3, 2023 ▶ 1:02:39 Beating GPT-4 with Open Source Models - with Michael Royzen of Phind
Dec 5, 2023 bullish
Prediction Didn’t hold up
Patel: AI inference will deploy more GPUs than training by 2024
“LLM inference will be bigger than training, or multimodal, whatever, blah, blah, blah inference will be bigger than training, you know, probably next year, in fact at least in terms of GPUs deployed,”
Dylan Patel Dec 5, 2023 ▶ 14:40 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Feb 8, 2024 bullish
Disclosure
Prakash: Together AI operates a fleet of 7,000 to 8,000 GPUs
“We have close to seven to 8000 GPUs today. It's growing monthly.”
Vipul Ved Prakash Feb 8, 2024 ▶ 27:33 Building an open AI company - with Ce and Vipul of Together AI
Feb 8, 2024 neutral
Prediction Held up
Prakash predicts up to 5 million AI GPUs will sell in 2024
“There is four to five million GPUs that will be sold this year. NVIDIA and others.”
Vipul Ved Prakash Feb 8, 2024 ▶ 30:43 Building an open AI company - with Ce and Vipul of Together AI
Feb 19, 2024 bearish
Prediction Not checkable as stated
VCs Subsidizing AI Inference Will Not See Their Expected Returns
“In the end, like, I don't think VCs will have the return they expected. Like, you know, in these things, but guess who's going to benefit? Like, you know, it's the consumers, right? Like someone's like reaping that the value of this. And that's, I think an ama…”
Erik Bernhardsson Feb 19, 2024 ▶ 50:15 Truly Serverless Infra for AI Engineers - with Erik Bernhardsson of Modal
Feb 19, 2024 positive
Assertion Supported
Modal Can Fan Out Workloads to Thousands of GPUs in Minutes
“It's like pretty easy in Moto, like, fan out to like, you know, at least like a hundred GPUs, like in a few seconds, and you know, if you give it like a couple of minutes, like we can, you know, you can fan out to like thousands of GPUs.”
Erik Bernhardsson Feb 19, 2024 ▶ 42:23 Truly Serverless Infra for AI Engineers - with Erik Bernhardsson of Modal
Feb 28, 2024 neutral
Assertion Not checkable as stated
Firshman: GPU demand is not currently outpacing supply at Replicate
“From our point of view, demand is not outpacing supply of GPUs. Like we have enough, from our point of view, we have enough GPUs to go around, but that might change for sure.”
Ben Firshman Feb 28, 2024 ▶ 1:05:09 A Brief History of the Open Source AI Hacker - with Ben Firshman of Replicate
Mar 6, 2024 neutral
Opinion
Chintala: Time and data constrain Meta LLM releases more than GPUs
“So, I think the, it's all a matter of time. I think time is the biggest bottleneck. It's like, when do you stop training the previous one, and when do you start training the next one? And how do you make those decisions? The data, do you have net new data, bet…”
Soumith Chintala Mar 6, 2024 ▶ 46:46 Open Source AI is AI we can Trust — with Soumith Chintala of Meta AI
May 31, 2024 negative
Assertion Not checkable as stated
Huang: Berkeley's JAX Ring Attention does not work well on GPUs
“The Jaxx implementation just does not work on, on GPUs very well. Like, any naive setup that you do, like, it just won't run out of the box very easily”
Mark Huang May 31, 2024 ▶ 26:02 How to train a Million Context LLM — with Mark Huang of Gradient.ai
Jun 25, 2024
Assertion Supported
Albrecht: 4K GPU clusters require 3-tier networking versus standard 1K 2-tier setups
“The normal, the like vanilla setup or, you know, these large clusters as vanilla as it can be is what's normally like a 127 node cluster. So closer to like 10, 24 GPUs instead of 4000. Here we have a larger cluster. As you start to get into the larger clusters…”
Josh Albrecht Jun 25, 2024 ▶ 13:50 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Jun 25, 2024
Assertion Not checkable as stated
Frankle: Most AI data centers are retrofitted, not built for high heat
“In data centers that for the most part were not built remotely for this kind of power or heat and have been retrofitted for this. Like failures happen on a good day with normal CPUs. And this is not a good day and not a normal CPU for the most part.”
Jonathan Frankle Jun 25, 2024 ▶ 12:41 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Jul 29, 2024
Insight
Eugene Yan: GPU floating-point math makes temperature-zero inference non-deterministic
“For GPUs with floating points, and you push it through so many calculations, and so many met miles, the floating points aren't just not gonna be precise. So that's why even if temperature is zero, it's not gonna be the same throughout, ah, for multiple request…”
Eugene Yan Jul 29, 2024 ▶ 1:08:29 [LLM Paper Club] Llama 3.1 Paper: The Llama Family of Models
Oct 18, 2024 positive
Insight
Houston: Humans should act as CPUs orchestrating AI systems like GPUs
“Right now we have, like, the human CPU doing a lot of, you know, silicon CPU tasks, and so you really have to, like, redesign the work thoughtfully such that, you know, probably not that different from how it's evolved in computer architecture, where the CPU i…”
Drew Houston Oct 18, 2024 ▶ 47:22 Building the Silicon Brain - Drew Houston of Dropbox
Oct 19, 2024 negative
Assertion Supported
Hu: Single MLE-bench evaluation run with OpenAI o1-preview costs $4,000
“Just for one seed, For one run of these things cost 4000 dollars all in with the GPU plus the tokens. And a bulk of the cost was actually the token, so even if you cut the GPU out, it'll still cost you three grand to run on one preview.”
Jesse Hu Oct 19, 2024 ▶ 50:38 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
Oct 19, 2024 neutral
Assertion Supported
Hu: GPU setups showed virtually no agent performance gain over CPU-only
“They compared a CPU only setup to a GPU setup to a multi GPU setup, and it kind of made no difference really.”
Jesse Hu Oct 19, 2024 ▶ 49:40 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
Dec 7, 2024 negative
Assertion Supported
GPUs cannot handle unstructured sparsity as efficiently as Cerebras hardware
“So both cerebris and GPUs can handle structured sparsity But GPUs are not designed to handle unstructured sparsity, whereas what I've just mentioned before is able to handle this unstructured sparsity.”
Sarah Chieng Dec 7, 2024 ▶ 38:26 [Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras)
Dec 23, 2024
Insight
Soldani: Frontier LLM pre-training requires at least 50,000 GPUs
“To give you a sense of, like, how I personally think about research budget for each part of the language model pipeline is, like, on the pre-training side, you can maybe do something with a thousand GPUs. Really, you want 10,000. And, like, if you want real es…”
Luca Soldani Dec 23, 2024 ▶ 11:11 Best of 2024: Open Models [LS LIVE! at NeurIPS 2024]
Dec 23, 2024
Insight
Soldani: Replicating OpenAI's o1 requires roughly 10,000 GPUs
“If you're interested in you know, your, Open replication of what OpenAI's O-one is you're gonna be on the 10 K spectrum of our GPUs.”
Luca Soldani Dec 23, 2024 ▶ 12:08 Best of 2024: Open Models [LS LIVE! at NeurIPS 2024]
Apr 11, 2025 positive
Insight
Conrad: CoreWeave's debt-financed long-term contract model is optimal for GPUs
“So that means that the best way to make money in GPUs was to do basically exactly what CoreWeave did which is go out and sign only long-term contracts, pretty much ignore the bottom end of the market completely, and then maximize your long-term contracts with …”
Evan Conrad Apr 11, 2025 ▶ 4:48 SF Compute: Commoditizing Compute
Apr 11, 2025 bullish
Prediction Didn’t hold up
Conrad: GPU market will likely return to a shortage by winter
“My general prediction is that like by the winter we will be back towards shortage, but then also this very much depends on The rollout of future chips.”
Evan Conrad Apr 11, 2025 ▶ 30:48 SF Compute: Commoditizing Compute
Apr 11, 2025
Insight
Conrad: Software margins on GPU clusters drive customers to build in-house
“So if you have a 10% margin increase because you have great software on your billion dollars, the customers are that price sensitive. They will immediately switch off if they can, because why wouldn't you? You would just take that hundred million dollars, you'…”
Evan Conrad Apr 11, 2025 ▶ 4:28 SF Compute: Commoditizing Compute
Apr 11, 2025
Insight
Conrad: Incremental GPUs always drive model performance and revenue, unlike CPUs
“Gusto isn't going to make like, you know, five percent more money. They're going to make zero, like literally zero money from every incremental GPU or CPU after a certain point. This is not the case for anyone who is training models. And it's not the case for …”
Evan Conrad Apr 11, 2025 ▶ 3:08 SF Compute: Commoditizing Compute
Jun 13, 2025
Assertion Supported
CPU latency in KV cache management bottlenecks GPU utilization
“Like your eviction policy runs on a CPU. Like that radix hashing algorithm and block hashing and all that stuff happens like primarily CPU. That's really important for performance because if you have latency in these steps, like you're not keeping your GPU uti…”
Chris Lattner Jun 13, 2025 ▶ 1:00:09 The Shape of Compute (Chris Lattner of Modular)
Jul 2, 2025 neutral
Insight
Morris: Small models should be defined as runnable on a single GPU
“I think that we should establish the definition of small model as being a model that a grad student can inference at reasonable time on a single GPU. Which is probably like seven B maybe. I don't think 27 is small under any reasonable.”
Jack Morris Jul 2, 2025 ▶ 52:05 Information Theory for Language Models: Jack Morris
Jul 2, 2025 positive
Insight
Morris: Deep understanding of GPU architecture makes engineers exceptionally hireable
“That said, if you do it, you're, you've gotta be one of the most hireable people in the world. Like if you like, Really deeply understand the architecture of the new GPUs coming out and how to control it. You're in a very small handful of people and like every…”
Jack Morris Jul 2, 2025 ▶ 12:13 Information Theory for Language Models: Jack Morris
Jul 28, 2025 neutral
Insight
Mohan: GPU container sharing limitations leave hardware heavily idle
“For most people, one of the things about CPUs that's really nice is with containers, right? You can end up having a single node and you can place many containers on them and all the containers will slowly start eating the compute. It's not really the same with…”
Varun Mohan Jul 28, 2025 ▶ 5:13 🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
Jul 28, 2025 positive
Assertion Not checkable as stated
Fanelli: ExaFunction customer cut compute costs 97% on single GPU
“And I saw one of your customers, they went from. 30 clients to just one single GPU and they cut costs by 97%.”
Alessio Fanelli Jul 28, 2025 ▶ 6:28 🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
Jul 31, 2025 positive
Insight
Lambert: Top AI talent is dramatically cheaper than GPU clusters
“Talent is cheaper than GPUs by a dramatic margin, and At the end of the day, it's like, okay, if we're spending this much, they go to the room and they stare in the mirror and you're like, wait, it might not actually be that ridiculous to spend this money on t…”
Nathan Lambert Jul 31, 2025 ▶ 1:13:46 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Jul 31, 2025 negative
Assertion Not checkable as stated
Lambert: Long inference generations break RL infrastructure and require more GPUs
“The inference, high inference length generations definitely just, like, kind of breaks all infrastructure, because there's just so many tokens, there's more opportunity for out of memory or other things to go wrong. So it's like, just on a default, all of your…”
Nathan Lambert Jul 31, 2025 ▶ 1:00:39 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Oct 1, 2025 bullish
Assertion Supported
Feldman: Cerebras provides 2,625x more memory bandwidth than traditional GPUs
“And we have 2625 times more memory bandwidth than the GPU does.”
Andrew Feldman Oct 1, 2025 ▶ 6:22 ⚡️Raising $1.1b to build the fastest LLM Chips on Earth — Andrew Feldman, Cerebras
Oct 30, 2025 negative
Insight
Sands: High inference costs make friendly fraud existentially threatening for AI startups
“Now we're in the world where GPUs are expensive, inference costs are high, and free trial abuse or refund abuse or general non-payment abuse, right, you rack up these charges and you never pay, is like existentially threatening for AI businesses.”
Emily Glassberg Sands Oct 30, 2025 ▶ 10:46 The Agents Economy Backbone - with Emily Glassberg Sands, Head of Data & AI at Stripe
Nov 25, 2025 neutral
Assertion Not checkable as stated
Johnson: Academic labs can no longer train state-of-the-art AI on few GPUs
“Like five or 10 years ago, you really could train state-of-the-art models in the lab even with just a couple of GPUs. But, you know, because that technology was so successful and scaled up so much, then you can't train state-of-the-art models with a couple of …”
Justin Johnson Nov 25, 2025 ▶ 9:51 After LLMs: Spatial Intelligence and World Models — Fei-Fei Li & Justin Johnson, World Labs
Nov 25, 2025 neutral
Assertion Supported
Johnson: AI compute per model has scaled one million-fold since 2012
“And if you think about, you know, AlexNet required this jump from CPUs to GPUs, but even from AlexNet to today, we're getting about a thousand times more performance per card than we had in AlexNet days. And now it's common to train models, not just on one GPU…”
Justin Johnson Nov 25, 2025 ▶ 4:13 After LLMs: Spatial Intelligence and World Models — Fei-Fei Li & Justin Johnson, World Labs
Dec 18, 2025 positive
Assertion Supported
Zhang: SAM 3 achieves real-time tracking across objects via multi-GPU parallelism
“Even for video, if you can't afford the kind of GPUs, pretty many, very kind of, do the kind of parallel inference algorithm. So even you have a lot of object to track, you can still get real-time tracking performance as long as you scale up the GPUs there.”
Pengchuan Zhang Dec 18, 2025 ▶ 9:50 SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow)
Dec 31, 2025 bullish
Insight
LLMs are commoditizing like raw compute, shifting value to abstraction layers
“Language models themselves are more like compute or GPU a generation ago, where what can we build at the layer above? And in software systems, we've traditionally thought of VMware being a great example. You have the operating system and the underlying archite…”
Andy Konwinski Dec 31, 2025 ▶ 7:04 [State of Research Funding] Beyond NSF, Slingshots, Open Frontiers — Andy Konwinski, Laude Institute
Jan 28, 2026 positive
Opinion
Anderson: $1M+ salaries can make sense when GPU spend dominates burn
“In terms of like total spend on GPUs, it can still be a total, a small fraction of your burn. So sometimes it kind of makes sense.”
Brandon Anderson Jan 28, 2026 ▶ 53:14 🔬 From Red Teaming GPT-4 to Automating Drug Discovery: The Future of AI in Science — Andrew White
Feb 19, 2026
Insight
Wang: AI startups face a dilemma balancing AGI research with product revenue
“I think the best researchers in the world have this dilemma of, okay, I want to go all in on AGI, but it's the product usage revenue flywheel that keeps the revenue in the house to power all the GPUs to get to AGI. And so it does make you know, I think it sets…”
Sarah Wang Feb 19, 2026 ▶ 13:09 Inside AI’s $10B+ Capital Flywheel — Martin Casado & Sarah Wang of a16z
Feb 19, 2026 bullish
Assertion Not checkable as stated
Casado: There is no compute supply overhang or 'dark GPUs'
“But we don't have a supply overhang. Like, there's no dark GPUs, right?”
Martin Casado Feb 19, 2026 ▶ 4:13 Inside AI’s $10B+ Capital Flywheel — Martin Casado & Sarah Wang of a16z
Feb 25, 2026 neutral
Assertion Not checkable as stated
Welling: Semiconductor Scaling Limits Require Materials Innovation for Hardware Advances
“More or less we've reached the limits of You know, scaling things down, and now we are trying to improve further by new materials, so that's the fundamental materials problem.”
Max Welling Feb 25, 2026 ▶ 12:03 🔬Max Welling: Materials Underlie Everything
Feb 26, 2026 bearish
Prediction Open · timeframe Dec 2027
Patel: Google will buy tons of GPUs through 2027 due to TPU limits
“When we look in 26, Google would buy a lot more TPUs, but they can't ramp production fast enough, right? And so they have to buy tons of GPUs. And we go to 27, it applies again, right? Google simply cannot buy enough TPUs, and they have to buy tons of GPUs.”
Dylan Patel Feb 26, 2026 ▶ 49:51 Dylan Patel Explains the AI War While Cooking | In-Context Cooking
Mar 24, 2026 bearish
Assertion Not checkable as stated
Unnamed materials foundation model is only 5x faster than DFT and unreliable
“It's only in my hands the one I'm still not naming is only about five times faster than my fastest DFT calculation on a GPU, and it also doesn't work all the time.”
Heather Kulik Mar 24, 2026 ▶ 20:32 🔬There Is No AlphaFold for Materials — AI for Materials Discovery with Heather Kulik
Apr 7, 2026 neutral
Insight
Lopopolo: Synchronous human attention is the only scarce resource in agentic software engineering
“The model is trivially paralyzable, right? As many GPUs and tokens as I am willing to spend, I can have capacity to work with a code base. The only fundamentally scarce thing is the synchronous human attention of my team.”
Ryan Lopopolo Apr 7, 2026 ▶ 9:38 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
May 21, 2026 neutral
Insight
Burazin: CPU environments must spin up instantly to prevent costly GPU idle time
“The reason why a lot of people come to us is because GPUs are more expensive than CPUs, right? So you want your GPU running at what? A hundred percent the entire time. And so when you're running runs on CPUs, when the CPU cycle is like down and spinning up the…”
Ivan Burazin May 21, 2026 ▶ 27:41 AI Agents Need Computers: 74% MoM Growth, 850K/Day Runs, & New Agent Cloud — Ivan Burazin, Daytona
May 24, 2026 positive
Assertion Supported
Gemma 4 E2B loads only 2B of 5B parameters into GPU
“So the GEMA for model is a E to B. That means that it effectively has two billion parameters loaded into the GPU. It actually has almost five billion parameters, but those three billion parameters can be in the CPU, they can be in the disk, which means that yo…”
Omar Sanseviero May 24, 2026 ▶ 0:52 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Jul 8, 2026
Insight
Bubna: Transferring RL weights is fundamentally an OS memory problem
“Like the way you move around your KV cache and how efficiently you can do it, how efficiently you move your weights from your training GPUs to your inference GPUs in RL is, there's a lot of degrees of freedom, and it is basically a systems problem of Moving me…”
Akshat Bubna Jul 8, 2026 ▶ 31:40 The Future of AI Infra: from Kubernetes to Agent Sandboxes — Akshat Bubna, Modal CTO
Jul 8, 2026
Assertion Not checkable as stated
Modal CTO: Production scale requires elastically scaling 1,000 to 1,500 GPUs quickly
“There it's not about scaling from zero to one, but it's how do we scale really elastically from, like, thousand to 1500 GPUs very quickly in, in a given region.”
Akshat Bubna Jul 8, 2026 ▶ 12:44 The Future of AI Infra: from Kubernetes to Agent Sandboxes — Akshat Bubna, Modal CTO
Jul 8, 2026
Assertion Not checkable as stated
Bubna: Modal's custom reliability layer insulates users from GPU hardware drops
“That's why it's something we've invested a lot of time in is actually building our own reliability layer on top. So if the GPU falls off the bus or something happens, we user workloads are not affected.”
Akshat Bubna Jul 8, 2026 ▶ 26:27 The Future of AI Infra: from Kubernetes to Agent Sandboxes — Akshat Bubna, Modal CTO
Jul 16, 2026 neutral
Assertion Supported
Beam: Reinforcement learning achieves only 5% to 6% GPU FLOP utilization
“And for reinforcement learning, it's always somewhere, like, around five to, like, six percent. So, said differently, that means that we're getting, like, five percent of the actual GPU computing power that we're paying for.”
Andy Beam Jul 16, 2026 ▶ 1:38:42 🔬 RL with Verifiable Rewards, but the Verifier is a Lab — Lila Sciences
Aug 3, 2026
Assertion Contradicted
All modern AI models requiring multi-GPU parallelization are Mixture-of-Experts
“Effectively, all models today are MOE models that are, you know, at least all models large enough that you would care to parallelize them across multiple GPUs.”
Philip Kiely Aug 3, 2026 ▶ 52:42 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Aug 11, 2026 negative
Assertion Supported
McPartlon: Triangle layers run inefficiently on modern GPU architectures
“These layers are pretty costly and like that kind of limits what you can do with the architectures. They're not like, Not only are they like costly in terms of compute, they're just like not efficient on modern GPUs either. You have small hidden dimensions, la…”
Matt McPartlon Aug 11, 2026 ▶ 1:12:11 🔬They Thought the Model Was Broken — Matt McPartlon & Neil Patil, Chai Discovery
Aug 26, 2026 bullish
Assertion Supported
FourCastNet matches supercomputer weather accuracy 10,000 times faster on consumer GPUs
“To our surprise, we found that it's not only, you know, accurate, it's almost as close to the what the traditional weather models can do accurately, but also tens of thousands of times faster. So what would take a big supercomputer to run can now be run. And w…”
Anima Anandkumar Aug 26, 2026 ▶ 33:56 🔬 Why Transformers Hit a Wall the Moment Physics Shows Up — Anima Anandkumar, Caltech
Aug 26, 2026 neutral
Assertion Supported
Lean faces CPU-bound scalability limits for verifying large neural networks
“So lean still has a lot of shortcomings there. It's CPU based and you know, it's not, Like, getting that onto the GPU has a lot of nuances there. So, you know, a lot of work needs to be done. So what we've started with is a framework, you know, making that mor…”
Anima Anandkumar Aug 26, 2026 ▶ 10:31 🔬 Why Transformers Hit a Wall the Moment Physics Shows Up — Anima Anandkumar, Caltech
Sep 2, 2026 bearish
Opinion
Sean Lie: Etched is not building anything better than a traditional GPU
“I think my reaction when I see pictures like this is that it's very impressive graphics design. But I also don't see them building anything beyond just, or trying to build something better than just, you know, a traditional GPU, right? You know, they've made c…”
Sean Lie Sep 2, 2026 ▶ 34:23 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Sep 2, 2026 bullish
Assertion Supported
Lie: Cerebras runs OpenAI's flagship model 14x faster than GPUs
“We're running you know, frontier level, one of the most intelligent models, right? OpenAI's largest, most capable, most intelligent model right now at 14 times faster than their normal, you know, GPU speeds.”
Sean Lie Sep 2, 2026 ▶ 8:37 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.