Sep 7, 2023 · 30m · no-priors

No Priors Ep. 31 | With Cerebras CEO Andrew Feldman

Andrew Feldman · 21m spoken Elad Gil · 3m spoken Sarah Guo · 2m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of No Priors, hosts Elad Gil and Sarah Guo interview Cerebras Systems CEO Andrew Feldman to discuss wafer-scale semiconductor architecture, solutions to the global GPU compute crunch, and the evolving economic realities of training and deploying foundation AI models.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 19.1% of the talking time here. How this is scored →

The hosts as informed peer 3.6 Guest teaching 4.6 Guest disagreement 1.9 The hosts pushing back 0.7
05100:0010:0020:0030:000:47–4:27 · The hosts as informed peer 4/10 The One Hundred Million Dollar Cerebras G42 Partnership The conversation opens warmly with Sarah reminiscing about Cerebras's earliest 2015 pitch meetings in her office. Andrew details the initial 36-exaflop G42 partnership and early architectural bets on wafer-scale chips. The dynamic is collaborative and celebratory rather than adversarial.4:27–8:24 · The hosts as informed peer 3/10 Real-World Training Efficiency versus Synthetic Hardware Benchmarks When asked about performance benchmarks, Andrew rejects synthetic benchmarks entirely, citing his background at AMD and Sun where entire teams were dedicated to gaming them. He educates the hosts on real-world training times, data parallelism, and how wafer-scale SRAM avoids GPU distributed compute overhead.8:24–11:43 · The hosts as informed peer 3/10 Disaggregating Compute and Memory to Outperform Standard GPUs Sarah prompts Andrew to explain the architectural differences between traditional GPUs and Cerebras hardware for large-scale ML. Andrew explains how coupling memory to compute on GPU interposers forces inefficient hardware buying, whereas disaggregating memory allows running arbitrary parameter sizes.11:44–15:30 · The hosts as informed peer 5/10 Multilingual Foundation Models, Cultural Nuance, and Tokenization Andrew highlights serving costs and multilingual models, particularly Arabic and low-resource languages. Elad and Sarah engage deeply on cultural encoding, byte tokenization biases, and alignment mechanics in pre-training datasets.15:30–17:49 · The hosts as informed peer 3/10 Engineering and Financial Realities of Starting Chip Companies Sarah asks about the structural hurdles of founding semiconductor startups. Andrew delivers a comprehensive breakdown of the sixty-million-dollar tapeout cycle, zero-margin-for-error QA ratios, and why hardware cannot operate on rapid SaaS feedback loops.17:49–23:26 · The hosts as informed peer 5/10 The AI Accelerator Landscape and Challenging NVIDIA Dominance Andrew pulls no punches regarding Nvidia, accusing them of exploiting monopoly conditions and extorting customers. Sarah probes Andrew on whether general sparse linear algebra can outlast specialized transformer architectures or separate training and inference silicon.23:27–26:09 · The hosts as informed peer 5/10 Supply Chain Inflexibility and the Ongoing GPU Shortage Andrew outlines the inflexible economics of semiconductor fabs, noting that TSMC twenty-billion-dollar fabs cannot pivot quickly when forecasts are blown. Elad adds industry context by noting Nvidia's previous inventory write-downs during the crypto market crash.26:10–29:42 · The hosts as informed peer 4/10 Economic Limits of Foundation Model Scaling and Deployment The hosts and guest discuss economic scaling limits, highlighting that inference costs will keep production models in the 3B-13B parameter range. Sarah and Andrew discuss high-value enterprise proprietary datasets from companies like Bloomberg and Reuters as the next major differentiator.29:44–30:00 · The hosts as informed peer 0/10 Podcast Conclusion, Social Channels, and Episode Transcripts Sarah delivers the closing housekeeping notes, social media channels, and transcript website links.0:47–4:27 · Guest teaching 3/10 The One Hundred Million Dollar Cerebras G42 Partnership The conversation opens warmly with Sarah reminiscing about Cerebras's earliest 2015 pitch meetings in her office. Andrew details the initial 36-exaflop G42 partnership and early architectural bets on wafer-scale chips. The dynamic is collaborative and celebratory rather than adversarial.4:27–8:24 · Guest teaching 6/10 Real-World Training Efficiency versus Synthetic Hardware Benchmarks When asked about performance benchmarks, Andrew rejects synthetic benchmarks entirely, citing his background at AMD and Sun where entire teams were dedicated to gaming them. He educates the hosts on real-world training times, data parallelism, and how wafer-scale SRAM avoids GPU distributed compute overhead.8:24–11:43 · Guest teaching 6/10 Disaggregating Compute and Memory to Outperform Standard GPUs Sarah prompts Andrew to explain the architectural differences between traditional GPUs and Cerebras hardware for large-scale ML. Andrew explains how coupling memory to compute on GPU interposers forces inefficient hardware buying, whereas disaggregating memory allows running arbitrary parameter sizes.11:44–15:30 · Guest teaching 4/10 Multilingual Foundation Models, Cultural Nuance, and Tokenization Andrew highlights serving costs and multilingual models, particularly Arabic and low-resource languages. Elad and Sarah engage deeply on cultural encoding, byte tokenization biases, and alignment mechanics in pre-training datasets.15:30–17:49 · Guest teaching 7/10 Engineering and Financial Realities of Starting Chip Companies Sarah asks about the structural hurdles of founding semiconductor startups. Andrew delivers a comprehensive breakdown of the sixty-million-dollar tapeout cycle, zero-margin-for-error QA ratios, and why hardware cannot operate on rapid SaaS feedback loops.17:49–23:26 · Guest teaching 5/10 The AI Accelerator Landscape and Challenging NVIDIA Dominance Andrew pulls no punches regarding Nvidia, accusing them of exploiting monopoly conditions and extorting customers. Sarah probes Andrew on whether general sparse linear algebra can outlast specialized transformer architectures or separate training and inference silicon.23:27–26:09 · Guest teaching 6/10 Supply Chain Inflexibility and the Ongoing GPU Shortage Andrew outlines the inflexible economics of semiconductor fabs, noting that TSMC twenty-billion-dollar fabs cannot pivot quickly when forecasts are blown. Elad adds industry context by noting Nvidia's previous inventory write-downs during the crypto market crash.26:10–29:42 · Guest teaching 4/10 Economic Limits of Foundation Model Scaling and Deployment The hosts and guest discuss economic scaling limits, highlighting that inference costs will keep production models in the 3B-13B parameter range. Sarah and Andrew discuss high-value enterprise proprietary datasets from companies like Bloomberg and Reuters as the next major differentiator.29:44–30:00 · Guest teaching 0/10 Podcast Conclusion, Social Channels, and Episode Transcripts Sarah delivers the closing housekeeping notes, social media channels, and transcript website links.0:47–4:27 · Guest disagreement 1/10 The One Hundred Million Dollar Cerebras G42 Partnership The conversation opens warmly with Sarah reminiscing about Cerebras's earliest 2015 pitch meetings in her office. Andrew details the initial 36-exaflop G42 partnership and early architectural bets on wafer-scale chips. The dynamic is collaborative and celebratory rather than adversarial.4:27–8:24 · Guest disagreement 4/10 Real-World Training Efficiency versus Synthetic Hardware Benchmarks When asked about performance benchmarks, Andrew rejects synthetic benchmarks entirely, citing his background at AMD and Sun where entire teams were dedicated to gaming them. He educates the hosts on real-world training times, data parallelism, and how wafer-scale SRAM avoids GPU distributed compute overhead.8:24–11:43 · Guest disagreement 2/10 Disaggregating Compute and Memory to Outperform Standard GPUs Sarah prompts Andrew to explain the architectural differences between traditional GPUs and Cerebras hardware for large-scale ML. Andrew explains how coupling memory to compute on GPU interposers forces inefficient hardware buying, whereas disaggregating memory allows running arbitrary parameter sizes.11:44–15:30 · Guest disagreement 1/10 Multilingual Foundation Models, Cultural Nuance, and Tokenization Andrew highlights serving costs and multilingual models, particularly Arabic and low-resource languages. Elad and Sarah engage deeply on cultural encoding, byte tokenization biases, and alignment mechanics in pre-training datasets.15:30–17:49 · Guest disagreement 2/10 Engineering and Financial Realities of Starting Chip Companies Sarah asks about the structural hurdles of founding semiconductor startups. Andrew delivers a comprehensive breakdown of the sixty-million-dollar tapeout cycle, zero-margin-for-error QA ratios, and why hardware cannot operate on rapid SaaS feedback loops.17:49–23:26 · Guest disagreement 4/10 The AI Accelerator Landscape and Challenging NVIDIA Dominance Andrew pulls no punches regarding Nvidia, accusing them of exploiting monopoly conditions and extorting customers. Sarah probes Andrew on whether general sparse linear algebra can outlast specialized transformer architectures or separate training and inference silicon.23:27–26:09 · Guest disagreement 2/10 Supply Chain Inflexibility and the Ongoing GPU Shortage Andrew outlines the inflexible economics of semiconductor fabs, noting that TSMC twenty-billion-dollar fabs cannot pivot quickly when forecasts are blown. Elad adds industry context by noting Nvidia's previous inventory write-downs during the crypto market crash.26:10–29:42 · Guest disagreement 1/10 Economic Limits of Foundation Model Scaling and Deployment The hosts and guest discuss economic scaling limits, highlighting that inference costs will keep production models in the 3B-13B parameter range. Sarah and Andrew discuss high-value enterprise proprietary datasets from companies like Bloomberg and Reuters as the next major differentiator.29:44–30:00 · Guest disagreement 0/10 Podcast Conclusion, Social Channels, and Episode Transcripts Sarah delivers the closing housekeeping notes, social media channels, and transcript website links.0:47–4:27 · The hosts pushing back 0/10 The One Hundred Million Dollar Cerebras G42 Partnership The conversation opens warmly with Sarah reminiscing about Cerebras's earliest 2015 pitch meetings in her office. Andrew details the initial 36-exaflop G42 partnership and early architectural bets on wafer-scale chips. The dynamic is collaborative and celebratory rather than adversarial.4:27–8:24 · The hosts pushing back 1/10 Real-World Training Efficiency versus Synthetic Hardware Benchmarks When asked about performance benchmarks, Andrew rejects synthetic benchmarks entirely, citing his background at AMD and Sun where entire teams were dedicated to gaming them. He educates the hosts on real-world training times, data parallelism, and how wafer-scale SRAM avoids GPU distributed compute overhead.8:24–11:43 · The hosts pushing back 0/10 Disaggregating Compute and Memory to Outperform Standard GPUs Sarah prompts Andrew to explain the architectural differences between traditional GPUs and Cerebras hardware for large-scale ML. Andrew explains how coupling memory to compute on GPU interposers forces inefficient hardware buying, whereas disaggregating memory allows running arbitrary parameter sizes.11:44–15:30 · The hosts pushing back 1/10 Multilingual Foundation Models, Cultural Nuance, and Tokenization Andrew highlights serving costs and multilingual models, particularly Arabic and low-resource languages. Elad and Sarah engage deeply on cultural encoding, byte tokenization biases, and alignment mechanics in pre-training datasets.15:30–17:49 · The hosts pushing back 0/10 Engineering and Financial Realities of Starting Chip Companies Sarah asks about the structural hurdles of founding semiconductor startups. Andrew delivers a comprehensive breakdown of the sixty-million-dollar tapeout cycle, zero-margin-for-error QA ratios, and why hardware cannot operate on rapid SaaS feedback loops.17:49–23:26 · The hosts pushing back 2/10 The AI Accelerator Landscape and Challenging NVIDIA Dominance Andrew pulls no punches regarding Nvidia, accusing them of exploiting monopoly conditions and extorting customers. Sarah probes Andrew on whether general sparse linear algebra can outlast specialized transformer architectures or separate training and inference silicon.23:27–26:09 · The hosts pushing back 1/10 Supply Chain Inflexibility and the Ongoing GPU Shortage Andrew outlines the inflexible economics of semiconductor fabs, noting that TSMC twenty-billion-dollar fabs cannot pivot quickly when forecasts are blown. Elad adds industry context by noting Nvidia's previous inventory write-downs during the crypto market crash.26:10–29:42 · The hosts pushing back 1/10 Economic Limits of Foundation Model Scaling and Deployment The hosts and guest discuss economic scaling limits, highlighting that inference costs will keep production models in the 3B-13B parameter range. Sarah and Andrew discuss high-value enterprise proprietary datasets from companies like Bloomberg and Reuters as the next major differentiator.29:44–30:00 · The hosts pushing back 0/10 Podcast Conclusion, Social Channels, and Episode Transcripts Sarah delivers the closing housekeeping notes, social media channels, and transcript website links.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 34.1% · guest 65.9%0:00 · the hosts 34.1% · guest 65.9%3:00 · the hosts 22% · guest 78%3:00 · the hosts 22% · guest 78%6:00 · the hosts 6.8% · guest 93.2%6:00 · the hosts 6.8% · guest 93.2%9:00 · the hosts 10.9% · guest 89.1%9:00 · the hosts 10.9% · guest 89.1%12:00 · the hosts 24.6% · guest 75.4%12:00 · the hosts 24.6% · guest 75.4%15:00 · the hosts 14.2% · guest 85.8%15:00 · the hosts 14.2% · guest 85.8%18:00 · the hosts 10.5% · guest 89.5%18:00 · the hosts 10.5% · guest 89.5%21:00 · the hosts 21% · guest 79%21:00 · the hosts 21% · guest 79%24:00 · the hosts 26.6% · guest 73.4%24:00 · the hosts 26.6% · guest 73.4%27:00 · the hosts 21.5% · guest 78.5%27:00 · the hosts 21.5% · guest 78.5%30:00 · the hosts 100% · guest 0%30:00 · the hosts 100% · guest 0%
Sharpest disagreement ▶ 18:04 Andrew attacks Nvidia's customer pricing and supply practices

Andrew aggressively frames Nvidia's market dominance as extortionate pricing and artificial dependency, drawing analogies to historic customer resentment against Intel.

Hardest push from the hosts ▶ 21:11 Sarah presses on internal bets between training and inference

Sarah directly questions Andrew on whether Cerebras is hedging its core architecture by making internal bets to separate training silicon from generative inference silicon.

Biggest teaching moment ▶ 15:46 Masterclass on chip development capital expenditure vs SaaS

Andrew educates the audience and hosts on the brutal economic reality of chip manufacturing, contrasting fast SaaS prototyping with multi-year, multi-million dollar tapeout cycles.

The host holds their own ▶ 25:39 Elad connects fab forecasting errors to Nvidia's crypto inventory crash

Elad brings in sharp market historical context, highlighting how Nvidia suffered severe supply over-allocation during the prior crypto downturn to support Andrew's forecasting argument.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
The One Hundred Million Dollar Cerebras G42 Partnership 4310 The conversation opens warmly with Sarah reminiscing about Cerebras's earliest 2015 pitch meetings in her office. Andrew details the initial 36-exaflop G42 partnership and early architectural bets on wafer-scale chips. The dynamic is collaborative and celebratory rather than adversarial.
Real-World Training Efficiency versus Synthetic Hardware Benchmarks 3641 When asked about performance benchmarks, Andrew rejects synthetic benchmarks entirely, citing his background at AMD and Sun where entire teams were dedicated to gaming them. He educates the hosts on real-world training times, data parallelism, and how wafer-scale SRAM avoids GPU distributed compute overhead.
Disaggregating Compute and Memory to Outperform Standard GPUs 3620 Sarah prompts Andrew to explain the architectural differences between traditional GPUs and Cerebras hardware for large-scale ML. Andrew explains how coupling memory to compute on GPU interposers forces inefficient hardware buying, whereas disaggregating memory allows running arbitrary parameter sizes.
Multilingual Foundation Models, Cultural Nuance, and Tokenization 5411 Andrew highlights serving costs and multilingual models, particularly Arabic and low-resource languages. Elad and Sarah engage deeply on cultural encoding, byte tokenization biases, and alignment mechanics in pre-training datasets.
Engineering and Financial Realities of Starting Chip Companies 3720 Sarah asks about the structural hurdles of founding semiconductor startups. Andrew delivers a comprehensive breakdown of the sixty-million-dollar tapeout cycle, zero-margin-for-error QA ratios, and why hardware cannot operate on rapid SaaS feedback loops.
The AI Accelerator Landscape and Challenging NVIDIA Dominance 5542 Andrew pulls no punches regarding Nvidia, accusing them of exploiting monopoly conditions and extorting customers. Sarah probes Andrew on whether general sparse linear algebra can outlast specialized transformer architectures or separate training and inference silicon.
Supply Chain Inflexibility and the Ongoing GPU Shortage 5621 Andrew outlines the inflexible economics of semiconductor fabs, noting that TSMC twenty-billion-dollar fabs cannot pivot quickly when forecasts are blown. Elad adds industry context by noting Nvidia's previous inventory write-downs during the crypto market crash.
Economic Limits of Foundation Model Scaling and Deployment 4411 The hosts and guest discuss economic scaling limits, highlighting that inference costs will keep production models in the 3B-13B parameter range. Sarah and Andrew discuss high-value enterprise proprietary datasets from companies like Bloomberg and Reuters as the next major differentiator.
Podcast Conclusion, Social Channels, and Episode Transcripts 0000 Sarah delivers the closing housekeeping notes, social media channels, and transcript website links.

Statements from this episode (25)

Disclosure
Feldman: Cerebras is building nine AI supercomputers totaling 36 exaflops with G42
“We did announce a strategic partnership with a group called G-forty-two, and we announced that we were building nine supercomputers. Each supercomputer would be four exaflops of AI compute, so in total, 36 exaflops of AI compute.”
Andrew Feldman Sep 7, 2023 ▶ 0:55
Assertion Not checkable as stated
Feldman: Cerebras received eight term sheets from eight pitches
“Early, early in, we went out, we did eight pitches. We got eight term sheets.”
Andrew Feldman Sep 7, 2023 ▶ 3:16
Assertion Not checkable as stated
Feldman: Sun Microsystems had 70 people dedicated to gaming benchmarks
“When our CTO was at Sun, they had a team of 70 whose job it was to gain benchmarks.”
Andrew Feldman Sep 7, 2023 ▶ 4:44
Assertion Supported
Feldman: Cerebras hardware runs strictly data parallel without complex distributed engineering
“And if you have to spend months doing distributed compute, doing tensor model parallel distributed compute, I mean, if you look at the back of some of these papers, they're crediting 20 or 30 people, sometimes more, who helped on the distributed compute. And i…”
Andrew Feldman Sep 7, 2023 ▶ 5:31
Assertion Supported
Feldman: Cerebras chip contains about 850,000 processor-memory tiles
“It's a data flow architecture. It's comprised of about 850,000 identical tiles. Each tile is a processor and memory, and so it's a fully distributed memory machine, and so you have huge amounts of memory bandwidth because the memory is speaking to A processor …”
Andrew Feldman Sep 7, 2023 ▶ 6:12
Assertion Not checkable as stated
Feldman: Cerebras eliminates memory bandwidth bottlenecks using on-wafer SRAM
“We keep a huge amount of SRAM on the wafer. All right, and so there are no memory bandwidth problems ever. That also allows us to harvest sparsity, which is something that others really struggle with.”
Andrew Feldman Sep 7, 2023 ▶ 7:30
Assertion Supported
Feldman: Cerebras clusters scale linearly up to 64 nodes
“The cluster we build keeps the parameters off chip in a parameter store, and it streams them in, and the result of this architecture is that, that we run strictly data parallel, which means even in a 64 node cluster, you run the exact same configuration on eac…”
Andrew Feldman Sep 7, 2023 ▶ 7:45
Assertion Supported
Feldman: Cerebras can run a trillion-parameter model on a single system
“And that was an idea that came from supercomputing that we knew really well, that we could organize this so you could run an arbitrarily large, now a trillion parameter network on a single system.”
Andrew Feldman Sep 7, 2023 ▶ 9:06
Disclosure
Feldman: Cerebras and G42 to open-source largest Arabic LLM
“On Wednesday, we're announcing that we're open sourcing with G-forty-two's group called Inception and NBZ-UAI University, the largest Arabic LLM.”
Andrew Feldman Sep 7, 2023 ▶ 10:21
Assertion Not checkable as stated
Feldman: Redistributing AI training takes one keystroke on Cerebras vs GPUs
“In March, we put seven GPT models in the open source community. Everybody else was putting one. Why? Because it's really hard to redistribute work across a GPU cluster. For us, it's one keystroke.”
Andrew Feldman Sep 7, 2023 ▶ 11:13
Disclosure
Feldman: Cerebras is training multiple 30B-to-175B parameter AI models
“We are training right now a whole collection in the 30 to one 75 category.”
Andrew Feldman Sep 7, 2023 ▶ 12:19
Insight
Feldman: Poor tokenization schemes heavily bias foundation models toward English
“One of the things we learned was that it's really important taking into account The language that has fewer tokens, especially, you know, Arabic, Hebrew, Hindi, their characters sometimes require two bits rather than one bit, and so the tokenization scheme has…”
Andrew Feldman Sep 7, 2023 ▶ 14:44
Insight
Building a first chip takes two years and $50 million, says Feldman
“We're gonna spend two, two and a half, sometimes three years and 50 or sixty million dollars building our first part. And then we will learn if people like it or hate it. It has a huge gaping chasm between when you start and when you begin learning from custom…”
Andrew Feldman Sep 7, 2023 ▶ 16:03
Insight
Hardware bugs cost $20 million to fix, requiring huge QA ratios
“Often in chip companies for each developer, you have three or four QA people, DV people, and that's because a bug can cost you 20 or thirty million dollars. It's not like, yeah, I'll fix that in, in four dot one dot one. You're gonna have to re-spin the chip i…”
Andrew Feldman Sep 7, 2023 ▶ 16:51
Disclosure
Feldman: Cerebras is About 75% Software
“Now, we still have to build a huge amount of software, and we're about 75% software.”
Andrew Feldman Sep 7, 2023 ▶ 17:20
Opinion
NVIDIA is extorting customers amid AI chip shortages, says Cerebras CEO
“I think they're now in a situation where they're extorting customers. They're extremely expensive. They're unable to ship. And that has, among other things, opened the door for many of us who have alternatives.”
Andrew Feldman Sep 7, 2023 ▶ 18:10
Assertion Not checkable as stated
Feldman: Cerebras converged a model in 3.5 days after 60-day GPU failure
“We had a situation where they were trying to train on a GPU cluster and they were at 60 days and it wasn't converging and We stood it up, and three and a half days later, their model converged”
Andrew Feldman Sep 7, 2023 ▶ 19:04
Prediction Not checkable as stated
Feldman: Chip market will use separate silicon for training and inference
“Now, whether you will have different silicon for training and for inference, I think you will.”
Andrew Feldman Sep 7, 2023 ▶ 21:06
Insight
Feldman: Chip architects must solve hard problems generally, not guess layers
“The trick in architecture is to solve hard problems in a general way, so you don't have to rely on product management to sort of guess what, what's the next cool layer type, right?”
Andrew Feldman Sep 7, 2023 ▶ 21:24
Prediction Not checkable as stated
AI shift to single-shot learning would doom NVIDIA and Cerebras hardware
“If we go to a type of model that doesn't require very much data, if we go to single-shot learning, right, NVIDIA's totally out of luck, right? Us too, everybody.”
Andrew Feldman Sep 7, 2023 ▶ 21:38
Assertion Supported
Running inference on an eight-GPU H100 system costs half a million dollars
“I mean, people using. Eight h, 100 to do inference on a big model. I mean, that's. Half a million dollars.”
Andrew Feldman Sep 7, 2023 ▶ 23:19
Insight
Feldman: The chip industry has a profoundly inflexible supply chain
“The chip market has a profoundly inflexible supply chain.”
Andrew Feldman Sep 7, 2023 ▶ 23:49
Assertion Not checkable as stated
Feldman: Nvidia missed its own demand forecast ahead of AI crunch
“It's not just that Wall Street missed what NVIDIA would sell. NVIDIA missed it. They missed the forecast.”
Andrew Feldman Sep 7, 2023 ▶ 24:35
Assertion Contradicted
Feldman: Arista has 52-week lead times on switches
“I mean, Arista today is at 52 week lead times on switches.”
Andrew Feldman Sep 7, 2023 ▶ 26:04
Prediction Not checkable as stated
Feldman: Commercial AI will settle on 3B to 13B parameter models
“And so we're going to be down at three and at six billion and at thirteen billion, because that's, I can get pretty good, pretty darn good, and not break the bank with free inference.”
Andrew Feldman Sep 7, 2023 ▶ 27:28
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.