Mar 8, 2026 · 1h 25m · latent-space

Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup

Kyle Kranen · 30m spoken Nader Khalil · 25m spoken Shawn Wang · 12m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

NVIDIA engineering leader Kyle and Brev founder Nader Khalil join Latent Space to discuss NVIDIA's first-principles engineering culture, data center-scale inference with Dynamo, and the evolving architectures, interfaces, and security boundaries of autonomous AI agents.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 16.6% of the talking time here. How this is scored →

The hosts as informed peer 6.3 Guest teaching 5.7 Guest disagreement 1.9 The hosts pushing back 3.3
05100:0020:0040:001:00:001:20:000:00–9:32 · The hosts as informed peer 4/10 The Fundamental Security Dilemma of Autonomous AI Agents Swyx and Nader share nostalgic startup stories about Brev's early GTC marketing stunts and its acquisition by NVIDIA. The tone is highly collaborative and warm, with Swyx asking about product design choices and cloud hardware abstractions.9:33–19:12 · The hosts as informed peer 6/10 NVIDIA Developer Experience and the Speed of Light Philosophy Swyx probes into NVIDIA's internal 'Speed of Light' (SOL) operating philosophy, questioning whether anyone besides Jensen Huang can invoke it without derailing stability. Kyle and Nader explain how SOL functions as a first-principles physics baseline rather than mere pressure.19:12–27:17 · The hosts as informed peer 6/10 Recommenders, Graph Neural Networks, and Zero Billion Dollar Markets Kyle outlines his journey from recommendation systems (DLRM, Wide & Deep) and GNNs to modern LLMs, explaining NVIDIA's embrace of 'zero billion dollar markets'. Swyx playfully challenges whether the automotive market qualifies as zero billion, prompting clarification on emerging versus mature bets.27:17–44:35 · The hosts as informed peer 7/10 Architecting NVIDIA Dynamo for Data Center Scale Inference Kyle delivers a technical breakdown of NVIDIA Dynamo, distinguishing scale-up versus scale-out limits, NVLink vs InfiniBand interconnect bandwidths, and pre-fill versus decode disaggregation. Swyx and Vibo actively contribute relevant terminology and architectural context.44:36–54:39 · The hosts as informed peer 7/10 Context Window Limits, Hardware Co-Design, and Architectural Unhobblers Swyx pushes back against current context window scaling trajectories, arguing the million-token slope will not reach massive scale without fundamental changes. Kyle schools the room on hardware-context co-design, MLA, expert sparsity trade-offs in Kimi K2, and Leopold Aschenbrenner's concept of architectural unhobblers.54:39–1:11:42 · The hosts as informed peer 6/10 Enterprise Coding Agents and the Strategic Value of CLIs Swyx challenges the industry trend of wrapping software in CLIs rather than raw REST/MCP APIs for coding agents. Kyle and Nader explain that the sheer volume of pre-training bash data, sandboxed execution safety, and deterministic encapsulation make CLIs superior for LLMs.1:11:42–1:21:51 · The hosts as informed peer 8/10 Sub-Agent Architectures, Compute Costs, and Local GPU Hardware Swyx details sub-agent execution dynamics and compares human-equivalent work hours against real-world 20-45 minute agent runtimes cited by Anthropic and Meter. Kyle and Vibo discuss the trade-offs between local workstation hardware (RTX 6000 Ada/Blackwell Pro) and data center inference economies of scale.0:00–9:32 · Guest teaching 3/10 The Fundamental Security Dilemma of Autonomous AI Agents Swyx and Nader share nostalgic startup stories about Brev's early GTC marketing stunts and its acquisition by NVIDIA. The tone is highly collaborative and warm, with Swyx asking about product design choices and cloud hardware abstractions.9:33–19:12 · Guest teaching 4/10 NVIDIA Developer Experience and the Speed of Light Philosophy Swyx probes into NVIDIA's internal 'Speed of Light' (SOL) operating philosophy, questioning whether anyone besides Jensen Huang can invoke it without derailing stability. Kyle and Nader explain how SOL functions as a first-principles physics baseline rather than mere pressure.19:12–27:17 · Guest teaching 5/10 Recommenders, Graph Neural Networks, and Zero Billion Dollar Markets Kyle outlines his journey from recommendation systems (DLRM, Wide & Deep) and GNNs to modern LLMs, explaining NVIDIA's embrace of 'zero billion dollar markets'. Swyx playfully challenges whether the automotive market qualifies as zero billion, prompting clarification on emerging versus mature bets.27:17–44:35 · Guest teaching 8/10 Architecting NVIDIA Dynamo for Data Center Scale Inference Kyle delivers a technical breakdown of NVIDIA Dynamo, distinguishing scale-up versus scale-out limits, NVLink vs InfiniBand interconnect bandwidths, and pre-fill versus decode disaggregation. Swyx and Vibo actively contribute relevant terminology and architectural context.44:36–54:39 · Guest teaching 8/10 Context Window Limits, Hardware Co-Design, and Architectural Unhobblers Swyx pushes back against current context window scaling trajectories, arguing the million-token slope will not reach massive scale without fundamental changes. Kyle schools the room on hardware-context co-design, MLA, expert sparsity trade-offs in Kimi K2, and Leopold Aschenbrenner's concept of architectural unhobblers.54:39–1:11:42 · Guest teaching 6/10 Enterprise Coding Agents and the Strategic Value of CLIs Swyx challenges the industry trend of wrapping software in CLIs rather than raw REST/MCP APIs for coding agents. Kyle and Nader explain that the sheer volume of pre-training bash data, sandboxed execution safety, and deterministic encapsulation make CLIs superior for LLMs.1:11:42–1:21:51 · Guest teaching 6/10 Sub-Agent Architectures, Compute Costs, and Local GPU Hardware Swyx details sub-agent execution dynamics and compares human-equivalent work hours against real-world 20-45 minute agent runtimes cited by Anthropic and Meter. Kyle and Vibo discuss the trade-offs between local workstation hardware (RTX 6000 Ada/Blackwell Pro) and data center inference economies of scale.0:00–9:32 · Guest disagreement 1/10 The Fundamental Security Dilemma of Autonomous AI Agents Swyx and Nader share nostalgic startup stories about Brev's early GTC marketing stunts and its acquisition by NVIDIA. The tone is highly collaborative and warm, with Swyx asking about product design choices and cloud hardware abstractions.9:33–19:12 · Guest disagreement 2/10 NVIDIA Developer Experience and the Speed of Light Philosophy Swyx probes into NVIDIA's internal 'Speed of Light' (SOL) operating philosophy, questioning whether anyone besides Jensen Huang can invoke it without derailing stability. Kyle and Nader explain how SOL functions as a first-principles physics baseline rather than mere pressure.19:12–27:17 · Guest disagreement 2/10 Recommenders, Graph Neural Networks, and Zero Billion Dollar Markets Kyle outlines his journey from recommendation systems (DLRM, Wide & Deep) and GNNs to modern LLMs, explaining NVIDIA's embrace of 'zero billion dollar markets'. Swyx playfully challenges whether the automotive market qualifies as zero billion, prompting clarification on emerging versus mature bets.27:17–44:35 · Guest disagreement 1/10 Architecting NVIDIA Dynamo for Data Center Scale Inference Kyle delivers a technical breakdown of NVIDIA Dynamo, distinguishing scale-up versus scale-out limits, NVLink vs InfiniBand interconnect bandwidths, and pre-fill versus decode disaggregation. Swyx and Vibo actively contribute relevant terminology and architectural context.44:36–54:39 · Guest disagreement 3/10 Context Window Limits, Hardware Co-Design, and Architectural Unhobblers Swyx pushes back against current context window scaling trajectories, arguing the million-token slope will not reach massive scale without fundamental changes. Kyle schools the room on hardware-context co-design, MLA, expert sparsity trade-offs in Kimi K2, and Leopold Aschenbrenner's concept of architectural unhobblers.54:39–1:11:42 · Guest disagreement 2/10 Enterprise Coding Agents and the Strategic Value of CLIs Swyx challenges the industry trend of wrapping software in CLIs rather than raw REST/MCP APIs for coding agents. Kyle and Nader explain that the sheer volume of pre-training bash data, sandboxed execution safety, and deterministic encapsulation make CLIs superior for LLMs.1:11:42–1:21:51 · Guest disagreement 2/10 Sub-Agent Architectures, Compute Costs, and Local GPU Hardware Swyx details sub-agent execution dynamics and compares human-equivalent work hours against real-world 20-45 minute agent runtimes cited by Anthropic and Meter. Kyle and Vibo discuss the trade-offs between local workstation hardware (RTX 6000 Ada/Blackwell Pro) and data center inference economies of scale.0:00–9:32 · The hosts pushing back 2/10 The Fundamental Security Dilemma of Autonomous AI Agents Swyx and Nader share nostalgic startup stories about Brev's early GTC marketing stunts and its acquisition by NVIDIA. The tone is highly collaborative and warm, with Swyx asking about product design choices and cloud hardware abstractions.9:33–19:12 · The hosts pushing back 3/10 NVIDIA Developer Experience and the Speed of Light Philosophy Swyx probes into NVIDIA's internal 'Speed of Light' (SOL) operating philosophy, questioning whether anyone besides Jensen Huang can invoke it without derailing stability. Kyle and Nader explain how SOL functions as a first-principles physics baseline rather than mere pressure.19:12–27:17 · The hosts pushing back 3/10 Recommenders, Graph Neural Networks, and Zero Billion Dollar Markets Kyle outlines his journey from recommendation systems (DLRM, Wide & Deep) and GNNs to modern LLMs, explaining NVIDIA's embrace of 'zero billion dollar markets'. Swyx playfully challenges whether the automotive market qualifies as zero billion, prompting clarification on emerging versus mature bets.27:17–44:35 · The hosts pushing back 2/10 Architecting NVIDIA Dynamo for Data Center Scale Inference Kyle delivers a technical breakdown of NVIDIA Dynamo, distinguishing scale-up versus scale-out limits, NVLink vs InfiniBand interconnect bandwidths, and pre-fill versus decode disaggregation. Swyx and Vibo actively contribute relevant terminology and architectural context.44:36–54:39 · The hosts pushing back 5/10 Context Window Limits, Hardware Co-Design, and Architectural Unhobblers Swyx pushes back against current context window scaling trajectories, arguing the million-token slope will not reach massive scale without fundamental changes. Kyle schools the room on hardware-context co-design, MLA, expert sparsity trade-offs in Kimi K2, and Leopold Aschenbrenner's concept of architectural unhobblers.54:39–1:11:42 · The hosts pushing back 4/10 Enterprise Coding Agents and the Strategic Value of CLIs Swyx challenges the industry trend of wrapping software in CLIs rather than raw REST/MCP APIs for coding agents. Kyle and Nader explain that the sheer volume of pre-training bash data, sandboxed execution safety, and deterministic encapsulation make CLIs superior for LLMs.1:11:42–1:21:51 · The hosts pushing back 4/10 Sub-Agent Architectures, Compute Costs, and Local GPU Hardware Swyx details sub-agent execution dynamics and compares human-equivalent work hours against real-world 20-45 minute agent runtimes cited by Anthropic and Meter. Kyle and Vibo discuss the trade-offs between local workstation hardware (RTX 6000 Ada/Blackwell Pro) and data center inference economies of scale.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 39.2% · guest 60.8%0:00 · the hosts 39.2% · guest 60.8%3:00 · the hosts 12.3% · guest 87.7%3:00 · the hosts 12.3% · guest 87.7%6:00 · the hosts 26% · guest 74%6:00 · the hosts 26% · guest 74%9:00 · the hosts 18.6% · guest 81.4%9:00 · the hosts 18.6% · guest 81.4%12:00 · the hosts 8.8% · guest 91.2%12:00 · the hosts 8.8% · guest 91.2%15:00 · the hosts 14.7% · guest 85.3%15:00 · the hosts 14.7% · guest 85.3%18:00 · the hosts 30.1% · guest 69.9%18:00 · the hosts 30.1% · guest 69.9%21:00 · the hosts 7.7% · guest 92.3%21:00 · the hosts 7.7% · guest 92.3%24:00 · the hosts 7.9% · guest 92.1%24:00 · the hosts 7.9% · guest 92.1%27:00 · the hosts 5.1% · guest 94.9%27:00 · the hosts 5.1% · guest 94.9%30:00 · the hosts 5.4% · guest 94.6%30:00 · the hosts 5.4% · guest 94.6%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 23.7% · guest 76.3%36:00 · the hosts 23.7% · guest 76.3%39:00 · the hosts 3.1% · guest 96.9%39:00 · the hosts 3.1% · guest 96.9%42:00 · the hosts 23.2% · guest 76.8%42:00 · the hosts 23.2% · guest 76.8%45:00 · the hosts 3.1% · guest 96.9%45:00 · the hosts 3.1% · guest 96.9%48:00 · the hosts 36.6% · guest 63.4%48:00 · the hosts 36.6% · guest 63.4%51:00 · the hosts 3.1% · guest 96.9%51:00 · the hosts 3.1% · guest 96.9%54:00 · the hosts 11.2% · guest 88.8%54:00 · the hosts 11.2% · guest 88.8%57:00 · the hosts 16.2% · guest 83.8%57:00 · the hosts 16.2% · guest 83.8%1:00:00 · the hosts 8.7% · guest 91.3%1:00:00 · the hosts 8.7% · guest 91.3%1:03:00 · the hosts 36.2% · guest 63.8%1:03:00 · the hosts 36.2% · guest 63.8%1:06:00 · the hosts 7.8% · guest 92.2%1:06:00 · the hosts 7.8% · guest 92.2%1:09:00 · the hosts 12.7% · guest 87.3%1:09:00 · the hosts 12.7% · guest 87.3%1:12:00 · the hosts 2.4% · guest 97.6%1:12:00 · the hosts 2.4% · guest 97.6%1:15:00 · the hosts 35.4% · guest 64.6%1:15:00 · the hosts 35.4% · guest 64.6%1:18:00 · the hosts 33.1% · guest 66.9%1:18:00 · the hosts 33.1% · guest 66.9%1:21:00 · the hosts 22.2% · guest 77.8%1:21:00 · the hosts 22.2% · guest 77.8%1:24:00 · the hosts 41.4% · guest 58.6%1:24:00 · the hosts 41.4% · guest 58.6%
Sharpest disagreement ▶ 48:15 Kyle and Swyx Spar Over Training Harnesses into Models

Kyle insists that optimal agent performance requires baking specific execution harnesses into pre-training, directly resisting Swyx's counterargument that models should remain general-purpose.

Hardest push from the hosts ▶ 49:45 Swyx Flatly Rejects Linear Context Window Scaling

Swyx aggressively pushes back on conventional context-scaling assumptions, arguing that reaching 100 trillion tokens is mathematically impossible on current slopes without revolutionary unhobblers.

Biggest teaching moment ▶ 29:45 Kyle Explains Hardware Physics of Disaggregated Inference

Kyle methodically educates the hosts on the microarchitectural split between compute-bound quadratic prefill and memory-bound linear decode steps in cluster-scale serving.

The host holds their own ▶ 1:19:30 Swyx Clarifies Human-Equivalent vs Clock-Time Agent Benchmarks

Swyx draws upon his direct interviews with the Meter and Anthropic teams to correct inflated public impressions of agent longevity, citing measured 20-to-45-minute production traffic figures.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
The Fundamental Security Dilemma of Autonomous AI Agents 4312 Swyx and Nader share nostalgic startup stories about Brev's early GTC marketing stunts and its acquisition by NVIDIA. The tone is highly collaborative and warm, with Swyx asking about product design choices and cloud hardware abstractions.
NVIDIA Developer Experience and the Speed of Light Philosophy 6423 Swyx probes into NVIDIA's internal 'Speed of Light' (SOL) operating philosophy, questioning whether anyone besides Jensen Huang can invoke it without derailing stability. Kyle and Nader explain how SOL functions as a first-principles physics baseline rather than mere pressure.
Recommenders, Graph Neural Networks, and Zero Billion Dollar Markets 6523 Kyle outlines his journey from recommendation systems (DLRM, Wide & Deep) and GNNs to modern LLMs, explaining NVIDIA's embrace of 'zero billion dollar markets'. Swyx playfully challenges whether the automotive market qualifies as zero billion, prompting clarification on emerging versus mature bets.
Architecting NVIDIA Dynamo for Data Center Scale Inference 7812 Kyle delivers a technical breakdown of NVIDIA Dynamo, distinguishing scale-up versus scale-out limits, NVLink vs InfiniBand interconnect bandwidths, and pre-fill versus decode disaggregation. Swyx and Vibo actively contribute relevant terminology and architectural context.
Context Window Limits, Hardware Co-Design, and Architectural Unhobblers 7835 Swyx pushes back against current context window scaling trajectories, arguing the million-token slope will not reach massive scale without fundamental changes. Kyle schools the room on hardware-context co-design, MLA, expert sparsity trade-offs in Kimi K2, and Leopold Aschenbrenner's concept of architectural unhobblers.
Enterprise Coding Agents and the Strategic Value of CLIs 6624 Swyx challenges the industry trend of wrapping software in CLIs rather than raw REST/MCP APIs for coding agents. Kyle and Nader explain that the sheer volume of pre-training bash data, sandboxed execution safety, and deterministic encapsulation make CLIs superior for LLMs.
Sub-Agent Architectures, Compute Costs, and Local GPU Hardware 8624 Swyx details sub-agent execution dynamics and compares human-equivalent work hours against real-world 20-45 minute agent runtimes cited by Anthropic and Meter. Kyle and Vibo discuss the trade-offs between local workstation hardware (RTX 6000 Ada/Blackwell Pro) and data center inference economies of scale.

Statements from this episode (21)

Disclosure
NVIDIA mandates running OpenClaw in isolated Brev cloud VMs
“Internally people want to run this and we know we have to be really careful from the security implications. Do we let this run on the corporate network securities guidance was, Hey, run this on breath. It's in, you know, it's a VM. It's sitting in the cloud. …”
Nader Khalil Mar 8, 2026 ▶ 9:08
Disclosure
NVIDIA Brev launches DGX Spark hardware registration in early access
“And one thing we just launched at CES, it's in, it's still in like early access. We're ironing out some kinks, but it should be ready by GTC. You can register your Spark on Brev.”
Nader Khalil Mar 8, 2026 ▶ 13:29
Insight
NVIDIA uses 'Speed of Light' framework to define minimum viable goals
“SOL is a term that NVIDIA is used to sort of like instigate a compelling event. You say, this is done. How do we get there? What is the minimum, as much as necessary, as little as possible thing that it takes for us to get exactly here? And it helps you just b…”
Kyle Kranen Mar 8, 2026 ▶ 15:28
Assertion Supported
NVIDIA chip design begins three to five years before market release
“The design process starts like- Exactly. ...three to five years before the chip gets to the market.”
Kyle Kranen Mar 8, 2026 ▶ 22:39
Assertion Supported
Jensen Huang prioritizes strategic investments in 'zero billion dollar markets'
“Jensen, He says, we're completely happy investing in zero billion dollar markets. We don't care if this creates revenue. It's important for us to know about this market. We think it will be important in the future. It can be zero billion dollars for a while.”
Kyle Kranen Mar 8, 2026 ▶ 25:14
Assertion Supported
Scaling beyond eight NVIDIA H100 GPUs requires inter-GPU communication over InfiniBand
“The maximum NVLink domain, domain for H 100 for most DGX H 100 Is eight GPUs, right? So if you scaled up past that, You're gonna have to figure out ways to handle the fact that now for the GPUs to communicate, you have to do it over InfiniBand, which is still …”
Kyle Kranen Mar 8, 2026 ▶ 30:26
Assertion Supported
ServiceNow trained an in-house AI model using NVIDIA's Nemotron dataset
“There are companies like, ah, ServiceNow took the dataset and they trained their own model.”
Nader Khalil Mar 8, 2026 ▶ 38:33
Insight
LLM prefill remains compute-bound while decoding phases are strictly memory-bound
“So prefill typically, and this changes as model architecture changes, prefill is right now compute bound. Most of the time. If the sequence is sufficiently long, it's compute bound on the decode side because you're doing a full pass over all the weights and th…”
Kyle Kranen Mar 8, 2026 ▶ 41:14
Assertion Supported
NVIDIA announces Rubin CPX as a dedicated prefill-specific hardware accelerator
“And like with our future generations, generations of hardware, we actually announced like with Rubin, this new accelerator that is pre-fill specific. It's called Rubin CPX.”
Kyle Kranen Mar 8, 2026 ▶ 42:04
Assertion Supported
NVIDIA Dynamo dynamically sizes and schedules Kubernetes prefill and decode workers
“Dynamo has a set of components that A, tell you how to scale. It tells you how many pre-fill workers and decoded workers it thinks you should have. And also provides a scheduling API for Kubernetes that allows you to actually represent and affect this scheduli…”
Kyle Kranen Mar 8, 2026 ▶ 43:57
Assertion Supported
DeepSeek 128k context fits in 8GB KV cache versus Llama's 80GB
“For context like the total I think the total context length of DeepSeq is a 128,000 tokens, or it might be 256,000 with rope extension. That entire context, I think it's a 128,000, fits into eight gigabytes. And previously context, like I think the Lama four o…”
Kyle Kranen Mar 8, 2026 ▶ 52:08
Prediction Not checkable as stated
Algorithmic unhobblers could soon expand context windows to 100 million tokens
“I wouldn't be surprised if we do see the ability to, like, break through to, like, ten million, twenty million, a hundred million context through the, an unhobbler showing up.”
Kyle Kranen Mar 8, 2026 ▶ 52:37
Insight
Local document pre-fill with global sequence decode solves transformer quadratic scaling
“If pre-fill becomes local and decode is, is still global, you solve that pre-fill quadratic scaling problem because you have a bunch of like small chunks that you pre-fill independently.”
Kyle Kranen Mar 8, 2026 ▶ 53:49
Insight
Secure AI agents must restrict one of file, internet, or code access
“Agents can do three things. They can access your files, they can access the internet, and then now they can write custom code and execute it. And you really only let an agent do two of those three things. If you can access your files and you can write custom c…”
Nader Khalil Mar 8, 2026 ▶ 59:54
Assertion Not checkable as stated
NVIDIA's build.nvidia.com was internally the company's largest inference deployment
“At one point, there's a website called build.nvd.com, and also for us, inference.nvd.com, that is, allows people to try models. It gives an API service, you can call the model with like a REST API, and, you know, you get a response. I ran the model site for th…”
Kyle Kranen Mar 8, 2026 ▶ 1:01:02
Insight
Terminal access makes coding agents significantly more effective than general agents
“I feel like coding agents have been so much more effective than general purpose agents, and I think a large part of that is it just has access to the terminal, like you said, and that means it has access to everything that you've installed into your terminal.”
Nader Khalil Mar 8, 2026 ▶ 1:08:26
Disclosure
NVIDIA's Brev team plans to open-source business application command-line interfaces
“We're gonna open source all of this and like, yeah, all the, I mean, they're just, they're, yeah, CLIs for the business applications.”
Nader Khalil Mar 8, 2026 ▶ 1:08:58
Insight
LLMs favor CLI tools over APIs due to massive pre-training data volumes
“I think that in pre-training, there's just an enormous amount of command line data. Like even let's ignore, let's like, let's ignore RL. Like you're doing no harness post training. Just the amount of like CLI versus API documentation for just like navigating t…”
Kyle Kranen Mar 8, 2026 ▶ 1:10:19
Assertion Supported
NVIDIA releases DGX Spark model router to dynamically route local inference
“We actually, for CES, we just released the model router. So for DGX Spark, where you can have a local model that's running on the Spark and then also a foundation model. And then the model order decides when to send queries to which one.”
Nader Khalil Mar 8, 2026 ▶ 1:17:16
Prediction Open · timeframe Dec 2026
AI agents will achieve 24-hour self-consistent autonomous runtimes by late 2026
“We will see before the end of the year an agent that is capable of running for longer than 24 hours with like self consistency the entire time.”
Kyle Kranen Mar 8, 2026 ▶ 1:19:00
Assertion Supported
Anthropic production traffic reveals Claude Code's autonomous execution spans 20-45 minutes
“I think actually Enflopic released a more recent chart that showed cloud code autonomy from their production traffic numbers, and that was 20 to 45 minutes.”
Shawn Wang Mar 8, 2026 ▶ 1:20:26
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.