Assertion Supported AI assessment confidence: 95% certainty 4/5 debate potential 1/5

NVIDIA Dynamo dynamically sizes and schedules Kubernetes prefill and decode workers

Kyle Kranen · Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup · Mar 8, 2026 · at 43:57

Kyle Kranen explains how NVIDIA's open-source inference engine Dynamo integrates with Kubernetes to scale out disaggregated inference pipelines.

0:00 / 0:18exact quote · 18.6s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“Dynamo has a set of components that A, tell you how to scale. It tells you how many pre-fill workers and decoded workers it thinks you should have. And also provides a scheduling API for Kubernetes that allows you to actually represent and affect this scheduling on, on, on your actual hardware, on your computer infrastructure.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Kyle Kranen

Prediction Not checkable as stated
Algorithmic unhobblers could soon expand context windows to 100 million tokens
“I wouldn't be surprised if we do see the ability to, like, break through to, like, ten million, twenty million, a hundred million context through the, an unhobbler showing up.”
Kyle Kranen Mar 8, 2026 ▶ 52:37 Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
Insight
LLMs favor CLI tools over APIs due to massive pre-training data volumes
“I think that in pre-training, there's just an enormous amount of command line data. Like even let's ignore, let's like, let's ignore RL. Like you're doing no harness post training. Just the amount of like CLI versus API documentation for just like navigating t…”
Kyle Kranen Mar 8, 2026 ▶ 1:10:19 Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
Prediction Open · timeframe Dec 2026
AI agents will achieve 24-hour self-consistent autonomous runtimes by late 2026
“We will see before the end of the year an agent that is capable of running for longer than 24 hours with like self consistency the entire time.”
Kyle Kranen Mar 8, 2026 ▶ 1:19:00 Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
Insight
NVIDIA uses 'Speed of Light' framework to define minimum viable goals
“SOL is a term that NVIDIA is used to sort of like instigate a compelling event. You say, this is done. How do we get there? What is the minimum, as much as necessary, as little as possible thing that it takes for us to get exactly here? And it helps you just b…”
Kyle Kranen Mar 8, 2026 ▶ 15:28 Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
Assertion Supported
Jensen Huang prioritizes strategic investments in 'zero billion dollar markets'
“Jensen, He says, we're completely happy investing in zero billion dollar markets. We don't care if this creates revenue. It's important for us to know about this market. We think it will be important in the future. It can be zero billion dollars for a while.”
Kyle Kranen Mar 8, 2026 ▶ 25:14 Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
Assertion Supported
NVIDIA announces Rubin CPX as a dedicated prefill-specific hardware accelerator
“And like with our future generations, generations of hardware, we actually announced like with Rubin, this new accelerator that is pre-fill specific. It's called Rubin CPX.”
Kyle Kranen Mar 8, 2026 ▶ 42:04 Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.