“For context like the total I think the total context length of DeepSeq is a 128,000 tokens, or it might be 256,000 with rope extension. That entire context, I think it's a 128,000, fits into eight gigabytes. And previously context, like I think the Lama four or five B context of a similar size was like 40 or 80 gigabytes in the same precision.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Kyle Kranen
PredictionNot checkable as stated
Algorithmic unhobblers could soon expand context windows to 100 million tokens
“I wouldn't be surprised if we do see the ability to, like, break through to, like, ten million, twenty million, a hundred million context through the, an unhobbler showing up.”
Kyle KranenMar 8, 2026▶ 52:37Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
Insight
LLMs favor CLI tools over APIs due to massive pre-training data volumes
“I think that in pre-training, there's just an enormous amount of command line data. Like even let's ignore, let's like, let's ignore RL. Like you're doing no harness post training. Just the amount of like CLI versus API documentation for just like navigating t…”
Kyle KranenMar 8, 2026▶ 1:10:19Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
PredictionOpen · timeframe Dec 2026
AI agents will achieve 24-hour self-consistent autonomous runtimes by late 2026
“We will see before the end of the year an agent that is capable of running for longer than 24 hours with like self consistency the entire time.”
Kyle KranenMar 8, 2026▶ 1:19:00Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
Insight
NVIDIA uses 'Speed of Light' framework to define minimum viable goals
“SOL is a term that NVIDIA is used to sort of like instigate a compelling event. You say, this is done. How do we get there? What is the minimum, as much as necessary, as little as possible thing that it takes for us to get exactly here? And it helps you just b…”
Kyle KranenMar 8, 2026▶ 15:28Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
AssertionSupported
Jensen Huang prioritizes strategic investments in 'zero billion dollar markets'
“Jensen, He says, we're completely happy investing in zero billion dollar markets. We don't care if this creates revenue. It's important for us to know about this market. We think it will be important in the future. It can be zero billion dollars for a while.”
Kyle KranenMar 8, 2026▶ 25:14Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
AssertionSupported
NVIDIA announces Rubin CPX as a dedicated prefill-specific hardware accelerator
“And like with our future generations, generations of hardware, we actually announced like with Rubin, this new accelerator that is pre-fill specific. It's called Rubin CPX.”
Kyle KranenMar 8, 2026▶ 42:04Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 200 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.