VLLM

product on 6 shows · 14 statements across 9 episodes · said 115 times in 32 episodes since 2023

Latent Space 75 the a16z Podcast 25 the MAD Podcast 11 20VC 2 the Y Combinator Startup Podcast 1 No Priors 1

Mentions by year, every show

tap a year for its mentions
0025850152023202420252026episodesmentions
08152023202420252026episodes it came up in
002.57.55152023202420252026episodesmentions per episode

Latent Space 75the a16z Podcast 25the MAD Podcast 1120VC 2No Priors 1the Y Combinator Startup Podcast 1

2026 45 mentions in 9 episodes 5 per episode
2025 47 mentions in 11 episodes 4 per episode
2024 22 mentions in 11 episodes 2 per episode
2023 1 mention in 1 episode

every mention on every show, scene by scene, with the transcript →

14 statements about VLLM, every show

a16z Assertion Not checkable as stated
Burger: vLLM Runs on 500,000 GPUs at Any Moment
“Today we're here with Simon Moe, co-founder of Infraact, and a lead maintainer of VLLM, the open source inference engine, now running on half a million GPUs at any moment.”
Elena Burger Aug 5, 2026 ▶ 1:01 How Open Source Became AI's Backbone | Inferact with a16z
a16z Assertion Partly supported
Simon Mo: vLLM supports over 1,000 model architectures
“For VRM, we support more than a thousand model architecture up to today, and a lot of those are proprietary, but also a lot of those are open-weight, right?”
Simon Moe Aug 5, 2026 ▶ 8:54 How Open Source Became AI's Backbone | Inferact with a16z
a16z Assertion Supported
Mo: Major chipmakers use vLLM as an internal benchmark
“And additionally, VLM also work closely with all the hardware vendors. So that means across like NVIDIA, AMD, Google, and Amazon, Intel, and a lot more, their newest chip will make sure VLM can run on them. And then a lot of cases they use VLM as a benchmark t…”
Simon Moe Aug 5, 2026 ▶ 9:29 How Open Source Became AI's Backbone | Inferact with a16z
MAD Assertion Supported
Lambert: Kernel differences between vLLM and Hugging Face cause RL numerical instability
“VLLM and HuggingFace use different kernels to do the actual internal computation of the model. So these kernels are the things that make things like vLLM really fast. But these things, this then results in subtle numerical differences between the completions t…”
Nathan Lambert Nov 20, 2025 ▶ 1:15:29 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
LATENT SPACE Assertion Not checkable as stated
Ermon: Inception Labs built proprietary engine for production inference traffic
“So just like you would normally serve an LLM using a VLLM or SGLang or a Tensor or TLLM, we have built our own inference engine. And so we are supporting production traffic already with our own inference engine. We support continuous batch and quantization.”
Stefano Ermon Aug 4, 2025 ▶ 21:33 ⚡️Mercury: Ultra-Fast Diffusion LLMs — Estefano Ermon, CEO Inception Labs
Mohan: Codeium would not have survived if it relied on vLLM
“If we use VLLM, we would not be talking with you right now.”
Varun Mohan Jul 28, 2025 ▶ 2:34:14 🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
LATENT SPACE Prediction Not checkable as stated
Morris: vLLM and SGLang are here to stay and will grow more complex
“I also think, ah, VLLM and SGLang seem, like, really good and important and here to stay. Like, they'll probably just get larger and more complex to accommodate future systems”
Jack Morris Jul 2, 2025 ▶ 14:10 Information Theory for Language Models: Jack Morris
Lattner: Modular MAX is 'more open source' than vLLM
“This thing's more open source than VLM because VLM depends on all these crazy binary CUDA kernels and stuff like this that are just opaque blobs from NVIDIA, right?”
Chris Lattner Jun 13, 2025 ▶ 12:25 The Shape of Compute (Chris Lattner of Modular)
Lattner calls vLLM a 'hot mess' due to too many stakeholders
“VLM seems much more like a massive community with a lot of stakeholders, a lot of stuff going on, and it's kind of a hot mess.”
Chris Lattner Jun 13, 2025 ▶ 23:52 The Shape of Compute (Chris Lattner of Modular)
Zhang: SGLang outperforms vLLM and has better usability than TensorRT-LLM
“I think for the common use case, maybe not, not the DeepSeq VIII, for the common use case, I think SGLAN's performance is better than FLM, and its usability is better than TensorFlow TLM.”
Yining Zhang Jan 19, 2025 ▶ 26:57 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
LATENT SPACE Assertion Supported
Zhang: vLLM does not support DeepSeek MLA while SGLang does
“Something like DeepSeq V-II, they proposed attention parent named MLA, multi-latent attention, and I think SGLAN is the only framework to support that. Maybe LightLM and TRTM also support, but VLM doesn't support.”
Yining Zhang Jan 19, 2025 ▶ 27:23 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
LATENT SPACE Assertion Supported
Zhang: SGLang achieves higher cache hit rates using block size of one
“Redix cache, I think it's the technology of the prefix caching, and it is a special case for something like block size is one, you know, for VLM or for other frameworks, they use something block size, 32, and SGLAN use the block size one. I think if you use th…”
Yining Zhang Jan 19, 2025 ▶ 36:03 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Varun Mohan: Codeium would have failed if it used vLLM
“If we use VLLM, we would not be talking with you right now.”
Varun Mohan Dec 13, 2024 ▶ 57:27 Windsurf: The Enterprise AI IDE
LATENT SPACE Assertion Not checkable as stated
Jamil: Chunked prefill is experimental in vLLM, likely used by majors
“This is called the chunked pre-fill, and it's an experimental feature that has been recently introduced in VLLN, but it's probably used in more sophisticated inference engines at major companies.”
Umar Jamil Sep 19, 2024 ▶ 6:07 [Paper Club] Writing in the Margins: Chunked Prefill KV Caching for Long Context Retrieval

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.