vLLM

10 statements across 7 episodes · 3 bullish · 5 bearish · 6 people on the record · first statement Sep 19, 2024 by Umar Jamil · said 75 times in 23 episodes since 2024 · across every show →

Mentions by year

brought up most by Chris Lattner (11), Shawn Wang (10), Umar Jamil (6), Yining Zhang (5), William Beauchamp (3), Philip Kiely (3), Kyle Kranen (3), Omar Sanseviero (2)

tap a year for its mentions
002555010202420252026episodesmentions
0510202420252026episodes it came up in
0035610202420252026episodesmentions per episode
2026 12 mentions in 5 episodes 2 per episode
2025 44 mentions in 8 episodes 6 per episode
2024 19 mentions in 10 episodes 2 per episode

every mention, scene by scene, with the transcript →

Everything said about vLLM, oldest first

Sep 19, 2024 neutral
Assertion Not checkable as stated
Jamil: Chunked prefill is experimental in vLLM, likely used by majors
“This is called the chunked pre-fill, and it's an experimental feature that has been recently introduced in VLLN, but it's probably used in more sophisticated inference engines at major companies.”
Umar Jamil Sep 19, 2024 ▶ 6:07 [Paper Club] Writing in the Margins: Chunked Prefill KV Caching for Long Context Retrieval
Dec 13, 2024 bearish
What-if
Varun Mohan: Codeium would have failed if it used vLLM
“If we use VLLM, we would not be talking with you right now.”
Varun Mohan Dec 13, 2024 ▶ 57:27 Windsurf: The Enterprise AI IDE
Jan 19, 2025 positive
Opinion
Zhang: SGLang outperforms vLLM and has better usability than TensorRT-LLM
“I think for the common use case, maybe not, not the DeepSeq VIII, for the common use case, I think SGLAN's performance is better than FLM, and its usability is better than TensorFlow TLM.”
Yining Zhang Jan 19, 2025 ▶ 26:57 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Jan 19, 2025 negative
Assertion Supported
Zhang: vLLM does not support DeepSeek MLA while SGLang does
“Something like DeepSeq V-II, they proposed attention parent named MLA, multi-latent attention, and I think SGLAN is the only framework to support that. Maybe LightLM and TRTM also support, but VLM doesn't support.”
Yining Zhang Jan 19, 2025 ▶ 27:23 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Jan 19, 2025 positive
Assertion Supported
Zhang: SGLang achieves higher cache hit rates using block size of one
“Redix cache, I think it's the technology of the prefix caching, and it is a special case for something like block size is one, you know, for VLM or for other frameworks, they use something block size, 32, and SGLAN use the block size one. I think if you use th…”
Yining Zhang Jan 19, 2025 ▶ 36:03 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Jun 13, 2025 negative
Opinion
Lattner calls vLLM a 'hot mess' due to too many stakeholders
“VLM seems much more like a massive community with a lot of stakeholders, a lot of stuff going on, and it's kind of a hot mess.”
Chris Lattner Jun 13, 2025 ▶ 23:52 The Shape of Compute (Chris Lattner of Modular)
Jun 13, 2025 negative
Opinion
Lattner: Modular MAX is 'more open source' than vLLM
“This thing's more open source than VLM because VLM depends on all these crazy binary CUDA kernels and stuff like this that are just opaque blobs from NVIDIA, right?”
Chris Lattner Jun 13, 2025 ▶ 12:25 The Shape of Compute (Chris Lattner of Modular)
Jul 2, 2025 bullish
Prediction Not checkable as stated
Morris: vLLM and SGLang are here to stay and will grow more complex
“I also think, ah, VLLM and SGLang seem, like, really good and important and here to stay. Like, they'll probably just get larger and more complex to accommodate future systems”
Jack Morris Jul 2, 2025 ▶ 14:10 Information Theory for Language Models: Jack Morris
Jul 28, 2025 negative
What-if
Mohan: Codeium would not have survived if it relied on vLLM
“If we use VLLM, we would not be talking with you right now.”
Varun Mohan Jul 28, 2025 ▶ 2:34:14 🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
Aug 4, 2025 neutral
Assertion Not checkable as stated
Ermon: Inception Labs built proprietary engine for production inference traffic
“So just like you would normally serve an LLM using a VLLM or SGLang or a Tensor or TLLM, we have built our own inference engine. And so we are supporting production traffic already with our own inference engine. We support continuous batch and quantization.”
Stefano Ermon Aug 4, 2025 ▶ 21:33 ⚡️Mercury: Ultra-Fast Diffusion LLMs — Estefano Ermon, CEO Inception Labs
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.