vLLM, every mention

41 scenes · ← back to vLLM

tap a year for its mentions
002555010202420252026episodesmentions
0510202420252026episodes it came up in
0035610202420252026episodesmentions per episode

every year anyone Chris Lattner 11Shawn Wang 10Umar Jamil 6Yining Zhang 5William Beauchamp 3Philip Kiely 3Kyle Kranen 3Omar Sanseviero 2Jim Zemlin 2Jeremy Howard 2

Verbatim, from the transcripts: the passages where vLLM comes up

loading…

Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten Aug 3, 2026 · 5 mentions

  • ▶ 13:51 Philip Kiely Getting to the point of I can make a token out of this model is not that hard because generally the, um, open source inference engines, your VLMs, SGLangs of the world oftentimes even receive weights ahead of time maintainers do, or the…
  • ▶ 24:12 Ali Taha But if you were to switch to VLM, that isn't the case.
  • ▶ 42:12 Alessio Fanelli Put it behind VLO.
  • ▶ 1:00:40 Philip Kiely So if you look at, like, the original VLM and SGLang, or VLM especially, like, that was written targeting Ampure and then had to be updated for Hopper, updated for Blackwell. 2 times in the scene

The AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO Jul 13, 2026 · 1 mention

  • ▶ 44:40 Dan Biderman Leads at Databricks, and one of the core contributors at VLLM, and we're all kind of like, ah, systems inclined, but we think infrastructures, engineers, ah, people who know how to work, ah,

The Future of AI Infra: from Kubernetes to Agent Sandboxes — Akshat Bubna, Modal CTO Jul 8, 2026 · 1 mention

  • ▶ 21:30 unnamed speaker How does it compare to, say, I take the same model, GLM, 5.2 FBA, take off the shelf inference engine, VLM, SGLang, um, you know, get compute of similar capacity, similar cost.

⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind May 24, 2026 · 2 mentions

  • ▶ 3:42 Omar Sanseviero So for example, we work with Lama CPP, Olama, MLX, Hogan Faces, BLM, NVIDIA, AMD. 2 times in the scene

Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup Mar 8, 2026 · 3 mentions

  • ▶ 28:05 Kyle Kranen Dynoa sort of came about at NVIDIA because myself and a couple others were sort of talking about these concepts that like, you know, you have inference engines like VLM, SGLang, TensorRTLM, um, 3 times in the scene

One Year of MCP — with David Soria Parria and AAIF leads from OpenAI, Goose, Linux Foundation Dec 28, 2025 · 2 mentions

A Technical History of Generative Media Sep 8, 2025 · 1 mention

  • ▶ 11:36 unnamed speaker And there's no community effort, like, a VLM?

⚡️Mercury: Ultra-Fast Diffusion LLMs — Estefano Ermon, CEO Inception Labs Aug 4, 2025 · 1 mention

  • ▶ 21:33 Stefano Ermon So just like you would normally serve an LLM using a VLLM or SGLang or a Tensor or TLLM, we have built our own inference engine.

🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R) Jul 28, 2025 · 2 mentions

Information Theory for Language Models: Jack Morris Jul 2, 2025 · 1 mention

  • ▶ 14:10 Jack Morris I also think, ah, VLLM and SGLang seem, like, really good and important and here to stay.

The Shape of Compute (Chris Lattner of Modular) Jun 13, 2025 · 17 mentions

  • ▶ 11:34 Chris Lattner So it's not as good as something like VLLM because it's missing some features, and it only supports NVIDIA and AMD hardware, for example. 4 times in the scene
  • ▶ 23:09 Shawn Wang I'm curious if you have any views or insider takes on what's happening with VLM versus SGLang and everything coming out of Berkeley. 6 times in the scene
  • ▶ 35:59 Chris Lattner And so this is why we have things like VLM and SGLang because they're the black box that you can just hopefully build on top of and not have to know how any of that scary stuff works is because we haven't taught the industry how to do this… 2 times in the scene
  • ▶ 41:06 Chris Lattner Well, again, you get back into hacking the internals of VLM and PyTorch isn't really designed for KV cache optimizations and all the modern transformer features and things like this. 2 times in the scene
  • ▶ 56:02 Chris Lattner and it drew attention to that layer of the stack because it wasn't just like throwing layers on top of PyTorch, you know, or VLM, or it was like doing that fun amount of work. 3 times in the scene

Outlasting Noam Shazeer, Crowdsourcing Chai AI w/ 1.4m DAU — with William Beauchamp, Chai Research Jan 26, 2025 · 3 mentions

DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing) Jan 19, 2025 · 17 mentions

  • ▶ 20:36 unnamed speaker We have customers on, on based end that are using TensorFlow and we have ones that are using VL and we have a growing number that are using SGLang too.
  • ▶ 26:35 unnamed speaker So you have SGLang, TRTLLM, VLLM. 4 times in the scene
  • ▶ 32:01 unnamed speaker And that's something that I'm seeing in the market that, like, people who are somewhat new to it, they're like, well, VLM equals equals production. 2 times in the scene
  • ▶ 33:52 Yining Zhang Equivalent with the FLM or with the Tencent RTLM. 4 times in the scene
  • ▶ 36:03 Yining Zhang Redix cache, I, I think it's the technology of the prefix caching, and it is a special case for something like block size is one, you know, for VLM or for other frameworks, they use something block size, 32, and SGLAN use the block size…
  • ▶ 43:17 unnamed speaker It seems like VLM obviously has the, the, it's one year older. 3 times in the scene
  • ▶ 50:55 unnamed speaker Honestly, I would go back to what I emphasized earlier, which was that I wish more people asked about what it takes to run mission critical inference workloads, because I see this in the market sometimes that they're like, well, I can just… 2 times in the scene

Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024] Dec 24, 2024 · 1 mention

Windsurf: The Enterprise AI IDE Dec 13, 2024 · 2 mentions

  • ▶ 57:22 unnamed speaker You can't just use VLLM and TensorFlow. 2 times in the scene

Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI Nov 25, 2024 · 2 mentions

  • ▶ 13:06 Shawn Wang Basically, it's just like the meta version of whatever HuggingFace offers, you know, or TensorRT, or BLM, or whatever the open source opportunity is.
  • ▶ 23:37 Shawn Wang I don't remember the, the speed numbers, but apparently much better than BLM, especially on a concurrency basis.

In the Arena: How LMSys changed LLM Benchmarking Forever Nov 1, 2024 · 1 mention

Building the Silicon Brain - Drew Houston of Dropbox Oct 18, 2024 · 1 mention

  • ▶ 13:51 Drew Houston I mean, it uses, the backend's somewhat interchangeable, so everything from, like, XLLAMA to VLLM or SGLANG, there's a bunch of these different backends you can use.

[Paper Club] Writing in the Margins: Chunked Prefill KV Caching for Long Context Retrieval Sep 19, 2024 · 6 mentions

  • ▶ 3:11 Umar Jamil Uh, so the first thing that the, the, the inference engine, which could be VLM, which could be Tessor RT, or any other inference, uh, um, uh, framework that you're using, the first thing that it does with your prompt is doing this… 2 times in the scene
  • ▶ 19:01 Umar Jamil You just, usually the KB cache allocation is a static, but even if you use VLLM, it's used, it's done using the, the so-called like pages attention.
  • ▶ 33:51 Umar Jamil Um, you can, you can actually, uh, or, well, okay, if you use, for example, VLM, they use this called, thing called the pages attention, so actually they prefill, uh, they allocate one entire page, which is actually a lot of tokens, so…
  • ▶ 48:30 Umar Jamil So as you, as you can see from VLLM, they have this, a feature is an experimental feature right now in VLLM. 2 times in the scene

Answer.ai & AI Magic with Jeremy Howard Aug 17, 2024 · 2 mentions

  • ▶ 22:43 Jeremy Howard Like, two or three weeks after we did FSTP, Qlory just popped up and said, okay, I've just converted the whole thing to Dora, and I've also created these VLLM extensions, and I've got all these benchmarks, and, you know, now I've got, um,…
  • ▶ 45:25 Jeremy Howard It's difficult, you know, because like, Karen's been doing a lot of work with, with VLM, for example.

[LLM Paper Club] Llama 3.1 Paper: The Llama Family of Models Jul 29, 2024 · 1 mention

  • ▶ 14:05 unnamed speaker Other tweets, George hots Carpathy tweeted, um, VLL supports it.

A Brief History of the Open Source AI Hacker - with Ben Firshman of Replicate Feb 28, 2024 · 2 mentions

  • ▶ 1:05:49 unnamed speaker Were you referencing VLM? 2 times in the scene

Building an open AI company - with Ce and Vipul of Together AI Feb 8, 2024 · 1 mention

  • ▶ 1:10:55 Ce Zhang And, and all the way to motion learning systems, right, if you want to, like, like to hack over, like, VRM, TGI, right, that's great.
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.