vLLM, every mention
41 scenes · ← back to vLLM
tap a year for its mentions
every year anyone Chris Lattner 11Shawn Wang 10Umar Jamil 6Yining Zhang 5William Beauchamp 3Philip Kiely 3Kyle Kranen 3Omar Sanseviero 2Jim Zemlin 2Jeremy Howard 2
Verbatim, from the transcripts: the passages where vLLM comes up
Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
- ▶ 13:51 Philip Kiely Getting to the point of I can make a token out of this model is not that hard because generally the, um, open source inference engines, your VLMs, SGLangs of the world oftentimes even receive weights ahead of time maintainers do, or the…
- ▶ 24:12 Ali Taha But if you were to switch to VLM, that isn't the case.
- ▶ 42:12 Alessio Fanelli Put it behind VLO.
- ▶ 1:00:40 Philip Kiely So if you look at, like, the original VLM and SGLang, or VLM especially, like, that was written targeting Ampure and then had to be updated for Hopper, updated for Blackwell. 2 times in the scene
The AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO
- ▶ 44:40 Dan Biderman Leads at Databricks, and one of the core contributors at VLLM, and we're all kind of like, ah, systems inclined, but we think infrastructures, engineers, ah, people who know how to work, ah,
The Future of AI Infra: from Kubernetes to Agent Sandboxes — Akshat Bubna, Modal CTO
- ▶ 21:30 unnamed speaker How does it compare to, say, I take the same model, GLM, 5.2 FBA, take off the shelf inference engine, VLM, SGLang, um, you know, get compute of similar capacity, similar cost.
⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
- ▶ 3:42 Omar Sanseviero So for example, we work with Lama CPP, Olama, MLX, Hogan Faces, BLM, NVIDIA, AMD. 2 times in the scene
Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
- ▶ 28:05 Kyle Kranen Dynoa sort of came about at NVIDIA because myself and a couple others were sort of talking about these concepts that like, you know, you have inference engines like VLM, SGLang, TensorRTLM, um, 3 times in the scene
One Year of MCP — with David Soria Parria and AAIF leads from OpenAI, Goose, Linux Foundation
- ▶ 1:25:17 Jim Zemlin Now you've got interesting technology, VLLM, Ray, 2 times in the scene
A Technical History of Generative Media
- ▶ 11:36 unnamed speaker And there's no community effort, like, a VLM?
⚡️Mercury: Ultra-Fast Diffusion LLMs — Estefano Ermon, CEO Inception Labs
- ▶ 21:33 Stefano Ermon So just like you would normally serve an LLM using a VLLM or SGLang or a Tensor or TLLM, we have built our own inference engine.
🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
- ▶ 2:34:10 Shawn Wang You can't just use VLLM and TensorFlow. 2 times in the scene
Information Theory for Language Models: Jack Morris
- ▶ 14:10 Jack Morris I also think, ah, VLLM and SGLang seem, like, really good and important and here to stay.
The Shape of Compute (Chris Lattner of Modular)
- ▶ 11:34 Chris Lattner So it's not as good as something like VLLM because it's missing some features, and it only supports NVIDIA and AMD hardware, for example. 4 times in the scene
- ▶ 23:09 Shawn Wang I'm curious if you have any views or insider takes on what's happening with VLM versus SGLang and everything coming out of Berkeley. 6 times in the scene
- ▶ 35:59 Chris Lattner And so this is why we have things like VLM and SGLang because they're the black box that you can just hopefully build on top of and not have to know how any of that scary stuff works is because we haven't taught the industry how to do this… 2 times in the scene
- ▶ 41:06 Chris Lattner Well, again, you get back into hacking the internals of VLM and PyTorch isn't really designed for KV cache optimizations and all the modern transformer features and things like this. 2 times in the scene
- ▶ 56:02 Chris Lattner and it drew attention to that layer of the stack because it wasn't just like throwing layers on top of PyTorch, you know, or VLM, or it was like doing that fun amount of work. 3 times in the scene
Outlasting Noam Shazeer, Crowdsourcing Chai AI w/ 1.4m DAU — with William Beauchamp, Chai Research
- ▶ 1:07:54 William Beauchamp We were using, we were running VLM for a while and VLM is really fantastic. 3 times in the scene
DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
- ▶ 20:36 unnamed speaker We have customers on, on based end that are using TensorFlow and we have ones that are using VL and we have a growing number that are using SGLang too.
- ▶ 26:35 unnamed speaker So you have SGLang, TRTLLM, VLLM. 4 times in the scene
- ▶ 32:01 unnamed speaker And that's something that I'm seeing in the market that, like, people who are somewhat new to it, they're like, well, VLM equals equals production. 2 times in the scene
- ▶ 33:52 Yining Zhang Equivalent with the FLM or with the Tencent RTLM. 4 times in the scene
- ▶ 36:03 Yining Zhang Redix cache, I, I think it's the technology of the prefix caching, and it is a special case for something like block size is one, you know, for VLM or for other frameworks, they use something block size, 32, and SGLAN use the block size…
- ▶ 43:17 unnamed speaker It seems like VLM obviously has the, the, it's one year older. 3 times in the scene
- ▶ 50:55 unnamed speaker Honestly, I would go back to what I emphasized earlier, which was that I wish more people asked about what it takes to run mission critical inference workloads, because I see this in the market sometimes that they're like, well, I can just… 2 times in the scene
Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
- ▶ 2:44 Loubna Ben Allal Some examples are VLM TGI and sensor RT.
Windsurf: The Enterprise AI IDE
- ▶ 57:22 unnamed speaker You can't just use VLLM and TensorFlow. 2 times in the scene
Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI
- ▶ 13:06 Shawn Wang Basically, it's just like the meta version of whatever HuggingFace offers, you know, or TensorRT, or BLM, or whatever the open source opportunity is.
- ▶ 23:37 Shawn Wang I don't remember the, the speed numbers, but apparently much better than BLM, especially on a concurrency basis.
In the Arena: How LMSys changed LLM Benchmarking Forever
- ▶ 9:46 Anastasios Angelopoulos They came to use VLLM inference, right?
Building the Silicon Brain - Drew Houston of Dropbox
- ▶ 13:51 Drew Houston I mean, it uses, the backend's somewhat interchangeable, so everything from, like, XLLAMA to VLLM or SGLANG, there's a bunch of these different backends you can use.
[Paper Club] Writing in the Margins: Chunked Prefill KV Caching for Long Context Retrieval
- ▶ 3:11 Umar Jamil Uh, so the first thing that the, the, the inference engine, which could be VLM, which could be Tessor RT, or any other inference, uh, um, uh, framework that you're using, the first thing that it does with your prompt is doing this… 2 times in the scene
- ▶ 19:01 Umar Jamil You just, usually the KB cache allocation is a static, but even if you use VLLM, it's used, it's done using the, the so-called like pages attention.
- ▶ 33:51 Umar Jamil Um, you can, you can actually, uh, or, well, okay, if you use, for example, VLM, they use this called, thing called the pages attention, so actually they prefill, uh, they allocate one entire page, which is actually a lot of tokens, so…
- ▶ 48:30 Umar Jamil So as you, as you can see from VLLM, they have this, a feature is an experimental feature right now in VLLM. 2 times in the scene
Answer.ai & AI Magic with Jeremy Howard
- ▶ 22:43 Jeremy Howard Like, two or three weeks after we did FSTP, Qlory just popped up and said, okay, I've just converted the whole thing to Dora, and I've also created these VLLM extensions, and I've got all these benchmarks, and, you know, now I've got, um,…
- ▶ 45:25 Jeremy Howard It's difficult, you know, because like, Karen's been doing a lot of work with, with VLM, for example.
[LLM Paper Club] Llama 3.1 Paper: The Llama Family of Models
- ▶ 14:05 unnamed speaker Other tweets, George hots Carpathy tweeted, um, VLL supports it.
A Brief History of the Open Source AI Hacker - with Ben Firshman of Replicate
- ▶ 1:05:49 unnamed speaker Were you referencing VLM? 2 times in the scene