vLLM, every mention
12 scenes · ← back to vLLM
tap a year for its mentions
every year anyone Simon Moe 12Elena Burger 9Matt Bornstein 1Ion Stoica 1Dylan Patel 1Arthur Mensch 1
Verbatim, from the transcripts: the passages where vLLM comes up
How Open Source Became AI's Backbone | Inferact with a16z
- ▶ 0:17 Elena Burger Can you talk about where VLOM sits in that stack? 4 times in the scene
- ▶ 1:01 Elena Burger Today we're here with Simon Moe, co-founder of Infraact, and a lead maintainer of VLLM, the open source inference engine, now running on half a million GPUs at any moment.
- ▶ 1:27 Elena Burger So VLLM actually has its origins, kind of,
- ▶ 4:47 Matt Bornstein And so, yeah, so, so look, I mean, um, you know, BERT was an early language model, um, that, uh, you know, newer models are much bigger, much more sophisticated, take up a lot more memory, a lot more compute, and, and, and, um, you know,…
- ▶ 8:31 Simon Moe Just about everybody uses VLLM. 10 times in the scene
- ▶ 19:13 Simon Moe Like what I'm talking about here, of course, is VLON's own fast mode, getting up to 405 hundred tokens per second, because it is really a big step change from like, especially when developers are interacting with the model, they can see,
- ▶ 30:11 Elena Burger and, like, looking back at not just the history of VLLM and Infraact, but also a company like Open Router or even Olama, all of these different teams kind of got started around twenty-twenty-two and twenty-twenty-three. 2 times in the scene
- ▶ 33:51 Simon Moe And so a lot of our developer within Infrax and for VLM are like retreating from using Fable five because you, you have a two hour job and you trigger the, the red line, which is false positive.
- ▶ 36:40 Elena Burger Just about Infraact and, you know, running the company, and I know that Ian Stoick of Databricks is an advisor and a co-founder of Infraact, and I'm just curious what you've learned from him, ah, in terms of taking an open source project…
Dylan Patel on GPT-5’s Router Moment, GPUs vs TPUs, Monetization
- ▶ 25:33 Dylan Patel But, but, like, you guys don't have one of these, like, you know, base 10 or any of these, like, sort of, like, API investments because you think, um, this is from someone on the infra team that you guys think it'll get commoditized…
Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
- ▶ 45:28 Ion Stoica Again, since then, there are many other things we've done here, and open source, like inference, uh, LLM inference, like VLLM and HLANG, but I do think that what happens in this kind of, this sense in, um, in, um, in industry,
Safety in Numbers: Keeping AI Open
- ▶ 13:07 Arthur Mensch You do need to do inferencing efficiently, and that's also the reason why we released an open source package based on VLLM so that the community can, can take also this code and modify it and see how that works.