vLLM, every mention

12 scenes · ← back to vLLM

tap a year for its mentions
001312522023202420252026episodesmentions
0122023202420252026episodes it came up in
001312522023202420252026episodesmentions per episode

every year anyone Simon Moe 12Elena Burger 9Matt Bornstein 1Ion Stoica 1Dylan Patel 1Arthur Mensch 1

Verbatim, from the transcripts: the passages where vLLM comes up

loading…

How Open Source Became AI's Backbone | Inferact with a16z Aug 5, 2026 · 22 mentions

  • ▶ 0:17 Elena Burger Can you talk about where VLOM sits in that stack? 4 times in the scene
  • ▶ 1:01 Elena Burger Today we're here with Simon Moe, co-founder of Infraact, and a lead maintainer of VLLM, the open source inference engine, now running on half a million GPUs at any moment.
  • ▶ 1:27 Elena Burger So VLLM actually has its origins, kind of,
  • ▶ 4:47 Matt Bornstein And so, yeah, so, so look, I mean, um, you know, BERT was an early language model, um, that, uh, you know, newer models are much bigger, much more sophisticated, take up a lot more memory, a lot more compute, and, and, and, um, you know,…
  • ▶ 8:31 Simon Moe Just about everybody uses VLLM. 10 times in the scene
  • ▶ 19:13 Simon Moe Like what I'm talking about here, of course, is VLON's own fast mode, getting up to 405 hundred tokens per second, because it is really a big step change from like, especially when developers are interacting with the model, they can see,
  • ▶ 30:11 Elena Burger and, like, looking back at not just the history of VLLM and Infraact, but also a company like Open Router or even Olama, all of these different teams kind of got started around twenty-twenty-two and twenty-twenty-three. 2 times in the scene
  • ▶ 33:51 Simon Moe And so a lot of our developer within Infrax and for VLM are like retreating from using Fable five because you, you have a two hour job and you trigger the, the red line, which is false positive.
  • ▶ 36:40 Elena Burger Just about Infraact and, you know, running the company, and I know that Ian Stoick of Databricks is an advisor and a co-founder of Infraact, and I'm just curious what you've learned from him, ah, in terms of taking an open source project…

Dylan Patel on GPT-5’s Router Moment, GPUs vs TPUs, Monetization Aug 18, 2025 · 1 mention

  • ▶ 25:33 Dylan Patel But, but, like, you guys don't have one of these, like, you know, base 10 or any of these, like, sort of, like, API investments because you think, um, this is from someone on the infra team that you guys think it'll get commoditized…

Beyond Leaderboards: LMArena’s Mission to Make AI Reliable May 29, 2025 · 1 mention

  • ▶ 45:28 Ion Stoica Again, since then, there are many other things we've done here, and open source, like inference, uh, LLM inference, like VLLM and HLANG, but I do think that what happens in this kind of, this sense in, um, in, um, in industry,

Safety in Numbers: Keeping AI Open Dec 28, 2023 · 1 mention

  • ▶ 13:07 Arthur Mensch You do need to do inferencing efficiently, and that's also the reason why we released an open source package based on VLLM so that the community can, can take also this code and modify it and see how that works.
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.