SGLang, every mention

27 scenes · ← back to SGLang

tap a year for its mentions
00303605202420252026episodesmentions
035202420252026episodes it came up in
007.52.5155202420252026episodesmentions per episode

every year anyone Yining Zhang 6Shawn Wang 4Philip Kiely 4Kyle Kranen 3Mark Bissell 2Chris Lattner 2Ali Taha 2Stefano Ermon 1Jack Morris 1Drew Houston 1

Verbatim, from the transcripts: the passages where SGLang comes up

loading…

Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten Aug 3, 2026 · 6 mentions

  • ▶ 0:04 Ali Taha We had a GLM-Five-to endpoint that we were using to, like, that we plugged in in our cloud code harness, so every engineer on the same uses, like, our GLM-Five-to, and it will do a forward pass on the GLM-Five-to instance of the, you know,…
  • ▶ 13:51 Philip Kiely Getting to the point of I can make a token out of this model is not that hard because generally the, um, open source inference engines, your VLMs, SGLangs of the world oftentimes even receive weights ahead of time maintainers do, or the…
  • ▶ 24:09 Ali Taha Or oftentimes this will only happen in an inference engine that you're using, like SGLang.
  • ▶ 1:00:40 Philip Kiely So if you look at, like, the original VLM and SGLang, or VLM especially, like, that was written targeting Ampure and then had to be updated for Hopper, updated for Blackwell.
  • ▶ 1:33:15 Philip Kiely Yeah, I mean, that's not exactly a model optimizing its own inference so much as a model like being able to read the SGLang docs, but 2 times in the scene

The Future of AI Infra: from Kubernetes to Agent Sandboxes — Akshat Bubna, Modal CTO Jul 8, 2026 · 2 mentions

  • ▶ 21:30 unnamed speaker How does it compare to, say, I take the same model, GLM, 5.2 FBA, take off the shelf inference engine, VLM, SGLang, um, you know, get compute of similar capacity, similar cost. 2 times in the scene

Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup Mar 8, 2026 · 3 mentions

  • ▶ 28:05 Kyle Kranen Dynoa sort of came about at NVIDIA because myself and a couple others were sort of talking about these concepts that like, you know, you have inference engines like VLM, SGLang, TensorRTLM, um, 3 times in the scene

Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell Feb 5, 2026 · 2 mentions

  • ▶ 23:51 Mark Bissell It's got a forked version of our, uh, of the SGLang code base that we've been working on. 2 times in the scene

[State of Research Funding] Beyond NSF, Slingshots, Open Frontiers — Andy Konwinski, Laude Institute Dec 31, 2025 · 1 mention

  • ▶ 21:21 Andy Konwinski All the continual learning, prompt optimization, BLM, SGL, breakthroughs in inference, breakthroughs in evaluations, terminal bench, and a lot of other, I just heard for the first time about this benchmark called Impossible Bench.

⚡️Mercury: Ultra-Fast Diffusion LLMs — Estefano Ermon, CEO Inception Labs Aug 4, 2025 · 1 mention

  • ▶ 21:33 Stefano Ermon So just like you would normally serve an LLM using a VLLM or SGLang or a Tensor or TLLM, we have built our own inference engine.

Information Theory for Language Models: Jack Morris Jul 2, 2025 · 1 mention

  • ▶ 14:10 Jack Morris I also think, ah, VLLM and SGLang seem, like, really good and important and here to stay.

The Shape of Compute (Chris Lattner of Modular) Jun 13, 2025 · 6 mentions

  • ▶ 23:09 Shawn Wang I'm curious if you have any views or insider takes on what's happening with VLM versus SGLang and everything coming out of Berkeley. 4 times in the scene
  • ▶ 35:59 Chris Lattner And so this is why we have things like VLM and SGLang because they're the black box that you can just hopefully build on top of and not have to know how any of that scary stuff works is because we haven't taught the industry how to do this…
  • ▶ 42:39 Chris Lattner I'd love to see SGLang adopt to max.

DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing) Jan 19, 2025 · 43 mentions

  • ▶ 0:52 unnamed speaker You also are very involved in SGLang, and that was actually one of the reasons we were discussing an episode with you even before 2 times in the scene
  • ▶ 18:36 unnamed speaker And then we can talk about SG Lang 6 times in the scene
  • ▶ 24:30 unnamed speaker And then just to maybe tie this into SGLang, how do you kind of think about the hidden magic? 3 times in the scene
  • ▶ 26:35 unnamed speaker So you have SGLang, TRTLLM, VLLM. 8 times in the scene
  • ▶ 32:13 Yining Zhang I agree with Amir, because I think it's open source libraries such as VLM, SGLAN, LatLM, or TansRTM.
  • ▶ 32:27 unnamed speaker SGLang unique thing. 8 times in the scene
  • ▶ 35:43 unnamed speaker Let's run through maybe the three main techniques behind SGLang. 2 times in the scene
  • ▶ 38:51 Yining Zhang I think SGLAN support concentrated decoding, and it also support jump forward, and we use something like outline or the X grammar to do the, something like change the, convert the schema from, from JSON to the FSM, the state machine, and… 4 times in the scene
  • ▶ 42:51 unnamed speaker Tracing this human path, I'm pretty sure I know the answer, but is there a reason, like, big projects like, uh, Grok, you know, like XAI also use SGLang? 7 times in the scene
  • ▶ 50:38 Yining Zhang So I, I think as SGLAN grows faster and the features optimization we, we iterate so fast, and I, I think there will be more users from different companies, from different teams to, to use it.
  • ▶ 55:36 unnamed speaker I think this is a really good dive into both BaseNet and SG Lang, and a little bit of DeepSig V three, which people are very interested in.

In the Arena: How LMSys changed LLM Benchmarking Forever Nov 1, 2024 · 1 mention

  • ▶ 38:23 Anastasios Angelopoulos Sort of Chatbot Arena has, of course, like, kind of become its own thing, and Lianmin and Ying, who are, you know, created LMSYS, have kind of, like, moved on to working on SGLang, and now

Building the Silicon Brain - Drew Houston of Dropbox Oct 18, 2024 · 1 mention

  • ▶ 13:51 Drew Houston I mean, it uses, the backend's somewhat interchangeable, so everything from, like, XLLAMA to VLLM or SGLANG, there's a bunch of these different backends you can use.
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.