TensorRT-LLM, every mention
13 scenes · ← back to TensorRT-LLM
tap a year for its mentions
every year anyone Yining Zhang 10Kyle Kranen 2Chris Lattner 2Stefano Ermon 1Ben Firshman 1Ali Taha 1
Verbatim, from the transcripts: the passages where TensorRT-LLM comes up
Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
- ▶ 28:05 Kyle Kranen Dynoa sort of came about at NVIDIA because myself and a couple others were sort of talking about these concepts that like, you know, you have inference engines like VLM, SGLang, TensorRTLM, um, 2 times in the scene
⚡️Mercury: Ultra-Fast Diffusion LLMs — Estefano Ermon, CEO Inception Labs
- ▶ 21:33 Stefano Ermon So just like you would normally serve an LLM using a VLLM or SGLang or a Tensor or TLLM, we have built our own inference engine.
The Shape of Compute (Chris Lattner of Modular)
- ▶ 24:05 Chris Lattner Well, and, and I mean, it's, as far as I know, it's like crushing TRT LLM and some of the other older systems and things like that. 2 times in the scene
DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
- ▶ 3:31 Yining Zhang So I think currently even the tensor RTM doesn't support the fp-eight.
- ▶ 18:51 unnamed speaker It has very native and deep support for TensorRTLM. 6 times in the scene
- ▶ 26:24 unnamed speaker Less so about, hey, here's my model, make sure you run it with TRT-LM or make sure we run with SGLank.
- ▶ 26:35 unnamed speaker So you have SGLang, TRTLLM, VLLM. 3 times in the scene
- ▶ 32:13 Yining Zhang I agree with Amir, because I think it's open source libraries such as VLM, SGLAN, LatLM, or TansRTM.
- ▶ 33:52 Yining Zhang Equivalent with the FLM or with the Tencent RTLM. 3 times in the scene
- ▶ 41:22 Yining Zhang And I, I think Xgrammar's performance is better than the outline's, and also in the TensorFlow RTLM, the latest release, TensorFlow RTLM also integrates Xgrammar as the backend for the constructed coding. 2 times in the scene
- ▶ 43:43 Yining Zhang I think if you care about the performance, maybe TensorFlow RTM is the best solution for now, especially for the latency sensitive scenery, TensorFlow RTM doing well. 3 times in the scene
A Brief History of the Open Source AI Hacker - with Ben Firshman of Replicate
- ▶ 1:06:35 Ben Firshman Um, we've, um, we've had success using inference servers like VLM and TRT LLM, um, and we're using those kind of things to serve language models.