TensorRT-LLM

2 statements across 1 episodes · 1 bullish · 0 bearish · 1 people on the record · first statement Jan 19, 2025 by Yining Zhang · said 27 times in 6 episodes since 2024 · across every show →

Mentions by year

brought up most by Yining Zhang (10), Kyle Kranen (2), Chris Lattner (2), Stefano Ermon (1), Ben Firshman (1), Ali Taha (1)

tap a year for its mentions
00132253202420252026episodesmentions
023202420252026episodes it came up in
0041.583202420252026episodesmentions per episode

every mention, scene by scene, with the transcript →

Everything said about TensorRT-LLM, oldest first

Jan 19, 2025 positive
Opinion
Zhang: SGLang outperforms vLLM and has better usability than TensorRT-LLM
“I think for the common use case, maybe not, not the DeepSeq VIII, for the common use case, I think SGLAN's performance is better than FLM, and its usability is better than TensorFlow TLM.”
Yining Zhang Jan 19, 2025 ▶ 26:57 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Jan 19, 2025 neutral
Assertion Supported
Zhang: TensorRT-LLM supports Eagle 1 speculative decoding, not Eagle 2
“Currently, even use the TanzRTM, it only supported Eagle One, not Eagle Two.”
Yining Zhang Jan 19, 2025 ▶ 45:36 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.