SGLAN

product on 1 show · 2 statements across 1 episodes

Latent Space

2 statements about SGLAN, every show

LATENT SPACE Assertion Supported
Zhang: SGLang achieved 3x throughput over vLLM in mid-2024 benchmarks
“At that time, I think its performance is maybe three times, is throughout, put it, three times than FLM.”
Yining Zhang Jan 19, 2025 ▶ 34:05 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
LATENT SPACE Assertion Supported
SGLang was the first framework to support prefix caching
“At 2024, January, they support something like Redix cache. It's a prefix caching technology. I think SGLAN is the first framework that supports prefix cache.”
Yining Zhang Jan 19, 2025 ▶ 33:14 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.