“The reason why Lama open-sourced the MOE model, because I think they tried to train our MOE model, but they failed. So that, that's why they didn't open source MOE mode for Lama series.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Yining Zhang
AssertionSupported
Zhang: XGrammar outperforms Outlines and is integrated into TensorRT-LLM
“And I think Xgrammar's performance is better than the outline's, and also in the TensorFlow RTLM, the latest release, TensorFlow RTLM also integrates Xgrammar as the backend for the constructed coding.”
Yining ZhangJan 19, 2025▶ 41:22DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
“Yeah, I think even they use the FP-A to quantization, the benchmark result is very good, such as something like GSM-HK. The score is nearly 94.6. It's so high, you know. I think it's higher than every other open source AIM, even the LAMA 400 zero five billion.”
Yining ZhangJan 19, 2025▶ 12:38DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
AssertionNot checkable as stated
Zhang: Baidu and ByteDance Internal Models Use DeepSeek-Like MoE Architectures
“As far as I know, some companies such as Baidu or Baidu Dance, they are internal, the dominant AOM, they use the MOE architecture, and their ways, I think, is similar to the DeepSeq MOE model.”
Yining ZhangJan 19, 2025▶ 13:41DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Opinion
Zhang: SGLang outperforms vLLM and has better usability than TensorRT-LLM
“I think for the common use case, maybe not, not the DeepSeq VIII, for the common use case, I think SGLAN's performance is better than FLM, and its usability is better than TensorFlow TLM.”
Yining ZhangJan 19, 2025▶ 26:57DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
AssertionSupported
Zhang: SGLang achieved 3x throughput over vLLM in mid-2024 benchmarks
“At that time, I think its performance is maybe three times, is throughout, put it, three times than FLM.”
Yining ZhangJan 19, 2025▶ 34:05DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Opinion
Zhang: DeepSeek-V3 is currently the leading open-source LLM
“Yeah, because DeepSeq VIII, I think, is currently considered the leading open source LLMs based on the benchmark results and the chat area results.”
Yining ZhangJan 19, 2025▶ 1:22DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 200 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.