Yining Zhang

Senior Director, Inference, Together AI · 1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

engineerexecutiveoperator@zhyncs42 ↗LinkedIn ↗zhyncs.com ↗

Zhang is a core maintainer of the open-source serving framework SGLang and co-creator of TokenSpeed, specializing in GPU kernel optimization and LLM serving. Previously a lead model performance engineer at Baseten, he currently serves as Senior Director of Inference at Together AI.

18statements → 15claims → 10claims resolved → 90%fully supported → 3.44/5average certainty → 1.94/5average debate potential →

9 supported 0 partly supported 1 contradicted 5 not checkable as stated how the 15 claims stand · each chip opens the sources

1 prediction · 14 assertions · 3 opinions · every statement was checked. The prediction and assertions are the 15 claims: statements the public record can support or contradict. 10 are resolved, and 5 name no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Yining argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
Zhang: XGrammar outperforms Outlines and is integrated into TensorRT-LLM
“And I think Xgrammar's performance is better than the outline's, and also in the TensorFlow RTLM, the latest release, TensorFlow RTLM also integrates Xgrammar as the backend for the constructed coding.”
Yining Zhang Jan 19, 2025 ▶ 41:22 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)

Their most notable contradicted claim

Assertion Contradicted
Yining Zhang: DeepSeek V3 scores 94.6 on GSM8K, outperforming Llama 405B
“Yeah, I think even they use the FP-A to quantization, the benchmark result is very good, such as something like GSM-HK. The score is nearly 94.6. It's so high, you know. I think it's higher than every other open source AIM, even the LAMA 400 zero five billion.”
Yining Zhang Jan 19, 2025 ▶ 12:38 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)

Expressed certainty vs assessment result

none yet certainty 1
none yet certainty 2
100% certainty 3
83% certainty 4
none yet certainty 5

weighted support: a fully supported claim counts one, a partly supported claim counts half. Each filled bar is clickable and opens exactly those claims; "none yet" means nothing said at that certainty level has resolved yet

Everything Yining Zhang said on Latent Space that made the record, most notable first. Filter by type, assessment or year in the ledger →

Assertion Not checkable as stated
Zhang: Meta Failed at Training MoE Models for Llama Series
“The reason why Lama open-sourced the MOE model, because I think they tried to train our MOE model, but they failed. So that, that's why they didn't open source MOE mode for Lama series.”
Yining Zhang Jan 19, 2025 ▶ 14:53 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Assertion Supported
Zhang: XGrammar outperforms Outlines and is integrated into TensorRT-LLM
“And I think Xgrammar's performance is better than the outline's, and also in the TensorFlow RTLM, the latest release, TensorFlow RTLM also integrates Xgrammar as the backend for the constructed coding.”
Yining Zhang Jan 19, 2025 ▶ 41:22 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Assertion Contradicted
Yining Zhang: DeepSeek V3 scores 94.6 on GSM8K, outperforming Llama 405B
“Yeah, I think even they use the FP-A to quantization, the benchmark result is very good, such as something like GSM-HK. The score is nearly 94.6. It's so high, you know. I think it's higher than every other open source AIM, even the LAMA 400 zero five billion.”
Yining Zhang Jan 19, 2025 ▶ 12:38 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Assertion Not checkable as stated
Zhang: Baidu and ByteDance Internal Models Use DeepSeek-Like MoE Architectures
“As far as I know, some companies such as Baidu or Baidu Dance, they are internal, the dominant AOM, they use the MOE architecture, and their ways, I think, is similar to the DeepSeq MOE model.”
Yining Zhang Jan 19, 2025 ▶ 13:41 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Opinion
Zhang: SGLang outperforms vLLM and has better usability than TensorRT-LLM
“I think for the common use case, maybe not, not the DeepSeq VIII, for the common use case, I think SGLAN's performance is better than FLM, and its usability is better than TensorFlow TLM.”
Yining Zhang Jan 19, 2025 ▶ 26:57 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Assertion Supported
Zhang: SGLang achieved 3x throughput over vLLM in mid-2024 benchmarks
“At that time, I think its performance is maybe three times, is throughout, put it, three times than FLM.”
Yining Zhang Jan 19, 2025 ▶ 34:05 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Opinion
Zhang: DeepSeek-V3 is currently the leading open-source LLM
“Yeah, because DeepSeq VIII, I think, is currently considered the leading open source LLMs based on the benchmark results and the chat area results.”
Yining Zhang Jan 19, 2025 ▶ 1:22 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Assertion Not checkable as stated
Zhang: Llama 405B sees very few enterprise users compared to 70B
“I think at the base time, something like LAMA-Seventy-B is more common. I think LAMA-Seventy-B has released the 400 zero five billion weights, but I think there are just a few users use that.”
Yining Zhang Jan 19, 2025 ▶ 4:39 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Assertion Supported
Zhang: vLLM does not support DeepSeek MLA while SGLang does
“Something like DeepSeq V-II, they proposed attention parent named MLA, multi-latent attention, and I think SGLAN is the only framework to support that. Maybe LightLM and TRTM also support, but VLM doesn't support.”
Yining Zhang Jan 19, 2025 ▶ 27:23 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Assertion Supported
Zhang: DeepSeek team officially recommends SGLang as its inference engine
“That's why SGLAN is the recommended LLM engine by the DeepSeq team.”
Yining Zhang Jan 19, 2025 ▶ 28:06 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Assertion Supported
SGLang was the first framework to support prefix caching
“At 2024, January, they support something like Redix cache. It's a prefix caching technology. I think SGLAN is the first framework that supports prefix cache.”
Yining Zhang Jan 19, 2025 ▶ 33:14 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Opinion
Zhang: TensorRT-LLM is blazing fast but difficult to extend
“And TensorFlow RTM, I think it's blazing fast. Its performance is so good, but it's not easy to do some secondary development. If you want to add some new feature, it's a little hard.”
Yining Zhang Jan 19, 2025 ▶ 35:17 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Assertion Supported
Zhang: SGLang achieves higher cache hit rates using block size of one
“Redix cache, I think it's the technology of the prefix caching, and it is a special case for something like block size is one, you know, for VLM or for other frameworks, they use something block size, 32, and SGLAN use the block size one. I think if you use th…”
Yining Zhang Jan 19, 2025 ▶ 36:03 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Assertion Supported
Zhang: DeepSeek-V3 cannot run on a single 8xH100 GPU node
“You need, I think 671 gigabytes for the weights, and you also need an extra memory for the KV cache, so it's not possible to run that on H-one hundred.”
Yining Zhang Jan 19, 2025 ▶ 2:46 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Prediction Not checkable as stated
Zhang: MoE Inference Optimization Will Be Essential in 2025
“So I think at this new year, the MOE inference optimization will be very essential, and yeah, Yeah, so important.”
Yining Zhang Jan 19, 2025 ▶ 13:59 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Assertion Supported
Zhang: SGLang creators Lianmin Zheng and Ying Sheng work at xAI
“Lian Ming and Yin are the XAI's member of the technical staff.”
Yining Zhang Jan 19, 2025 ▶ 43:05 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Assertion Supported
Zhang: TensorRT-LLM supports Eagle 1 speculative decoding, not Eagle 2
“Currently, even use the TanzRTM, it only supported Eagle One, not Eagle Two.”
Yining Zhang Jan 19, 2025 ▶ 45:36 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Assertion Not checkable as stated
Zhang: Cursor Team Contacted SGLang Over DeepSeek-V3 Support
“And when we released the DeepSeq feed story support some employee from the Cursor team Also very interested in our implementation and ask, reach out and ask some questions from us.”
Yining Zhang Jan 19, 2025 ▶ 50:25 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)

Appearances (1)

EpisodeDateSpeaking time
DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pr Jan 19, 2025 13m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.