Llama 405B, every mention
4 scenes · ← back to Llama 405B
tap a year for its mentions
every year anyone Yining Zhang 4Sarah Chieng 1
Verbatim, from the transcripts: the passages where Llama 405B comes up
DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
- ▶ 4:45 Yining Zhang I think LAMA-Seventy-B has released the 400 zero five billion weights, but I think there are just a few users use that. 3 times in the scene
- ▶ 12:52 Yining Zhang I, I think it's higher than every other open source AIM, even the LAMA 400 zero five billion.
- ▶ 14:14 unnamed speaker Like, I think, well, Lama four or five B is dense.
[Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras)
- ▶ 3:06 Sarah Chieng And, you know, as Swix mentioned in the chat a few weeks ago, Cerebris came out that the wafer scale engine three can serve llama 70 B at 2.1 thousand, um, sorry, 202,100 tokens per second and serves llama four or five B at nearly 1000…