Llama 70B, every mention
8 scenes · ← back to Llama 70B
tap a year for its mentions
every year anyone Yining Zhang 3William Beauchamp 3Rishabh Agarwal 2Sarah Chieng 1Dylan Patel 1
Verbatim, from the transcripts: the passages where Llama 70B comes up
The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
- ▶ 17:12 Rishabh Agarwal And for a fixed number of tokens, if you have a much smaller model, so let's say lama-seventyb versus lama-seventyb, you can generate 10 times more data in principle, right? 2 times in the scene
Outlasting Noam Shazeer, Crowdsourcing Chai AI w/ 1.4m DAU — with William Beauchamp, Chai Research
- ▶ 56:30 William Beauchamp What the money let them do was, if they wanted to fine-tune Alarma-seventyb on eight H-one-hundreds overnight, if you give them money, then they can do it.
- ▶ 1:03:09 William Beauchamp Right, and then you can give that to, like, a Llama-seventyb. 2 times in the scene
DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
- ▶ 4:39 Yining Zhang I think at the base time, something like LAMA-Seventy-B is more common. 3 times in the scene
- ▶ 10:08 unnamed speaker And so when it comes to the quantization question, uh, where it matters is that we would never quantize the model behind, you know, the user's back and say, look at us, there's a faster Lama, uh, 70 B that has been, you know, uh, somehow…
[Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras)
- ▶ 3:06 Sarah Chieng And, you know, as Swix mentioned in the chat a few weeks ago, Cerebris came out that the wafer scale engine three can serve llama 70 B at 2.1 thousand, um, sorry, 202,100 tokens per second and serves llama four or five B at nearly 1000…
LLM Asia Paper Club Survey Round
- ▶ 46:10 unnamed speaker So, I think we have a huge model, like a Lama-seventyb, and we have a smaller model, which we call a draft model, called a Lama-seventyb. 2 times in the scene
The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
- ▶ 12:12 Dylan Patel Of, of Lama, 70 B inference, right?