Llama 70B, every mention

8 scenes · ← back to Llama 70B

tap a year for its mentions
0052103202320242025episodesmentions
023202320242025episodes it came up in
001.51.533202320242025episodesmentions per episode

every year anyone Yining Zhang 3William Beauchamp 3Rishabh Agarwal 2Sarah Chieng 1Dylan Patel 1

Verbatim, from the transcripts: the passages where Llama 70B comes up

loading…

The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind Mar 23, 2025 · 2 mentions

  • ▶ 17:12 Rishabh Agarwal And for a fixed number of tokens, if you have a much smaller model, so let's say lama-seventyb versus lama-seventyb, you can generate 10 times more data in principle, right? 2 times in the scene

Outlasting Noam Shazeer, Crowdsourcing Chai AI w/ 1.4m DAU — with William Beauchamp, Chai Research Jan 26, 2025 · 3 mentions

  • ▶ 56:30 William Beauchamp What the money let them do was, if they wanted to fine-tune Alarma-seventyb on eight H-one-hundreds overnight, if you give them money, then they can do it.
  • ▶ 1:03:09 William Beauchamp Right, and then you can give that to, like, a Llama-seventyb. 2 times in the scene

DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing) Jan 19, 2025 · 4 mentions

  • ▶ 4:39 Yining Zhang I think at the base time, something like LAMA-Seventy-B is more common. 3 times in the scene
  • ▶ 10:08 unnamed speaker And so when it comes to the quantization question, uh, where it matters is that we would never quantize the model behind, you know, the user's back and say, look at us, there's a faster Lama, uh, 70 B that has been, you know, uh, somehow…

[Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras) Dec 7, 2024 · 1 mention

  • ▶ 3:06 Sarah Chieng And, you know, as Swix mentioned in the chat a few weeks ago, Cerebris came out that the wafer scale engine three can serve llama 70 B at 2.1 thousand, um, sorry, 202,100 tokens per second and serves llama four or five B at nearly 1000…

LLM Asia Paper Club Survey Round May 22, 2024 · 2 mentions

  • ▶ 46:10 unnamed speaker So, I think we have a huge model, like a Lama-seventyb, and we have a smaller model, which we call a draft model, called a Lama-seventyb. 2 times in the scene

The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis Dec 5, 2023 · 1 mention

Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.