Llama 405B, every mention

4 scenes · ← back to Llama 405B

tap a year for its mentions
00315120242025episodesmentions
01120242025episodes it came up in
002.50.55120242025episodesmentions per episode

every year anyone Yining Zhang 4Sarah Chieng 1

Verbatim, from the transcripts: the passages where Llama 405B comes up

loading…

DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing) Jan 19, 2025 · 5 mentions

  • ▶ 4:45 Yining Zhang I think LAMA-Seventy-B has released the 400 zero five billion weights, but I think there are just a few users use that. 3 times in the scene
  • ▶ 12:52 Yining Zhang I, I think it's higher than every other open source AIM, even the LAMA 400 zero five billion.
  • ▶ 14:14 unnamed speaker Like, I think, well, Lama four or five B is dense.

[Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras) Dec 7, 2024 · 1 mention

  • ▶ 3:06 Sarah Chieng And, you know, as Swix mentioned in the chat a few weeks ago, Cerebris came out that the wafer scale engine three can serve llama 70 B at 2.1 thousand, um, sorry, 202,100 tokens per second and serves llama four or five B at nearly 1000…
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.