Llama 70B

part of Llama

3 statements across 3 episodes · 1 bullish · 1 bearish · 3 people on the record · first statement Dec 5, 2023 by Dylan Patel · said 13 times in 6 episodes since 2023 · across every show →

Mentions by year

brought up most by Yining Zhang (3), William Beauchamp (3), Rishabh Agarwal (2), Sarah Chieng (1), Dylan Patel (1)

tap a year for its mentions
0052103202320242025episodesmentions
023202320242025episodes it came up in
001.51.533202320242025episodesmentions per episode

every mention, scene by scene, with the transcript →

Everything said about Llama 70B, oldest first

Dec 5, 2023 positive
Assertion Supported
Patel: Google TPU v5e is half the size of TPU v5p
“TPUv-V-V-E is, like the new one, but it's mostly, mostly an inference chip. It's a small chip. It's a little bit, it's about half the size of a TPUv-V-V. That chip, you know, you can get very good performance on, like, Of Lama, 70 B inference, right?”
Dylan Patel Dec 5, 2023 ▶ 11:49 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Jan 19, 2025 negative
Assertion Not checkable as stated
Zhang: Llama 405B sees very few enterprise users compared to 70B
“I think at the base time, something like LAMA-Seventy-B is more common. I think LAMA-Seventy-B has released the 400 zero five billion weights, but I think there are just a few users use that.”
Yining Zhang Jan 19, 2025 ▶ 4:39 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Jul 13, 2026 neutral
Assertion Partly supported
Biderman: Processing a Wikipedia article in Llama 70B consumes 80GB HBM
“If you take a Lama, a 70 B model, and you load one article from Wikipedia, which is a few tens of kilobytes, and you have the model read this The brain state of the model when reading this few tens of kilobytes is like, 80 gigabytes. 80 gigabytes on, on the HB…”
Dan Biderman Jul 13, 2026 ▶ 22:39 The AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.