Llama

also referred to as: llama models

includes Llama 3, Llama 2, Llama 4, Llama 3.1, Llama 1, Llama 70B, Llama 405B, Llama Stack, Llama 3.2, Llama 3.3, Llama 7B, Llama Guard and 5 more

23 statements across 21 episodes · 7 bullish · 7 bearish · 20 people on the record · first statement Dec 5, 2023 by Dylan Patel · said 455 times in 94 episodes since 2023 · across every show →

Mentions by year, the whole family

brought up most by Shawn Wang (63), Alessio Fanelli (35), Thomas Scialom (22), Nathan Lambert (21), Soumith Chintala (15), George Hotz (13), Yining Zhang (9), Lin Qiao (7)

tap a year for its mentions
0015025300502023202420252026episodesmentions
025502023202420252026episodes it came up in
004258502023202420252026episodesmentions per episode
2026 28 mentions in 15 episodes 2 per episode
2025 90 mentions in 28 episodes 3 per episode
2024 299 mentions in 43 episodes 7 per episode
2023 38 mentions in 8 episodes 5 per episode

every mention, scene by scene, with the transcript →

Everything said about Llama, oldest first

Dec 5, 2023 negative
Assertion Supported
Patel: Hugging Face libraries achieve only 15% MBU for inference
“Hugging Face's libraries are actually very inefficient, like incredibly inefficient for inference. You get like, 15% MBU on, on, on, on some configurations, like eight, eight, eight, eight, eight, eight, 100, and LLAMA-seventy-beat, you get like, 15%, which is…”
Dylan Patel Dec 5, 2023 ▶ 18:40 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Dec 5, 2023
Assertion Supported
Patel: Running LLaMA-70B at reading speed requires 2.1 TB/s memory bandwidth
“Hey, to run Llama's seventy billion requires two terabytes a second of memory bandwidth, 2.1, at reading, human reading speed.”
Dylan Patel Dec 5, 2023 ▶ 26:06 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Jan 11, 2024 neutral
Insight
Lambert: RLHF Performance Depends on Data and Systems Over RL Details
“It really ends up being kind of, like, gibberish that I think is less important now, because it's more about data and infrastructure than RL details, than, like, value functions and everything. A lot of the papers have different terms in the equations. I think…”
Nathan Lambert Jan 11, 2024 ▶ 57:58 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Feb 7, 2024 bullish
Prediction Not checkable as stated
Future Llama models will match frontier GPT models if progress slows
“As AI progress slows down, so if we get like Llama-IV, Llama-V for example, maybe it's a comparable at that point, like GPT-V or GPT-VI, like, It made it to the point where it was like, look, I just want to use Lama. Like, it's, you know, safe for me to, you k…”
David Hsu Feb 7, 2024 ▶ 50:30 The State of AI in production — with David Hsu of Retool
Mar 6, 2024 neutral
Opinion
Chintala: Time and data constrain Meta LLM releases more than GPUs
“So, I think the, it's all a matter of time. I think time is the biggest bottleneck. It's like, when do you stop training the previous one, and when do you start training the next one? And how do you make those decisions? The data, do you have net new data, bet…”
Soumith Chintala Mar 6, 2024 ▶ 46:46 Open Source AI is AI we can Trust — with Soumith Chintala of Meta AI
Mar 27, 2024 bearish
Prediction Not checkable as stated
Pure-play foundation model companies will be commoditized by Llama and big tech
“I think pure play foundation model companies are just gonna be pinched by how good The next couple of llamas are going to be, and the next, like, what next good open source thing, and then seeing the really big players put ridiculous amounts of compute behind …”
David Luan Mar 27, 2024 ▶ 42:31 Why Google failed to make GPT-3 -- with David Luan of Adept
Jun 11, 2024 negative
Assertion Not checkable as stated
Conover: Commercial LLMs struggle to generate 5,000 output tokens in one generation
“There is a characteristic output length for these models. Let's say it's about 1200 tokens. Like it is very difficult to get any of the commercial LMs or LLAMA to write 5000 tokens.”
Mike Conover Jun 11, 2024 ▶ 14:33 How AI is Eating Finance - with Mike Conover of Brightwave
Jul 5, 2024 neutral
Insight
Yi Tay: Meta's Llama is corporate open weights, not grassroots open source
“To me, Lama Tree is like... Meta has an org that is hypothetically very similar to Gemini or something but they just decide to release the weights It's open weights It's open weights and everything”
Yi Tay Jul 5, 2024 ▶ 1:59:19 The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
Oct 18, 2024
Disclosure
Drew Houston brings an external GPU on planes to run Llama locally
“When I'm on a plane or something or where, like, you don't have access or the Internet's not reliable, I actually bring a gaming laptop on the plane with me. It's, like, a little, like, blue briefcase-looking thing, and then I, like, literally hook up a GPU, l…”
Drew Houston Oct 18, 2024 ▶ 13:18 Building the Silicon Brain - Drew Houston of Dropbox
Nov 1, 2024 neutral
Assertion Supported
Angelopoulos: The Chatbot Arena leaderboard is currently not an apples-to-apples comparison
“None of the leaderboard currently is apples to apples, because you have, like, Gemini Flash, you have, you know, all sorts of tiny models, like Llama Like, eight B and four or five B are not apples to apples.”
Anastasios Angelopoulos Nov 1, 2024 ▶ 28:03 In the Arena: How LMSys changed LLM Benchmarking Forever
Dec 7, 2024 positive
Assertion Supported
Cerebras WSE-3 runs Llama inference 70x faster than NVIDIA GPUs
“Cerebris came out that the wafer scale engine three can serve llama 70 B at 2.1 thousand sorry, 202,100 tokens per second and serves llama four or five B at nearly 1000 tokens per second. So this, you know, to give you an understanding, like this is about 70 t…”
Sarah Chieng Dec 7, 2024 ▶ 3:06 [Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras)
Dec 23, 2024 negative
Assertion Supported
Soldani: Llama and Qwen models fail OSI open source AI definition
“Under this definition, for example, Lama or some of the Quen models are not open source because the license says you can, you can't use this model for this, or it says if you use this model, you have to name the output this way or derivative needs to be named …”
Luca Soldani Dec 23, 2024 ▶ 7:07 Best of 2024: Open Models [LS LIVE! at NeurIPS 2024]
Dec 24, 2024 bullish
Assertion Supported
Ben Allal: Hugging Face SmolLM2-1.7B outperforms Llama 3.2 models
“So it's a series of three models, which are the best in class in each model size. For example, our 1.7 B model outperforms Lama one B and also .2.”
Loubna Ben Allal Dec 24, 2024 ▶ 22:31 Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
Dec 24, 2024 neutral
Assertion Supported
Ben Allal: LLaMA 3 used 15x more pre-training tokens than original LLaMA
“LAMA was trained on one trillion tokens, but LAMA-III was trained on 15 trillion tokens.”
Loubna Ben Allal Dec 24, 2024 ▶ 20:33 Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
Jan 19, 2025 negative
Assertion Not checkable as stated
Zhang: Meta Failed at Training MoE Models for Llama Series
“The reason why Lama open-sourced the MOE model, because I think they tried to train our MOE model, but they failed. So that, that's why they didn't open source MOE mode for Lama series.”
Yining Zhang Jan 19, 2025 ▶ 14:53 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Mar 23, 2025 positive
Insight
Agarwal: Synthetic data distillation bypasses model vocabulary and tokenizer mismatches
“The one nice thing about this kind of distillation is it doesn't matter if you have a vocabulary mismatch, because we're not using the next token distribution or probability labels. You can distill from one model which uses some random tokenizer to another mod…”
Rishabh Agarwal Mar 23, 2025 ▶ 14:26 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Jul 14, 2025 bearish
Assertion Supported
Meta Llama 3.3 and Llama 4 perform poorly on agent benchmarks
“Another, of course, the other surprise was that all the Lama models were not performing well on our benchmark. 3.3 and even the Lama four all were really performing extremely poor.”
Pratik Bhavsar Jul 14, 2025 ▶ 10:08 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
Jul 31, 2025 negative
Opinion
Lambert: Meta withholding its leading benchmark model is bad execution
“But to be a model that claims to be open and then not release the model that is your leading claim is just, like, that is, like, bad execution.”
Nathan Lambert Jul 31, 2025 ▶ 1:13:36 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Feb 5, 2026 positive
Disclosure
Goodfire AI: We replicated code error and malicious features in Llama
“We replicated a lot of these features in, in our llama models as well.”
Myra Deng Feb 5, 2026 ▶ 46:01 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Mar 8, 2026 positive
Assertion Supported
DeepSeek 128k context fits in 8GB KV cache versus Llama's 80GB
“For context like the total I think the total context length of DeepSeq is a 128,000 tokens, or it might be 256,000 with rope extension. That entire context, I think it's a 128,000, fits into eight gigabytes. And previously context, like I think the Lama four o…”
Kyle Kranen Mar 8, 2026 ▶ 52:08 Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
Mar 30, 2026 positive
What-if
Lample: Post-training breakthroughs like DPO were impossible without open LLaMA
“And if you look at many of the techniques that were developed after, for instance, Temma was open source, like all these post-training approaches like even DPOD, like performance optimization, all of this were done by people that had access to this model, and …”
Guillaume Lample Mar 30, 2026 ▶ 38:05 Mistral: Voxtral TTS, Forge, Leanstral, & Mistral 4 — w/ Pavan Kumar Reddy & Guillaume Lample
Jun 1, 2026 neutral
Disclosure
Ethan He: NVIDIA Cosmos uses a 7B video model with a larger LLM rewriter
“I think in in Cosmos, we use Lama or we use mix, mix through. And the Cosmos video model itself is only seven B, and the model, the language model is a prompt rewriter. It's bigger than that.”
Ethan He Jun 1, 2026 ▶ 1:15:37 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jul 8, 2026 neutral
Assertion Open · timeframe Jul 2026
Bubna: Ramp trained custom tokenizers to swap into LLaMA
“Ramp actually early in the day was training their own tokenizer and, like, Swapping out the tokenizer in Lama and whatnot.”
Akshat Bubna Jul 8, 2026 ▶ 51:11 The Future of AI Infra: from Kubernetes to Agent Sandboxes — Akshat Bubna, Modal CTO
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.