Llama 2

part of Llama includes Llama 2 70B, Llama 2 13B, Llama 2 7B

13 statements across 5 episodes · 8 bullish · 1 bearish · 5 people on the record · first statement Jan 11, 2024 by Nathan Lambert · said 113 times in 28 episodes since 2023 · across every show →

Mentions by year, the whole family

brought up most by Shawn Wang (17), Thomas Scialom (10), Nathan Lambert (8), Soumith Chintala (7), Dylan Patel (6), Ben Firshman (5), Jerry Liu (4), Alessio Fanelli (4)

tap a year for its mentions
0040880152023202420252026episodesmentions
08152023202420252026episodes it came up in
0047.58152023202420252026episodesmentions per episode
2026 4 mentions in 2 episodes 2 per episode
2025 7 mentions in 6 episodes 1 per episode
2024 80 mentions in 13 episodes 6 per episode
2023 22 mentions in 7 episodes 3 per episode

every mention, scene by scene, with the transcript →

Everything said about Llama 2, oldest first

Jan 11, 2024 positive
Assertion Supported
Lambert: Meta Used Rejection Sampling to Bootstrapping Llama 2 RLHF
“Llama started their RLHF process with this to get some signal out of preference data. That preference data went into a reward model, and then the reward model did a good enough ranking that it was, like, essentially superpowered instruction tuning based on rew…”
Nathan Lambert Jan 11, 2024 ▶ 1:03:01 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 neutral
Assertion Not checkable as stated
Lambert: Meta spent roughly $6M to $8M on Llama 2 preference data
“So I would say, still say, like, six to eight million is safe to say that they're spending, if not more, they're probably also buying other types of data and or throwing out data that they don't like.”
Nathan Lambert Jan 11, 2024 ▶ 46:51 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Feb 28, 2024 positive
Assertion Partly supported
Firshman: Llama 2 costs $25M to train but $50 to fine-tune
“Lama II as a base model is that, like, yeah, it costs twenty five million dollars to train to start with, but then you can fine tune it for, like, 50 bucks.”
Ben Firshman Feb 28, 2024 ▶ 1:14:10 A Brief History of the Open Source AI Hacker - with Ben Firshman of Replicate
Feb 28, 2024 bullish
Disclosure
Firshman: Llama 2 release was Replicate's biggest week of growth ever
“Llama II was, like, our biggest week of growth ever, because, like, tons of people wanted to tinker with it and run it.”
Ben Firshman Feb 28, 2024 ▶ 39:20 A Brief History of the Open Source AI Hacker - with Ben Firshman of Replicate
May 31, 2024 negative
Assertion Supported
Huang: Training CodeLlama on Llama 2 caused catastrophic language forgetting
“We do have historical precedent where CodeLlama was, you know, trained further from the original CodeLlama was trained further from Lama II, and it just lost, All its language capabilities, basically, right?”
Mark Huang May 31, 2024 ▶ 33:03 How to train a Million Context LLM — with Mark Huang of Gradient.ai
May 31, 2024 positive
Assertion Supported
Huang: Curriculum context expansion outperforms full-length training from scratch
“If you train a model on a shorter context and you progressively increase that context to, like, You know, the final limit that you have, like, 32 K is usually the limit of Lama two was that long. It actually performs better than if you try to train 32 K the …”
Mark Huang May 31, 2024 ▶ 15:39 How to train a Million Context LLM — with Mark Huang of Gradient.ai
Jul 23, 2024 bullish
Assertion Supported
Scialom: Llama 3 Scaled Pre-Training to 15 Trillion Tokens
“It's the same recipe done in terms of architectures and training than LAMA-II, but we put so much effort on scaling the data and the quality of data. There's now 15 triant tokens compared to two triants, so it's another magnitude there as well, including for t…”
Thomas Scialom Jul 23, 2024 ▶ 18:15 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Jul 23, 2024 neutral
Disclosure
Scialom: Llama 1 and 2 flagship size was chosen to reproduce Chinchilla
“Lama two, maybe I would say it's like Lama one. We had a flagship model, which was seven TB. It's also because the project was taking some routes to reproducing a chinchilla, which was a seven TB.”
Thomas Scialom Jul 23, 2024 ▶ 9:23 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Jul 23, 2024 positive
Disclosure
Meta's Llama 3 post-training uses almost entirely synthetic data
“So what we did is that we generated all the data on the prompts with LAMA-II, and we applied, like, basically the last round of LAMA-II we had to kick off and start LAMA-III post-training. So Lama-free post-training doesn't have any, like, human-written answer…”
Thomas Scialom Jul 23, 2024 ▶ 33:40 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Jul 23, 2024 positive
Disclosure
Scialom: Meta Used Llama 2 to Filter and Tag Llama 3 Pre-Training Data
“LAMA was the best, at the time, before LAMA Free, the best model we had access to legally, to labelize the web and select what are the good tokens and the bad tokens. The additional thing is that it also enabled to have a topic tag, Like, is it about law? Is i…”
Thomas Scialom Jul 23, 2024 ▶ 19:16 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Jul 23, 2024 positive
Disclosure
Scialom: Meta expanded Llama 3's vocabulary size to support multilingual capabilities
“Lama III compared to Lama II is multilingual, has multilingual capabilities. We worked on that. And so, because you have languages that are not just Latin languages like English, there's a lot of different characters you want to include them to represent, like…”
Thomas Scialom Jul 23, 2024 ▶ 55:24 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Jul 23, 2024
Disclosure
Meta skipped coding, reasoning, and multilingual annotations for Llama 2
“And we didn't annotate at all for code, neither for reasoning or multinguity.”
Thomas Scialom Jul 23, 2024 ▶ 25:17 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Oct 20, 2025 neutral
Assertion Supported
Bakouch: DeepSeek-V3 uses the same Adam optimizer parameters as Llama 2
“And for example, a good a good way to view that is that DeepSeq rig three is still using the same Adam parameter than Lama two.”
Elie Bakouch Oct 20, 2025 ▶ 8:54 ⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.