Llama 2

product on 14 shows · 22 statements across 13 episodes · said 192 times in 70 episodes since 2023

Latent Space 104 Big Technology 21 the MAD Podcast 15 All-In 13 20VC 13 the a16z Podcast 11 No Priors 5 BG2 Pod 2 We Live to Build 2 Acquired 2 In Depth 1 the Official SaaStr Podcast 1 How I Built This 1 A Product Market Fit Show 1

Mentions by year, every show

tap a year for its mentions
007520150402023202420252026episodesmentions
020402023202420252026episodes it came up in
002.5205402023202420252026episodesmentions per episode

Latent Space 104Big Technology 21the MAD Podcast 1520VC 13All-In 13the a16z Podcast 11No Priors 5Acquired 26 more shows

2026 5 mentions in 3 episodes 2 per episode
2025 13 mentions in 11 episodes 1 per episode
2024 102 mentions in 24 episodes 4 per episode
2023 72 mentions in 32 episodes 2 per episode

every mention on every show, scene by scene, with the transcript →

22 statements about Llama 2, every show

LATENT SPACE Assertion Supported
Bakouch: DeepSeek-V3 uses the same Adam optimizer parameters as Llama 2
“And for example, a good a good way to view that is that DeepSeq rig three is still using the same Adam parameter than Lama two.”
Elie Bakouch Oct 20, 2025 ▶ 8:54 ⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
LATENT SPACE Disclosure
Scialom: Llama 1 and 2 flagship size was chosen to reproduce Chinchilla
“Lama two, maybe I would say it's like Lama one. We had a flagship model, which was seven TB. It's also because the project was taking some routes to reproducing a chinchilla, which was a seven TB.”
Thomas Scialom Jul 23, 2024 ▶ 9:23 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
LATENT SPACE Assertion Supported
Scialom: Llama 3 Scaled Pre-Training to 15 Trillion Tokens
“It's the same recipe done in terms of architectures and training than LAMA-II, but we put so much effort on scaling the data and the quality of data. There's now 15 triant tokens compared to two triants, so it's another magnitude there as well, including for t…”
Thomas Scialom Jul 23, 2024 ▶ 18:15 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
LATENT SPACE Disclosure
Scialom: Meta Used Llama 2 to Filter and Tag Llama 3 Pre-Training Data
“LAMA was the best, at the time, before LAMA Free, the best model we had access to legally, to labelize the web and select what are the good tokens and the bad tokens. The additional thing is that it also enabled to have a topic tag, Like, is it about law? Is i…”
Thomas Scialom Jul 23, 2024 ▶ 19:16 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
LATENT SPACE Disclosure
Meta skipped coding, reasoning, and multilingual annotations for Llama 2
“And we didn't annotate at all for code, neither for reasoning or multinguity.”
Thomas Scialom Jul 23, 2024 ▶ 25:17 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
LATENT SPACE Disclosure
Meta's Llama 3 post-training uses almost entirely synthetic data
“So what we did is that we generated all the data on the prompts with LAMA-II, and we applied, like, basically the last round of LAMA-II we had to kick off and start LAMA-III post-training. So Lama-free post-training doesn't have any, like, human-written answer…”
Thomas Scialom Jul 23, 2024 ▶ 33:40 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
LATENT SPACE Disclosure
Scialom: Meta expanded Llama 3's vocabulary size to support multilingual capabilities
“Lama III compared to Lama II is multilingual, has multilingual capabilities. We worked on that. And so, because you have languages that are not just Latin languages like English, there's a lot of different characters you want to include them to represent, like…”
Thomas Scialom Jul 23, 2024 ▶ 55:24 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
LATENT SPACE Assertion Supported
Huang: Curriculum context expansion outperforms full-length training from scratch
“If you train a model on a shorter context and you progressively increase that context to, like, You know, the final limit that you have, like, 32 K is usually the limit of Lama two was that long. It actually performs better than if you try to train 32 K the …”
Mark Huang May 31, 2024 ▶ 15:39 How to train a Million Context LLM — with Mark Huang of Gradient.ai
LATENT SPACE Assertion Supported
Huang: Training CodeLlama on Llama 2 caused catastrophic language forgetting
“We do have historical precedent where CodeLlama was, you know, trained further from the original CodeLlama was trained further from Lama II, and it just lost, All its language capabilities, basically, right?”
Mark Huang May 31, 2024 ▶ 33:03 How to train a Million Context LLM — with Mark Huang of Gradient.ai
BIG TECHNOLOGY Assertion Partly supported
Meta used 100x more compute to train Llama 3 than Llama 2
“So actually, I think it's I believe it's a hundred times more compute.”
Meta's Generative AI Head Apr 22, 2024 ▶ 10:04 Meta's Generative AI Head: How We Trained Llama 3
BIG TECHNOLOGY Disclosure
Ahmad Al-Dahle admits Meta over-leveraged alignment tools in Llama 2
“In Lama two, it was we definitely I think over leveraged some of the alignment tools to discourage answering those kinds of questions.”
Meta's Generative AI Head Apr 22, 2024 ▶ 14:50 Meta's Generative AI Head: How We Trained Llama 3
HOW I BUILT THIS Assertion Supported
Harris: Safety controls on open-source AI models can be removed for $100
“I won't go into technical details, but basically for about a hundred dollars, you can retrain all the safety controls off of it.”
Tristan Harris Feb 29, 2024 ▶ 23:37 The peril (and promise) of AI with Tristan Harris: Part 2
LATENT SPACE Disclosure
Firshman: Llama 2 release was Replicate's biggest week of growth ever
“Llama II was, like, our biggest week of growth ever, because, like, tons of people wanted to tinker with it and run it.”
Ben Firshman Feb 28, 2024 ▶ 39:20 A Brief History of the Open Source AI Hacker - with Ben Firshman of Replicate
LATENT SPACE Assertion Partly supported
Firshman: Llama 2 costs $25M to train but $50 to fine-tune
“Lama II as a base model is that, like, yeah, it costs twenty five million dollars to train to start with, but then you can fine tune it for, like, 50 bucks.”
Ben Firshman Feb 28, 2024 ▶ 1:14:10 A Brief History of the Open Source AI Hacker - with Ben Firshman of Replicate
LATENT SPACE Assertion Not checkable as stated
Lambert: Meta spent roughly $6M to $8M on Llama 2 preference data
“So I would say, still say, like, six to eight million is safe to say that they're spending, if not more, they're probably also buying other types of data and or throwing out data that they don't like.”
Nathan Lambert Jan 11, 2024 ▶ 46:51 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
LATENT SPACE Assertion Supported
Lambert: Meta Used Rejection Sampling to Bootstrapping Llama 2 RLHF
“Llama started their RLHF process with this to get some signal out of preference data. That preference data went into a reward model, and then the reward model did a good enough ranking that it was, like, essentially superpowered instruction tuning based on rew…”
Nathan Lambert Jan 11, 2024 ▶ 1:03:01 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
ALL-IN Opinion
Chamath: Mistral and open-source models are superior to Llama 2
“Before I give all the credit to Facebook, I'd rather say that I think there are a lot of open source alternatives, including Mistral, that I think are much better. And free and open and in the clear where you can have unencumbered growth, not dependent on anyb…”
Chamath Palihapitiya Nov 3, 2023 ▶ 1:04:38 E152: Real estate chaos, WeWork bankruptcy, Biden regulates AI, Ukraine's “Cronkite Moment” & more
MAD Assertion Not checkable as stated
Fine-tuned Llama 2 achieves performance comparable to GPT-3.5 and GPT-4
“Straight out of the bat, if you just use Lama tool directly, I don't think you could get like comparable performance, you know, with GPD, 3.5 or four today. But like with fine tuning, if you make that investment in curating your data set in running that fine t…”
Shreya Rajpal Sep 27, 2023 ▶ 22:48 Guardrails AI: The Playbook for Safer, Hallucination-Free LLMs — Shreya Rajpal Explains
a16z Prediction Not checkable as stated
Scott: Meta's Llama 2 will be a key building block for AI software
“I think Lama too is just like Lama going to be an important building block that people are going to want to use to build AI software.”
Kevin Scott Sep 25, 2023 ▶ 5:28 AI Copilots and the Future of Knowledge Work with Microsoft's Kevin Scott
MAD Assertion Supported
Shah: Llama 2 instruction tuning datasets averaged one instruction per sample
“The average number of instructions on, there's, they have like seven data sets they showed for their instruction tuning. The average number of instructions? One.”
Munjal Shah Aug 23, 2023 ▶ 18:26 Hippocratic AI’s Munjal Shah: Building the First Safety-First LLM for Healthcare
WE LIVE TO BUILD Assertion Contradicted
Reyes: Meta's 7B parameter Llama 2 model achieves GPT-4 level behavior
“Facebook open-sourced their model this JAMA that is a competitor to chat GPT, and with seven billion parameters, it reaches very similar behavior than GPT-IV.”
Ricardo Michel Reyes Aug 15, 2023 ▶ 31:47 550 Million People, One Language, Zero Access to Venture Capital
Wood: Open-source AI models require substantial deployment infrastructure and tooling
“Lama II is a excellent, very capable model, but there is a long way to go from having the model weights which are what comprises the neural network to actually building out an artificial intelligence system. And just having the model weights is super useful, b…”
Matt Wood Aug 3, 2023 ▶ 2:16 Amazon Reveals Its AI Master Plan — With Matt Wood

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.