Apr 22, 2024 · 30m · big-technology

Meta's Generative AI Head: How We Trained Llama 3

Meta's Generative AI Head · 19m spoken Alex Kantrowitz · 8m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Meta's Generative AI Head sits down with Alex Kantrowitz to break down the engineering, training infrastructure, and synthetic data innovations behind Llama 3, as well as Meta's strategy for integrating frontier models directly into consumer products while maintaining open-source safety standards.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Alex holds 30.9% of the talking time here. How this is scored →

Alex as informed peer 4.4 Guest teaching 3.0 Guest disagreement 1.5 Alex pushing back 2.1
05100:0010:0020:0030:000:00–2:03 · Alex as informed peer 3/10 Announcing Llama 3 and Meta AI Integration Kantrowitz opens with a broad question about the Llama 3 launch. The guest provides a high-level overview of releasing the 8B and 70B models and integrating with Meta AI without any friction.2:03–4:28 · Alex as informed peer 2/10 Explaining Model Parameters, Weights, and Next-Token Prediction Kantrowitz asks fundamental clarifying questions about parameters and weights. The guest delivers an educational breakdown of matrix multiplication, token prediction, and knowledge encoding.4:28–7:52 · Alex as informed peer 6/10 Gradient Descent Optimization and GPU Cluster Infrastructure The guest explains gradient descent after asking Kantrowitz to clarify his prompt. Kantrowitz demonstrates solid research by quoting Meta's cluster publications and computing the financial scale of 24,000 GPUs.7:52–10:15 · Alex as informed peer 4/10 Scaling Architecture from Llama 2 to Llama 3 Kantrowitz asks about the scaling leap between generations. The guest explains the architecture scaling and gently corrects Kantrowitz's summary by clarifying that compute increased by 100x rather than 10x.10:15–12:42 · Alex as informed peer 6/10 Navigating Data Scarcity and Synthetic Data Generation Kantrowitz presses hard on data constraints by citing a New York Times report detailing internal discussions on buying publishing houses. The guest pushes back against the premise of a hard data wall by highlighting synthetic data breakthroughs.12:42–16:11 · Alex as informed peer 4/10 Refining Alignment, Boundary Sampling, and Refusal Tone Kantrowitz asks how Meta made the model more of a 'cowboy' to avoid false refusals. The guest rejects the cowboy framing and educates on boundary sampling techniques using Linux command examples.16:11–20:36 · Alex as informed peer 4/10 Meta AI User Experience and Real-Time Image Generation The exchange shifts to product UX and real-time generation speed. Kantrowitz relates practical user experience anecdotes from image generation workflows.20:36–24:07 · Alex as informed peer 5/10 Operationalizing Frontier Models Through Cross-Team Collaboration and Red Teaming Kantrowitz brings up red teaming and references Google's Gemini issues, then questions Meta AI's discoverability in messaging apps. The guest details cross-functional launch orchestration.24:07–27:15 · Alex as informed peer 6/10 Open Source Philosophy, Model Safety, and Guardrails Kantrowitz challenges the guest on Meta's commitment to open-sourcing the upcoming 400B model. The guest calls out the host's attempt to force a binary answer while the model is still training, prompting Kantrowitz to push back on cybersecurity risks.27:15–29:54 · Alex as informed peer 4/10 Evaluating Scaling Laws and Resisting Frontier Speculation Kantrowitz asks if anything surprised the team during training. The guest explains scaling laws and firmly rejects speculative timeline predictions by drawing parallels to autonomous driving hype.0:00–2:03 · Guest teaching 1/10 Announcing Llama 3 and Meta AI Integration Kantrowitz opens with a broad question about the Llama 3 launch. The guest provides a high-level overview of releasing the 8B and 70B models and integrating with Meta AI without any friction.2:03–4:28 · Guest teaching 5/10 Explaining Model Parameters, Weights, and Next-Token Prediction Kantrowitz asks fundamental clarifying questions about parameters and weights. The guest delivers an educational breakdown of matrix multiplication, token prediction, and knowledge encoding.4:28–7:52 · Guest teaching 4/10 Gradient Descent Optimization and GPU Cluster Infrastructure The guest explains gradient descent after asking Kantrowitz to clarify his prompt. Kantrowitz demonstrates solid research by quoting Meta's cluster publications and computing the financial scale of 24,000 GPUs.7:52–10:15 · Guest teaching 4/10 Scaling Architecture from Llama 2 to Llama 3 Kantrowitz asks about the scaling leap between generations. The guest explains the architecture scaling and gently corrects Kantrowitz's summary by clarifying that compute increased by 100x rather than 10x.10:15–12:42 · Guest teaching 3/10 Navigating Data Scarcity and Synthetic Data Generation Kantrowitz presses hard on data constraints by citing a New York Times report detailing internal discussions on buying publishing houses. The guest pushes back against the premise of a hard data wall by highlighting synthetic data breakthroughs.12:42–16:11 · Guest teaching 4/10 Refining Alignment, Boundary Sampling, and Refusal Tone Kantrowitz asks how Meta made the model more of a 'cowboy' to avoid false refusals. The guest rejects the cowboy framing and educates on boundary sampling techniques using Linux command examples.16:11–20:36 · Guest teaching 1/10 Meta AI User Experience and Real-Time Image Generation The exchange shifts to product UX and real-time generation speed. Kantrowitz relates practical user experience anecdotes from image generation workflows.20:36–24:07 · Guest teaching 2/10 Operationalizing Frontier Models Through Cross-Team Collaboration and Red Teaming Kantrowitz brings up red teaming and references Google's Gemini issues, then questions Meta AI's discoverability in messaging apps. The guest details cross-functional launch orchestration.24:07–27:15 · Guest teaching 3/10 Open Source Philosophy, Model Safety, and Guardrails Kantrowitz challenges the guest on Meta's commitment to open-sourcing the upcoming 400B model. The guest calls out the host's attempt to force a binary answer while the model is still training, prompting Kantrowitz to push back on cybersecurity risks.27:15–29:54 · Guest teaching 3/10 Evaluating Scaling Laws and Resisting Frontier Speculation Kantrowitz asks if anything surprised the team during training. The guest explains scaling laws and firmly rejects speculative timeline predictions by drawing parallels to autonomous driving hype.0:00–2:03 · Guest disagreement 0/10 Announcing Llama 3 and Meta AI Integration Kantrowitz opens with a broad question about the Llama 3 launch. The guest provides a high-level overview of releasing the 8B and 70B models and integrating with Meta AI without any friction.2:03–4:28 · Guest disagreement 0/10 Explaining Model Parameters, Weights, and Next-Token Prediction Kantrowitz asks fundamental clarifying questions about parameters and weights. The guest delivers an educational breakdown of matrix multiplication, token prediction, and knowledge encoding.4:28–7:52 · Guest disagreement 1/10 Gradient Descent Optimization and GPU Cluster Infrastructure The guest explains gradient descent after asking Kantrowitz to clarify his prompt. Kantrowitz demonstrates solid research by quoting Meta's cluster publications and computing the financial scale of 24,000 GPUs.7:52–10:15 · Guest disagreement 1/10 Scaling Architecture from Llama 2 to Llama 3 Kantrowitz asks about the scaling leap between generations. The guest explains the architecture scaling and gently corrects Kantrowitz's summary by clarifying that compute increased by 100x rather than 10x.10:15–12:42 · Guest disagreement 2/10 Navigating Data Scarcity and Synthetic Data Generation Kantrowitz presses hard on data constraints by citing a New York Times report detailing internal discussions on buying publishing houses. The guest pushes back against the premise of a hard data wall by highlighting synthetic data breakthroughs.12:42–16:11 · Guest disagreement 2/10 Refining Alignment, Boundary Sampling, and Refusal Tone Kantrowitz asks how Meta made the model more of a 'cowboy' to avoid false refusals. The guest rejects the cowboy framing and educates on boundary sampling techniques using Linux command examples.16:11–20:36 · Guest disagreement 0/10 Meta AI User Experience and Real-Time Image Generation The exchange shifts to product UX and real-time generation speed. Kantrowitz relates practical user experience anecdotes from image generation workflows.20:36–24:07 · Guest disagreement 0/10 Operationalizing Frontier Models Through Cross-Team Collaboration and Red Teaming Kantrowitz brings up red teaming and references Google's Gemini issues, then questions Meta AI's discoverability in messaging apps. The guest details cross-functional launch orchestration.24:07–27:15 · Guest disagreement 6/10 Open Source Philosophy, Model Safety, and Guardrails Kantrowitz challenges the guest on Meta's commitment to open-sourcing the upcoming 400B model. The guest calls out the host's attempt to force a binary answer while the model is still training, prompting Kantrowitz to push back on cybersecurity risks.27:15–29:54 · Guest disagreement 3/10 Evaluating Scaling Laws and Resisting Frontier Speculation Kantrowitz asks if anything surprised the team during training. The guest explains scaling laws and firmly rejects speculative timeline predictions by drawing parallels to autonomous driving hype.0:00–2:03 · Alex pushing back 0/10 Announcing Llama 3 and Meta AI Integration Kantrowitz opens with a broad question about the Llama 3 launch. The guest provides a high-level overview of releasing the 8B and 70B models and integrating with Meta AI without any friction.2:03–4:28 · Alex pushing back 0/10 Explaining Model Parameters, Weights, and Next-Token Prediction Kantrowitz asks fundamental clarifying questions about parameters and weights. The guest delivers an educational breakdown of matrix multiplication, token prediction, and knowledge encoding.4:28–7:52 · Alex pushing back 1/10 Gradient Descent Optimization and GPU Cluster Infrastructure The guest explains gradient descent after asking Kantrowitz to clarify his prompt. Kantrowitz demonstrates solid research by quoting Meta's cluster publications and computing the financial scale of 24,000 GPUs.7:52–10:15 · Alex pushing back 2/10 Scaling Architecture from Llama 2 to Llama 3 Kantrowitz asks about the scaling leap between generations. The guest explains the architecture scaling and gently corrects Kantrowitz's summary by clarifying that compute increased by 100x rather than 10x.10:15–12:42 · Alex pushing back 5/10 Navigating Data Scarcity and Synthetic Data Generation Kantrowitz presses hard on data constraints by citing a New York Times report detailing internal discussions on buying publishing houses. The guest pushes back against the premise of a hard data wall by highlighting synthetic data breakthroughs.12:42–16:11 · Alex pushing back 2/10 Refining Alignment, Boundary Sampling, and Refusal Tone Kantrowitz asks how Meta made the model more of a 'cowboy' to avoid false refusals. The guest rejects the cowboy framing and educates on boundary sampling techniques using Linux command examples.16:11–20:36 · Alex pushing back 0/10 Meta AI User Experience and Real-Time Image Generation The exchange shifts to product UX and real-time generation speed. Kantrowitz relates practical user experience anecdotes from image generation workflows.20:36–24:07 · Alex pushing back 3/10 Operationalizing Frontier Models Through Cross-Team Collaboration and Red Teaming Kantrowitz brings up red teaming and references Google's Gemini issues, then questions Meta AI's discoverability in messaging apps. The guest details cross-functional launch orchestration.24:07–27:15 · Alex pushing back 6/10 Open Source Philosophy, Model Safety, and Guardrails Kantrowitz challenges the guest on Meta's commitment to open-sourcing the upcoming 400B model. The guest calls out the host's attempt to force a binary answer while the model is still training, prompting Kantrowitz to push back on cybersecurity risks.27:15–29:54 · Alex pushing back 2/10 Evaluating Scaling Laws and Resisting Frontier Speculation Kantrowitz asks if anything surprised the team during training. The guest explains scaling laws and firmly rejects speculative timeline predictions by drawing parallels to autonomous driving hype.

speaking balance: gold is Alex, purple is the guest (3 minute bins)

0:00 · Alex 9% · guest 91%0:00 · Alex 9% · guest 91%3:00 · Alex 11.1% · guest 88.9%3:00 · Alex 11.1% · guest 88.9%6:00 · Alex 31.6% · guest 68.4%6:00 · Alex 31.6% · guest 68.4%9:00 · Alex 41.4% · guest 58.6%9:00 · Alex 41.4% · guest 58.6%12:00 · Alex 14% · guest 86%12:00 · Alex 14% · guest 86%15:00 · Alex 36.4% · guest 63.6%15:00 · Alex 36.4% · guest 63.6%18:00 · Alex 59.8% · guest 40.2%18:00 · Alex 59.8% · guest 40.2%21:00 · Alex 18.9% · guest 81.1%21:00 · Alex 18.9% · guest 81.1%24:00 · Alex 42.9% · guest 57.1%24:00 · Alex 42.9% · guest 57.1%27:00 · Alex 29.2% · guest 70.8%27:00 · Alex 29.2% · guest 70.8%30:00 · Alex 100% · guest 0%30:00 · Alex 100% · guest 0%
Sharpest disagreement ▶ 26:05 Guest directly rebuffs forced binary framing

The guest pushes back firmly against Kantrowitz's interpretation, telling him directly that he is trying to force a yes-or-no answer when the model is still in training.

Hardest push from Alex ▶ 25:57 Host challenges noncommittal open source response

Kantrowitz refuses to let the guest's ambiguous answer slide regarding whether Meta will open source its massive 400B model, pointing out the contrast with earlier releases.

Biggest teaching moment ▶ 10:04 Guest corrects host on compute scaling order of magnitude

The guest directly corrects Kantrowitz's 10x compute figure, clarifying that Llama 3 required a 100x increase in compute resources compared to Llama 2.

Alex holds their own ▶ 10:15 Host confronts guest with leaked NYT reporting

Kantrowitz quotes specific reporting from the New York Times regarding internal meetings, Simon & Schuster acquisition talks, and data ceiling limits.

the scores for every segment, with the reasoning behind each
ChapterTopicAlex as informed peerGuest teachingGuest disagreementAlex pushing backWhy
Announcing Llama 3 and Meta AI Integration 3100 Kantrowitz opens with a broad question about the Llama 3 launch. The guest provides a high-level overview of releasing the 8B and 70B models and integrating with Meta AI without any friction.
Explaining Model Parameters, Weights, and Next-Token Prediction 2500 Kantrowitz asks fundamental clarifying questions about parameters and weights. The guest delivers an educational breakdown of matrix multiplication, token prediction, and knowledge encoding.
Gradient Descent Optimization and GPU Cluster Infrastructure 6411 The guest explains gradient descent after asking Kantrowitz to clarify his prompt. Kantrowitz demonstrates solid research by quoting Meta's cluster publications and computing the financial scale of 24,000 GPUs.
Scaling Architecture from Llama 2 to Llama 3 4412 Kantrowitz asks about the scaling leap between generations. The guest explains the architecture scaling and gently corrects Kantrowitz's summary by clarifying that compute increased by 100x rather than 10x.
Navigating Data Scarcity and Synthetic Data Generation 6325 Kantrowitz presses hard on data constraints by citing a New York Times report detailing internal discussions on buying publishing houses. The guest pushes back against the premise of a hard data wall by highlighting synthetic data breakthroughs.
Refining Alignment, Boundary Sampling, and Refusal Tone 4422 Kantrowitz asks how Meta made the model more of a 'cowboy' to avoid false refusals. The guest rejects the cowboy framing and educates on boundary sampling techniques using Linux command examples.
Meta AI User Experience and Real-Time Image Generation 4100 The exchange shifts to product UX and real-time generation speed. Kantrowitz relates practical user experience anecdotes from image generation workflows.
Operationalizing Frontier Models Through Cross-Team Collaboration and Red Teaming 5203 Kantrowitz brings up red teaming and references Google's Gemini issues, then questions Meta AI's discoverability in messaging apps. The guest details cross-functional launch orchestration.
Open Source Philosophy, Model Safety, and Guardrails 6366 Kantrowitz challenges the guest on Meta's commitment to open-sourcing the upcoming 400B model. The guest calls out the host's attempt to force a binary answer while the model is still training, prompting Kantrowitz to push back on cybersecurity risks.
Evaluating Scaling Laws and Resisting Frontier Speculation 4332 Kantrowitz asks if anything surprised the team during training. The guest explains scaling laws and firmly rejects speculative timeline predictions by drawing parallels to autonomous driving hype.

Statements from this episode (16)

Disclosure
Meta releases 8-billion and 70-billion parameter Llama 3 models
“We are releasing an updated eight billion parameter model plus a seventy billion parameter model. And these are state of the art.”
Meta's Generative AI Head Apr 22, 2024 ▶ 0:09
Prediction Not checkable as stated
Ahmad Al-Dahle says Meta AI will be the best free assistant
“And meta AI is going to be the, you know one of the best if not the best assistant that's available for free.”
Meta's Generative AI Head Apr 22, 2024 ▶ 1:42
Insight
Ahmad Al-Dahle: 70B parameters is the magical number for billion-user scale
“The seventy billion is a magical number that allows us to scale to billions of people. And so we're really excited that we've been able to strike the right balance between intelligence and efficiency.”
Meta's Generative AI Head Apr 22, 2024 ▶ 1:54
Assertion Supported
Ahmad Al-Dahle: 8-billion parameter models can run locally on phones
“So an eight billion parameter model is small enough to sort of run, On phones or even laptops at the higher end of those.”
Meta's Generative AI Head Apr 22, 2024 ▶ 2:32
Assertion Supported
Meta trained Llama 3 8B and 70B models on 15 trillion tokens
“So if you look at something like the eight billion and seventy billion, They were trained on almost 15 trillion tokens and tokens roughly you can imagine as a word. So roughly like 15 trillion words, which is an incredible outcome.”
Meta's Generative AI Head Apr 22, 2024 ▶ 5:43
Disclosure
Ahmad Al-Dahle: Meta is currently training a 400-billion parameter Llama 3 model
“We also are talking a little bit about one of the larger models that we're training that is already achieving, you know exceptional performance which is a, it's a model that's over four hundred billion parameters.”
Meta's Generative AI Head Apr 22, 2024 ▶ 9:07
Assertion Partly supported
Meta used 100x more compute to train Llama 3 than Llama 2
“So actually, I think it's I believe it's a hundred times more compute.”
Meta's Generative AI Head Apr 22, 2024 ▶ 10:04
Disclosure
Meta leveraged synthetic data during Llama 3 post-training phase
“One of the things that we did with Lama three is in post-training we actually leveraged synthetic data.”
Meta's Generative AI Head Apr 22, 2024 ▶ 11:58
Opinion
Ahmad Al-Dahle doubts the AI industry will hit a data wall soon
“I don't think we know yet that, you know, we'll run out of data any, or that there's some like limiting factor here.”
Meta's Generative AI Head Apr 22, 2024 ▶ 12:29
Disclosure
Ahmad Al-Dahle admits Meta over-leveraged alignment tools in Llama 2
“In Lama two, it was we definitely I think over leveraged some of the alignment tools to discourage answering those kinds of questions.”
Meta's Generative AI Head Apr 22, 2024 ▶ 14:50
Insight
Ahmad Al-Dahle: AI models degrade user experience by over-moralizing refusals
“Some of these models, for example tend to do a lot of moralization or like really take a perspective or a point of view. And we worked on and I'm continuing to work on and innovate on how, how the model responds and how it refuses, which I think is also part o…”
Meta's Generative AI Head Apr 22, 2024 ▶ 15:22
Assertion Supported
Meta's new real-time image generation model produces outputs under one second
“And that's why we did all this model work to really get this thing very fast. It's under a second, which is like really, really exciting.”
Meta's Generative AI Head Apr 22, 2024 ▶ 19:58
Disclosure
Meta is integrating Meta AI into search suggestions and inboxes
“We have high level entry points directly in the inbox. We're integrating it into search. So we'll have also suggestions and type of heads as you sort of apply it in search. So we're, we've integrated it at a very prominent, in a very prominent way.”
Meta's Generative AI Head Apr 22, 2024 ▶ 23:31
Assertion Not checkable as stated
Ahmad Al-Dahle says every AI lab depends on open research
“I would say like every AI lab in the world today kind of has depended on openness and transparency in order to achieve the outcomes and the results and the improvements to these models that we have today.”
Meta's Generative AI Head Apr 22, 2024 ▶ 25:18
Assertion Supported
Meta releases safety model Llama Guard and open-sources cybersecurity evaluations
“So if you look at for example, in our release with Lama three, we were open, we're opening a model called Lama guard. And we've also been opening up cybersecurity evals to help understand this, the safety metrics for cybersecurity.”
Meta's Generative AI Head Apr 22, 2024 ▶ 26:49
Assertion Not checkable as stated
Ahmad Al-Dahle: Llama 3's performance exactly matched Meta's scaling law predictions
“I don't think anything about the model has really personally surprised me in terms of its performance. I think we kind of expected to be here. You know, we do a lot of like rigorous scaling laws and rigorous prediction of what we think the metrics will look li…”
Meta's Generative AI Head Apr 22, 2024 ▶ 27:55
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.