Llama 3

also referred to as: llama-3

part of Llama includes Llama 3 405B, Llama 3 70B, Llama 3 8B, Llama 3 Tokenizer

12 statements across 5 episodes · 9 bullish · 0 bearish · 5 people on the record · first statement May 31, 2024 by Mark Huang · said 118 times in 32 episodes since 2023 · across every show →

Mentions by year, the whole family

brought up most by Shawn Wang (24), Alessio Fanelli (15), Thomas Scialom (8), Soumith Chintala (4), Philip Kiely (3), Mark Huang (3), Loubna Ben Allal (3), Yi Tay (2)

tap a year for its mentions
007510150202023202420252026episodesmentions
010202023202420252026episodes it came up in
003106202023202420252026episodesmentions per episode
2026 7 mentions in 4 episodes 2 per episode
2025 9 mentions in 8 episodes 1 per episode
2024 101 mentions in 19 episodes 5 per episode
2023 1 mention in 1 episode

every mention, scene by scene, with the transcript →

Everything said about Llama 3, oldest first

May 31, 2024 positive
Disclosure
Gradient trained 1M context Llama-3 on Crusoe's Nvidia L40 GPUs
“It just made it really easy for us to, like, scale up with their L-Forties, and those are the specific GPU instances we used, and coordinating that effort with them to get, you know, that dedicated cluster first to do the project”
Mark Huang May 31, 2024 ▶ 13:47 How to train a Million Context LLM — with Mark Huang of Gradient.ai
Jul 23, 2024 bullish
Assertion Supported
Scialom: Llama 3 Scaled Pre-Training to 15 Trillion Tokens
“It's the same recipe done in terms of architectures and training than LAMA-II, but we put so much effort on scaling the data and the quality of data. There's now 15 triant tokens compared to two triants, so it's another magnitude there as well, including for t…”
Thomas Scialom Jul 23, 2024 ▶ 18:15 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Jul 23, 2024
Disclosure
Meta changed Llama 3's pre-training data mixture mid-training run
“What happened is we changed the data mix during the training of Lama three with some findings that happened in the... Training is long, so you have to do something while it's training. And what the team did, I was working on my side of motion post-training, bu…”
Thomas Scialom Jul 23, 2024 ▶ 23:12 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Jul 23, 2024 positive
Assertion Supported
Scialom: Llama 3 405B is the best open-source model ever released
“At a high level, it's the best open source model ever. It's Better than GPT-IV. I mean, what version? But, by far, compared to the version originally released even now, I think there's maybe the last cloud Sonya FF-V and GPT-IV-Zero that are performing it.”
Thomas Scialom Jul 23, 2024 ▶ 37:43 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Jul 23, 2024 positive
Disclosure
Meta's Llama 3 post-training uses almost entirely synthetic data
“So what we did is that we generated all the data on the prompts with LAMA-II, and we applied, like, basically the last round of LAMA-II we had to kick off and start LAMA-III post-training. So Lama-free post-training doesn't have any, like, human-written answer…”
Thomas Scialom Jul 23, 2024 ▶ 33:40 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Jul 23, 2024 positive
Disclosure
Scialom: Meta Used Llama 2 to Filter and Tag Llama 3 Pre-Training Data
“LAMA was the best, at the time, before LAMA Free, the best model we had access to legally, to labelize the web and select what are the good tokens and the bad tokens. The additional thing is that it also enabled to have a topic tag, Like, is it about law? Is i…”
Thomas Scialom Jul 23, 2024 ▶ 19:16 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Jul 23, 2024 positive
Assertion Supported
Scialom: Llama 3 natively supports state-of-the-art tool calling
“We have that from day one. Good news for the community. We are state of the art there. I think the model will be pretty good at that, we have a lot of gems about tools in the paper, but the model is fine-tuned to do tool usage, to zero-shot function calling”
Thomas Scialom Jul 23, 2024 ▶ 45:05 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Jul 23, 2024
Disclosure
Scialom: Multimodal Llama 3 will add parameters beyond 405B
“For the text text model only? Yes. A bit of additional parameters for the multimodal version that we come later.”
Thomas Scialom Jul 23, 2024 ▶ 0:36 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Jul 23, 2024 positive
Disclosure
Scialom: Meta expanded Llama 3's vocabulary size to support multilingual capabilities
“Lama III compared to Lama II is multilingual, has multilingual capabilities. We worked on that. And so, because you have languages that are not just Latin languages like English, there's a lot of different characters you want to include them to represent, like…”
Thomas Scialom Jul 23, 2024 ▶ 55:24 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Aug 22, 2024 positive
Opinion
Wang: Natural language-to-code translation is ripe for synthetic data generation
“I think that translation between natural language, English versus code and back and forth, I think is actually actually a really ripe source of synthetic data and Lama three specifically called out that, that they trained on that.”
Shawn Wang Aug 22, 2024 ▶ 48:12 Is finetuning GPT4o worth it?
Dec 23, 2024 bullish
Assertion Supported
Soldani: 2024 open models rival frontier performance of closed models
“You have models that are, you know, reveling frontier level performance of what you can get from closed models, from like Quen, from Deep Seek. We got Lama III.”
Luca Soldani Dec 23, 2024 ▶ 1:15 Best of 2024: Open Models [LS LIVE! at NeurIPS 2024]
Dec 24, 2024 neutral
Assertion Supported
Ben Allal: LLaMA 3 used 15x more pre-training tokens than original LLaMA
“LAMA was trained on one trillion tokens, but LAMA-III was trained on 15 trillion tokens.”
Loubna Ben Allal Dec 24, 2024 ▶ 20:33 Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.