Jul 23, 2024 · 1h 4m · latent-space

Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI

Thomas Scialom · 39m spoken Shawn Wang · 10m spoken Alessio Fanelli · 7m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the Latent Space podcast, Meta AI technical lead Thomas Scialom discusses the engineering breakthroughs, scaling principles, and post-training innovations behind the Llama 2 and Llama 3 model families. He provides deep technical insights into synthetic data generation, RLHF mechanics, architectural trade-offs, and Meta's future roadmap toward autonomous agentic intelligence.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 13% of the talking time here. How this is scored →

The hosts as informed peer 6.3 Guest teaching 6.2 Guest disagreement 2.0 The hosts pushing back 2.1
05100:0015:0030:0045:001:00:000:45–4:11 · The hosts as informed peer 6/10 Career Path: From Quant Trading to NLP and Meta Shawn demonstrates strong background knowledge of Thomas's publishing history, recalling specific papers on summarization, factual consistency, and language GANs. Thomas explains his transition from quant trading to NLP right before BERT, with both sides sharing a collegial peer dynamic.4:11–8:03 · The hosts as informed peer 5/10 The Origins of Llama 2: From Galactica to Scale Thomas schools the hosts on the internal history of Galactica, the unexpected public backlash, and how Galactica Instruct pioneered RLHF annotation at Meta before Llama 2. Shawn and Alessio ask probing questions regarding annotation scale and open research questions.8:05–12:17 · The hosts as informed peer 7/10 LLM Scaling Laws and Escaping the Chinchilla Trap Alessio presses Thomas on parameter scaling laws and his tweet warning researchers not to fall into the Chinchilla trap. Thomas provides an in-depth breakdown of inference-optimal training versus compute-optimal training for published benchmarks.12:17–17:36 · The hosts as informed peer 6/10 Strategic Vision for Llama 3 405B and Quantization Alessio challenges the accessibility of training a massive 405B dense model for community inference. Thomas justifies the decision by explaining how FP8 quantization enables single-node execution and how 405B acts as a teacher for smaller models.17:36–22:40 · The hosts as informed peer 7/10 Pre-Training Data Quality and Synthetic Filtering in Llama 3 Shawn and Alessio raise technical points about synthetic data filtering and compare pre-training synthetic data to computer vision augmentation. Thomas enthusiastically agrees while drawing a clear line between pre-training curation and post-training augmentation.22:40–27:43 · The hosts as informed peer 5/10 Post-Training Strategy, Expert Domains, and MoE Considerations Thomas explains Meta's continuous pre-training strategy across expert domains and clarifies why dense models were chosen over MoE for this generation. Alessio and Shawn ask targeted follow-ups about curriculum learning and synthetic models.27:46–37:20 · The hosts as informed peer 7/10 Demystifying RLHF: Human Discrimination vs Generation Alessio references Nathan Lambert's work and asks detailed questions about RLHF versus SFT impact. Thomas delivers a masterclass on human discrimination versus generation abilities, explaining how RLHF allows models to surpass human performance without human-written SFT data.37:21–45:38 · The hosts as informed peer 7/10 Evaluating Llama 3: Benchmarks, Calibration, and Tool Calling Alessio and Shawn cite specific evaluation numbers and discuss calibration, uncertainty estimation, and tool use. Thomas agrees on calibration deficiencies in post-trained models and proposes calibration-specific prompts.45:38–54:36 · The hosts as informed peer 7/10 Paving the Way for Llama 4: Agents and Latent Reasoning Shawn brings up Yann LeCun's JEPA architecture and Anthropic's hidden thinking prompt techniques. Thomas outlines the path toward Llama 4 and autonomous agents, discussing latent space reasoning and variable compute per token.54:45–59:15 · The hosts as informed peer 6/10 Tokenizer Architecture and Multilingual Vocabulary Scaling Alessio queries the mechanical impact of scaling vocabulary size to 128K, prompting Thomas to explain token compression, compute trade-offs for small versus large models, and potential pixel-level or character-level tokenizers.0:45–4:11 · Guest teaching 3/10 Career Path: From Quant Trading to NLP and Meta Shawn demonstrates strong background knowledge of Thomas's publishing history, recalling specific papers on summarization, factual consistency, and language GANs. Thomas explains his transition from quant trading to NLP right before BERT, with both sides sharing a collegial peer dynamic.4:11–8:03 · Guest teaching 6/10 The Origins of Llama 2: From Galactica to Scale Thomas schools the hosts on the internal history of Galactica, the unexpected public backlash, and how Galactica Instruct pioneered RLHF annotation at Meta before Llama 2. Shawn and Alessio ask probing questions regarding annotation scale and open research questions.8:05–12:17 · Guest teaching 7/10 LLM Scaling Laws and Escaping the Chinchilla Trap Alessio presses Thomas on parameter scaling laws and his tweet warning researchers not to fall into the Chinchilla trap. Thomas provides an in-depth breakdown of inference-optimal training versus compute-optimal training for published benchmarks.12:17–17:36 · Guest teaching 5/10 Strategic Vision for Llama 3 405B and Quantization Alessio challenges the accessibility of training a massive 405B dense model for community inference. Thomas justifies the decision by explaining how FP8 quantization enables single-node execution and how 405B acts as a teacher for smaller models.17:36–22:40 · Guest teaching 6/10 Pre-Training Data Quality and Synthetic Filtering in Llama 3 Shawn and Alessio raise technical points about synthetic data filtering and compare pre-training synthetic data to computer vision augmentation. Thomas enthusiastically agrees while drawing a clear line between pre-training curation and post-training augmentation.22:40–27:43 · Guest teaching 7/10 Post-Training Strategy, Expert Domains, and MoE Considerations Thomas explains Meta's continuous pre-training strategy across expert domains and clarifies why dense models were chosen over MoE for this generation. Alessio and Shawn ask targeted follow-ups about curriculum learning and synthetic models.27:46–37:20 · Guest teaching 8/10 Demystifying RLHF: Human Discrimination vs Generation Alessio references Nathan Lambert's work and asks detailed questions about RLHF versus SFT impact. Thomas delivers a masterclass on human discrimination versus generation abilities, explaining how RLHF allows models to surpass human performance without human-written SFT data.37:21–45:38 · Guest teaching 6/10 Evaluating Llama 3: Benchmarks, Calibration, and Tool Calling Alessio and Shawn cite specific evaluation numbers and discuss calibration, uncertainty estimation, and tool use. Thomas agrees on calibration deficiencies in post-trained models and proposes calibration-specific prompts.45:38–54:36 · Guest teaching 7/10 Paving the Way for Llama 4: Agents and Latent Reasoning Shawn brings up Yann LeCun's JEPA architecture and Anthropic's hidden thinking prompt techniques. Thomas outlines the path toward Llama 4 and autonomous agents, discussing latent space reasoning and variable compute per token.54:45–59:15 · Guest teaching 7/10 Tokenizer Architecture and Multilingual Vocabulary Scaling Alessio queries the mechanical impact of scaling vocabulary size to 128K, prompting Thomas to explain token compression, compute trade-offs for small versus large models, and potential pixel-level or character-level tokenizers.0:45–4:11 · Guest disagreement 1/10 Career Path: From Quant Trading to NLP and Meta Shawn demonstrates strong background knowledge of Thomas's publishing history, recalling specific papers on summarization, factual consistency, and language GANs. Thomas explains his transition from quant trading to NLP right before BERT, with both sides sharing a collegial peer dynamic.4:11–8:03 · Guest disagreement 2/10 The Origins of Llama 2: From Galactica to Scale Thomas schools the hosts on the internal history of Galactica, the unexpected public backlash, and how Galactica Instruct pioneered RLHF annotation at Meta before Llama 2. Shawn and Alessio ask probing questions regarding annotation scale and open research questions.8:05–12:17 · Guest disagreement 2/10 LLM Scaling Laws and Escaping the Chinchilla Trap Alessio presses Thomas on parameter scaling laws and his tweet warning researchers not to fall into the Chinchilla trap. Thomas provides an in-depth breakdown of inference-optimal training versus compute-optimal training for published benchmarks.12:17–17:36 · Guest disagreement 2/10 Strategic Vision for Llama 3 405B and Quantization Alessio challenges the accessibility of training a massive 405B dense model for community inference. Thomas justifies the decision by explaining how FP8 quantization enables single-node execution and how 405B acts as a teacher for smaller models.17:36–22:40 · Guest disagreement 2/10 Pre-Training Data Quality and Synthetic Filtering in Llama 3 Shawn and Alessio raise technical points about synthetic data filtering and compare pre-training synthetic data to computer vision augmentation. Thomas enthusiastically agrees while drawing a clear line between pre-training curation and post-training augmentation.22:40–27:43 · Guest disagreement 3/10 Post-Training Strategy, Expert Domains, and MoE Considerations Thomas explains Meta's continuous pre-training strategy across expert domains and clarifies why dense models were chosen over MoE for this generation. Alessio and Shawn ask targeted follow-ups about curriculum learning and synthetic models.27:46–37:20 · Guest disagreement 2/10 Demystifying RLHF: Human Discrimination vs Generation Alessio references Nathan Lambert's work and asks detailed questions about RLHF versus SFT impact. Thomas delivers a masterclass on human discrimination versus generation abilities, explaining how RLHF allows models to surpass human performance without human-written SFT data.37:21–45:38 · Guest disagreement 2/10 Evaluating Llama 3: Benchmarks, Calibration, and Tool Calling Alessio and Shawn cite specific evaluation numbers and discuss calibration, uncertainty estimation, and tool use. Thomas agrees on calibration deficiencies in post-trained models and proposes calibration-specific prompts.45:38–54:36 · Guest disagreement 2/10 Paving the Way for Llama 4: Agents and Latent Reasoning Shawn brings up Yann LeCun's JEPA architecture and Anthropic's hidden thinking prompt techniques. Thomas outlines the path toward Llama 4 and autonomous agents, discussing latent space reasoning and variable compute per token.54:45–59:15 · Guest disagreement 2/10 Tokenizer Architecture and Multilingual Vocabulary Scaling Alessio queries the mechanical impact of scaling vocabulary size to 128K, prompting Thomas to explain token compression, compute trade-offs for small versus large models, and potential pixel-level or character-level tokenizers.0:45–4:11 · The hosts pushing back 1/10 Career Path: From Quant Trading to NLP and Meta Shawn demonstrates strong background knowledge of Thomas's publishing history, recalling specific papers on summarization, factual consistency, and language GANs. Thomas explains his transition from quant trading to NLP right before BERT, with both sides sharing a collegial peer dynamic.4:11–8:03 · The hosts pushing back 2/10 The Origins of Llama 2: From Galactica to Scale Thomas schools the hosts on the internal history of Galactica, the unexpected public backlash, and how Galactica Instruct pioneered RLHF annotation at Meta before Llama 2. Shawn and Alessio ask probing questions regarding annotation scale and open research questions.8:05–12:17 · The hosts pushing back 3/10 LLM Scaling Laws and Escaping the Chinchilla Trap Alessio presses Thomas on parameter scaling laws and his tweet warning researchers not to fall into the Chinchilla trap. Thomas provides an in-depth breakdown of inference-optimal training versus compute-optimal training for published benchmarks.12:17–17:36 · The hosts pushing back 3/10 Strategic Vision for Llama 3 405B and Quantization Alessio challenges the accessibility of training a massive 405B dense model for community inference. Thomas justifies the decision by explaining how FP8 quantization enables single-node execution and how 405B acts as a teacher for smaller models.17:36–22:40 · The hosts pushing back 2/10 Pre-Training Data Quality and Synthetic Filtering in Llama 3 Shawn and Alessio raise technical points about synthetic data filtering and compare pre-training synthetic data to computer vision augmentation. Thomas enthusiastically agrees while drawing a clear line between pre-training curation and post-training augmentation.22:40–27:43 · The hosts pushing back 2/10 Post-Training Strategy, Expert Domains, and MoE Considerations Thomas explains Meta's continuous pre-training strategy across expert domains and clarifies why dense models were chosen over MoE for this generation. Alessio and Shawn ask targeted follow-ups about curriculum learning and synthetic models.27:46–37:20 · The hosts pushing back 2/10 Demystifying RLHF: Human Discrimination vs Generation Alessio references Nathan Lambert's work and asks detailed questions about RLHF versus SFT impact. Thomas delivers a masterclass on human discrimination versus generation abilities, explaining how RLHF allows models to surpass human performance without human-written SFT data.37:21–45:38 · The hosts pushing back 2/10 Evaluating Llama 3: Benchmarks, Calibration, and Tool Calling Alessio and Shawn cite specific evaluation numbers and discuss calibration, uncertainty estimation, and tool use. Thomas agrees on calibration deficiencies in post-trained models and proposes calibration-specific prompts.45:38–54:36 · The hosts pushing back 2/10 Paving the Way for Llama 4: Agents and Latent Reasoning Shawn brings up Yann LeCun's JEPA architecture and Anthropic's hidden thinking prompt techniques. Thomas outlines the path toward Llama 4 and autonomous agents, discussing latent space reasoning and variable compute per token.54:45–59:15 · The hosts pushing back 2/10 Tokenizer Architecture and Multilingual Vocabulary Scaling Alessio queries the mechanical impact of scaling vocabulary size to 128K, prompting Thomas to explain token compression, compute trade-offs for small versus large models, and potential pixel-level or character-level tokenizers.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 5.2% · guest 94.8%0:00 · the hosts 5.2% · guest 94.8%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 15.8% · guest 84.2%6:00 · the hosts 15.8% · guest 84.2%9:00 · the hosts 15.2% · guest 84.8%9:00 · the hosts 15.2% · guest 84.8%12:00 · the hosts 15.4% · guest 84.6%12:00 · the hosts 15.4% · guest 84.6%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 13.9% · guest 86.1%18:00 · the hosts 13.9% · guest 86.1%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 7.4% · guest 92.6%24:00 · the hosts 7.4% · guest 92.6%27:00 · the hosts 30.1% · guest 69.9%27:00 · the hosts 30.1% · guest 69.9%30:00 · the hosts 20.7% · guest 79.3%30:00 · the hosts 20.7% · guest 79.3%33:00 · the hosts 7.9% · guest 92.1%33:00 · the hosts 7.9% · guest 92.1%36:00 · the hosts 23.1% · guest 76.9%36:00 · the hosts 23.1% · guest 76.9%39:00 · the hosts 13.4% · guest 86.6%39:00 · the hosts 13.4% · guest 86.6%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 14.4% · guest 85.6%48:00 · the hosts 14.4% · guest 85.6%51:00 · the hosts 20.4% · guest 79.6%51:00 · the hosts 20.4% · guest 79.6%54:00 · the hosts 25.1% · guest 74.9%54:00 · the hosts 25.1% · guest 74.9%57:00 · the hosts 26.5% · guest 73.5%57:00 · the hosts 26.5% · guest 73.5%1:00:00 · the hosts 6.4% · guest 93.6%1:00:00 · the hosts 6.4% · guest 93.6%1:03:00 · the hosts 44.4% · guest 55.6%1:03:00 · the hosts 44.4% · guest 55.6%
Sharpest disagreement ▶ 18:58 Dismissing raw web data as low quality

Thomas forcefully rejects the idea that raw web scraping is sufficient, stating bluntly that the web is full of poor-quality text that wastes compute without strong filtering.

Hardest push from the hosts ▶ 12:20 Alessio challenges the practicality of a 405B model

Alessio directly pushes back against the massive 405B parameter size, pointing out that regular developers and community members cannot run it on local hardware or easily find cloud resources.

Biggest teaching moment ▶ 30:45 Explaining why RLHF achieves superhuman performance

Thomas thoroughly educates the hosts on the core mathematical and behavioral intuitions behind RLHF, illustrating why human discrimination produces better targets than human SFT generation.

The host holds their own ▶ 50:30 Swyx dissects stopgap reasoning tokens and Anthropic system prompts

Shawn showcases deep architectural understanding by analyzing pause tokens, Anthropic's hidden reasoning tokens, and the fundamental requirement for variable inference compute in latent space.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Career Path: From Quant Trading to NLP and Meta 6311 Shawn demonstrates strong background knowledge of Thomas's publishing history, recalling specific papers on summarization, factual consistency, and language GANs. Thomas explains his transition from quant trading to NLP right before BERT, with both sides sharing a collegial peer dynamic.
The Origins of Llama 2: From Galactica to Scale 5622 Thomas schools the hosts on the internal history of Galactica, the unexpected public backlash, and how Galactica Instruct pioneered RLHF annotation at Meta before Llama 2. Shawn and Alessio ask probing questions regarding annotation scale and open research questions.
LLM Scaling Laws and Escaping the Chinchilla Trap 7723 Alessio presses Thomas on parameter scaling laws and his tweet warning researchers not to fall into the Chinchilla trap. Thomas provides an in-depth breakdown of inference-optimal training versus compute-optimal training for published benchmarks.
Strategic Vision for Llama 3 405B and Quantization 6523 Alessio challenges the accessibility of training a massive 405B dense model for community inference. Thomas justifies the decision by explaining how FP8 quantization enables single-node execution and how 405B acts as a teacher for smaller models.
Pre-Training Data Quality and Synthetic Filtering in Llama 3 7622 Shawn and Alessio raise technical points about synthetic data filtering and compare pre-training synthetic data to computer vision augmentation. Thomas enthusiastically agrees while drawing a clear line between pre-training curation and post-training augmentation.
Post-Training Strategy, Expert Domains, and MoE Considerations 5732 Thomas explains Meta's continuous pre-training strategy across expert domains and clarifies why dense models were chosen over MoE for this generation. Alessio and Shawn ask targeted follow-ups about curriculum learning and synthetic models.
Demystifying RLHF: Human Discrimination vs Generation 7822 Alessio references Nathan Lambert's work and asks detailed questions about RLHF versus SFT impact. Thomas delivers a masterclass on human discrimination versus generation abilities, explaining how RLHF allows models to surpass human performance without human-written SFT data.
Evaluating Llama 3: Benchmarks, Calibration, and Tool Calling 7622 Alessio and Shawn cite specific evaluation numbers and discuss calibration, uncertainty estimation, and tool use. Thomas agrees on calibration deficiencies in post-trained models and proposes calibration-specific prompts.
Paving the Way for Llama 4: Agents and Latent Reasoning 7722 Shawn brings up Yann LeCun's JEPA architecture and Anthropic's hidden thinking prompt techniques. Thomas outlines the path toward Llama 4 and autonomous agents, discussing latent space reasoning and variable compute per token.
Tokenizer Architecture and Multilingual Vocabulary Scaling 6722 Alessio queries the mechanical impact of scaling vocabulary size to 128K, prompting Thomas to explain token compression, compute trade-offs for small versus large models, and potential pixel-level or character-level tokenizers.

Statements from this episode (26)

Disclosure
Scialom: Multimodal Llama 3 will add parameters beyond 405B
“For the text text model only? Yes. A bit of additional parameters for the multimodal version that we come later.”
Thomas Scialom Jul 23, 2024 ▶ 0:36
Insight
Scialom: Multilinguality in LLMs emerges naturally with very little data
“Multilinguality almost emerged naturally with very, very few data, which was really surprising and not expected at all for us at the time.”
Thomas Scialom Jul 23, 2024 ▶ 3:48
Assertion Supported
Scialom: Meta had to reinvent scaling RLHF without published frontier research
“You have just the basics, but then when it comes to, like, ChatGPT or GPT Instruct or Cloud, No one published the details there. And so we had to reinvent the wheel there in a very short amount of time.”
Thomas Scialom Jul 23, 2024 ▶ 7:53
Disclosure
Scialom: Llama 1 and 2 flagship size was chosen to reproduce Chinchilla
“Lama two, maybe I would say it's like Lama one. We had a flagship model, which was seven TB. It's also because the project was taking some routes to reproducing a chinchilla, which was a seven TB.”
Thomas Scialom Jul 23, 2024 ▶ 9:23
Opinion
Scialom: OpenAI likely understood scaling laws before Chinchilla was published
“To be fair, I think OpenAI knew that at the time of Chinchilla paper.”
Thomas Scialom Jul 23, 2024 ▶ 10:55
Insight
Scialom: Overtrain models beyond Chinchilla optimal to minimize inference costs
“And so, to be compute efficient at inference time, it's much better to train it much longer training time, even if it's an effort, an additional effort, than to have a bigger model. That's what I call, like, I refer to the chinchilla trap, Not that Chinchilla …”
Thomas Scialom Jul 23, 2024 ▶ 11:44
Disclosure
Scialom: Next-generation Llama model will be bigger than Llama 3
“As, like, Mark announced, we have more and more GPUs, so the next generation will be bigger.”
Thomas Scialom Jul 23, 2024 ▶ 13:15
Assertion Supported
Scialom: FP8-quantized Llama 3 405B runs on a single compute node
“But quantizing it to FPA can run on node, even with a long context of 128 tokens.”
Thomas Scialom Jul 23, 2024 ▶ 13:29
Disclosure
Scialom: Smaller Llama 3 models improved via distillation from 405B
“Having bigger models enables to collect better data, for instance, at RLHF stage, because that's the model we use for the annotation. And so we distillate straight forward, like those annotations from this better model to the other models. So I can guarantee y…”
Thomas Scialom Jul 23, 2024 ▶ 14:02
Assertion Supported
Scialom: Llama 3 Scaled Pre-Training to 15 Trillion Tokens
“It's the same recipe done in terms of architectures and training than LAMA-II, but we put so much effort on scaling the data and the quality of data. There's now 15 triant tokens compared to two triants, so it's another magnitude there as well, including for t…”
Thomas Scialom Jul 23, 2024 ▶ 18:15
Disclosure
Scialom: Meta Used Llama 2 to Filter and Tag Llama 3 Pre-Training Data
“LAMA was the best, at the time, before LAMA Free, the best model we had access to legally, to labelize the web and select what are the good tokens and the bad tokens. The additional thing is that it also enabled to have a topic tag, Like, is it about law? Is i…”
Thomas Scialom Jul 23, 2024 ▶ 19:16
Disclosure
Meta changed Llama 3's pre-training data mixture mid-training run
“What happened is we changed the data mix during the training of Lama three with some findings that happened in the... Training is long, so you have to do something while it's training. And what the team did, I was working on my side of motion post-training, bu…”
Thomas Scialom Jul 23, 2024 ▶ 23:12
Disclosure
Meta skipped coding, reasoning, and multilingual annotations for Llama 2
“And we didn't annotate at all for code, neither for reasoning or multinguity.”
Thomas Scialom Jul 23, 2024 ▶ 25:17
Opinion
Scialom: Bullish on synthetic data, bearish on dedicated synthetic data models
“I'm very bullish on synthetic data generation, but I think just gets better when you have a better model. I'm not really bullish on having like a model only for synthetic data generation.”
Thomas Scialom Jul 23, 2024 ▶ 26:19
Disclosure
Scialom: Meta is exploring Mixture of Experts architectures for future models
“So, it's just an hyperparameter we haven't optimized a lot yet, but we have some stuff ongoing, and that's an hyperparameter we will explore in the future.”
Thomas Scialom Jul 23, 2024 ▶ 27:35
Insight
Scialom: RLHF yields superhuman models because humans judge better than they generate
“And because of that, you can have a model that flats the bad outputs, and learns to only shift towards the best and better and better outputs. And you can even end to superhuman abilities, since that I'm bad at writing a poem, but I'm good at judging which one…”
Thomas Scialom Jul 23, 2024 ▶ 31:56
Disclosure
Meta's Llama 3 post-training uses almost entirely synthetic data
“So what we did is that we generated all the data on the prompts with LAMA-II, and we applied, like, basically the last round of LAMA-II we had to kick off and start LAMA-III post-training. So Lama-free post-training doesn't have any, like, human-written answer…”
Thomas Scialom Jul 23, 2024 ▶ 33:40
Assertion Supported
Scialom: Llama 3 405B is the best open-source model ever released
“At a high level, it's the best open source model ever. It's Better than GPT-IV. I mean, what version? But, by far, compared to the version originally released even now, I think there's maybe the last cloud Sonya FF-V and GPT-IV-Zero that are performing it.”
Thomas Scialom Jul 23, 2024 ▶ 37:43
Assertion Supported
Scialom: Llama 3 natively supports state-of-the-art tool calling
“We have that from day one. Good news for the community. We are state of the art there. I think the model will be pretty good at that, we have a lot of gems about tools in the paper, but the model is fine-tuned to do tool usage, to zero-shot function calling”
Thomas Scialom Jul 23, 2024 ▶ 45:05
Prediction Not checkable as stated
Scialom: Agentic systems will yield order-of-magnitude scaling gains over pre-training
“I expect some incremental and significant progress on pre-training and post-training, but I'm really hopeful that we can gain some order of magnitude of scaling by interconnecting well models into agents as a more complex system that can do planning, that can …”
Thomas Scialom Jul 23, 2024 ▶ 47:14
Assertion Supported
Swyx: Anthropic simulates latent thinking by hiding prompt thinking tokens in Claude Artifacts
“Anthropic actually cheats at this right now. If you look at the system prompt in, in the cloud artifacts, I actually have a thinking section that is explicitly removed from the output, which is, I mean, they're still spending the tokens, but like that is befor…”
Shawn Wang Jul 23, 2024 ▶ 50:32
Insight
Scialom: Fixed compute per token in transformers forces models to generate extra tokens to think
“We are lacking of flexibility for pre-training architecture, transformers, where We spend the same amount of compute per token. And so, because of that, how can you, like, mitigate this? By generating more tokens, so more thoughts, more compute, because you ha…”
Thomas Scialom Jul 23, 2024 ▶ 51:19
Disclosure
Scialom: Meta expanded Llama 3's vocabulary size to support multilingual capabilities
“Lama III compared to Lama II is multilingual, has multilingual capabilities. We worked on that. And so, because you have languages that are not just Latin languages like English, there's a lot of different characters you want to include them to represent, like…”
Thomas Scialom Jul 23, 2024 ▶ 55:24
Insight
Scialom: Larger tokenizers allow models to see more text per compute unit
“With a bigger vocabulary, for the same text, you have less tokens, right? And so you can train your model on the same amount of knowledge with fewer steps. So for the same compute, you can see more knowledge if you don't epoch.”
Thomas Scialom Jul 23, 2024 ▶ 56:46
Assertion Partly supported
Scialom: Llama 3's 128K tokenizer reduces token count by about 30%
“Eight K basically, or 128 K, now with this tokenizer means 30% about less text to encode.”
Thomas Scialom Jul 23, 2024 ▶ 57:27
Opinion
Scialom: AI has minted infrastructure unicorns but few successful application companies
“I see like now a lot of fundamental stacks that are like the unicorn of today. Foundational models, foundational like clusters, data notations, things like that. There's a lot, but less successful yet, for now at least, application company. And it's hard to bu…”
Thomas Scialom Jul 23, 2024 ▶ 1:02:32
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.