Thomas Scialom

Senior Staff Research Scientist, Meta AI · 1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

scientistinvestoracademicLinkedIn ↗ai.meta.com ↗

Thomas Scialom is an artificial intelligence researcher best known as a co-creator of Meta's Llama models, serving as the lead for Llama 2 and post-training lead for Llama 3. He co-authored influential papers such as Toolformer and Galactica, and invests in early-stage generative AI startups.

25statements → 7claims → 6claims resolved → 83%fully supported → 3.88/5average certainty → 1.92/5average debate potential → 4.2/5argument clarity · the sources → 10said about them ↓

5 supported 1 partly supported 0 contradicted 1 not checkable as stated how the 7 claims stand · each chip opens the sources

1 prediction · 6 assertions · 3 opinions · 5 insights · 10 disclosures · every statement was checked. The prediction and assertions are the 7 claims: statements the public record can support or contradict. 6 are resolved, and 1 names no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Thomas argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
Scialom: Llama 3 405B is the best open-source model ever released
“At a high level, it's the best open source model ever. It's Better than GPT-IV. I mean, what version? But, by far, compared to the version originally released even now, I think there's maybe the last cloud Sonya FF-V and GPT-IV-Zero that are performing it.”
Thomas Scialom Jul 23, 2024 ▶ 37:43 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI

Expressed certainty vs assessment result

none yet certainty 1
none yet certainty 2
50% certainty 3
100% certainty 4
none yet certainty 5

weighted support: a fully supported claim counts one, a partly supported claim counts half. Each filled bar is clickable and opens exactly those claims; "none yet" means nothing said at that certainty level has resolved yet

Argument clarity: do they answer the question? how? →

4.2 / 5 directness 4.3 · coherence 4.4 · precision 4.2 · compression 3.8

answered every one of 12 assessed questions directly

This is a score against a rubric. It is not a rank. Every host question → answer exchange is scored with names hidden on directness, coherence, precision and compression, 1–5 each, on meaning alone: disfluencies are ignored, and only raw unedited episodes count. This is the score that measures thought. Every scored exchange, scores shown → · The rubric and its checks →

How they sound: not measured why? →

We measure speaking style by listening to the audio itself, and a fair number needs at least 2,000 words from one person on tape we have measured. There is too little of Thomas Scialom on measured tape to publish a rate. This says nothing about how they speak.

Everything Thomas Scialom said on Latent Space that made the record, most notable first. Filter by type, assessment or year in the ledger →

Insight
Scialom: Overtrain models beyond Chinchilla optimal to minimize inference costs
“And so, to be compute efficient at inference time, it's much better to train it much longer training time, even if it's an effort, an additional effort, than to have a bigger model. That's what I call, like, I refer to the chinchilla trap, Not that Chinchilla …”
Thomas Scialom Jul 23, 2024 ▶ 11:44 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Disclosure
Meta's Llama 3 post-training uses almost entirely synthetic data
“So what we did is that we generated all the data on the prompts with LAMA-II, and we applied, like, basically the last round of LAMA-II we had to kick off and start LAMA-III post-training. So Lama-free post-training doesn't have any, like, human-written answer…”
Thomas Scialom Jul 23, 2024 ▶ 33:40 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Assertion Supported
Scialom: Llama 3 405B is the best open-source model ever released
“At a high level, it's the best open source model ever. It's Better than GPT-IV. I mean, what version? But, by far, compared to the version originally released even now, I think there's maybe the last cloud Sonya FF-V and GPT-IV-Zero that are performing it.”
Thomas Scialom Jul 23, 2024 ▶ 37:43 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Prediction Not checkable as stated
Scialom: Agentic systems will yield order-of-magnitude scaling gains over pre-training
“I expect some incremental and significant progress on pre-training and post-training, but I'm really hopeful that we can gain some order of magnitude of scaling by interconnecting well models into agents as a more complex system that can do planning, that can …”
Thomas Scialom Jul 23, 2024 ▶ 47:14 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Opinion
Scialom: AI has minted infrastructure unicorns but few successful application companies
“I see like now a lot of fundamental stacks that are like the unicorn of today. Foundational models, foundational like clusters, data notations, things like that. There's a lot, but less successful yet, for now at least, application company. And it's hard to bu…”
Thomas Scialom Jul 23, 2024 ▶ 1:02:32 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Insight
Scialom: Multilinguality in LLMs emerges naturally with very little data
“Multilinguality almost emerged naturally with very, very few data, which was really surprising and not expected at all for us at the time.”
Thomas Scialom Jul 23, 2024 ▶ 3:48 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Opinion
Scialom: OpenAI likely understood scaling laws before Chinchilla was published
“To be fair, I think OpenAI knew that at the time of Chinchilla paper.”
Thomas Scialom Jul 23, 2024 ▶ 10:55 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Disclosure
Scialom: Next-generation Llama model will be bigger than Llama 3
“As, like, Mark announced, we have more and more GPUs, so the next generation will be bigger.”
Thomas Scialom Jul 23, 2024 ▶ 13:15 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Assertion Supported
Scialom: FP8-quantized Llama 3 405B runs on a single compute node
“But quantizing it to FPA can run on node, even with a long context of 128 tokens.”
Thomas Scialom Jul 23, 2024 ▶ 13:29 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Disclosure
Scialom: Smaller Llama 3 models improved via distillation from 405B
“Having bigger models enables to collect better data, for instance, at RLHF stage, because that's the model we use for the annotation. And so we distillate straight forward, like those annotations from this better model to the other models. So I can guarantee y…”
Thomas Scialom Jul 23, 2024 ▶ 14:02 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Disclosure
Scialom: Meta Used Llama 2 to Filter and Tag Llama 3 Pre-Training Data
“LAMA was the best, at the time, before LAMA Free, the best model we had access to legally, to labelize the web and select what are the good tokens and the bad tokens. The additional thing is that it also enabled to have a topic tag, Like, is it about law? Is i…”
Thomas Scialom Jul 23, 2024 ▶ 19:16 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Disclosure
Meta changed Llama 3's pre-training data mixture mid-training run
“What happened is we changed the data mix during the training of Lama three with some findings that happened in the... Training is long, so you have to do something while it's training. And what the team did, I was working on my side of motion post-training, bu…”
Thomas Scialom Jul 23, 2024 ▶ 23:12 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Opinion
Scialom: Bullish on synthetic data, bearish on dedicated synthetic data models
“I'm very bullish on synthetic data generation, but I think just gets better when you have a better model. I'm not really bullish on having like a model only for synthetic data generation.”
Thomas Scialom Jul 23, 2024 ▶ 26:19 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Insight
Scialom: RLHF yields superhuman models because humans judge better than they generate
“And because of that, you can have a model that flats the bad outputs, and learns to only shift towards the best and better and better outputs. And you can even end to superhuman abilities, since that I'm bad at writing a poem, but I'm good at judging which one…”
Thomas Scialom Jul 23, 2024 ▶ 31:56 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Assertion Supported
Scialom: Llama 3 natively supports state-of-the-art tool calling
“We have that from day one. Good news for the community. We are state of the art there. I think the model will be pretty good at that, we have a lot of gems about tools in the paper, but the model is fine-tuned to do tool usage, to zero-shot function calling”
Thomas Scialom Jul 23, 2024 ▶ 45:05 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Assertion Supported
Scialom: Meta had to reinvent scaling RLHF without published frontier research
“You have just the basics, but then when it comes to, like, ChatGPT or GPT Instruct or Cloud, No one published the details there. And so we had to reinvent the wheel there in a very short amount of time.”
Thomas Scialom Jul 23, 2024 ▶ 7:53 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Assertion Supported
Scialom: Llama 3 Scaled Pre-Training to 15 Trillion Tokens
“It's the same recipe done in terms of architectures and training than LAMA-II, but we put so much effort on scaling the data and the quality of data. There's now 15 triant tokens compared to two triants, so it's another magnitude there as well, including for t…”
Thomas Scialom Jul 23, 2024 ▶ 18:15 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Disclosure
Meta skipped coding, reasoning, and multilingual annotations for Llama 2
“And we didn't annotate at all for code, neither for reasoning or multinguity.”
Thomas Scialom Jul 23, 2024 ▶ 25:17 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Disclosure
Scialom: Meta is exploring Mixture of Experts architectures for future models
“So, it's just an hyperparameter we haven't optimized a lot yet, but we have some stuff ongoing, and that's an hyperparameter we will explore in the future.”
Thomas Scialom Jul 23, 2024 ▶ 27:35 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Insight
Scialom: Fixed compute per token in transformers forces models to generate extra tokens to think
“We are lacking of flexibility for pre-training architecture, transformers, where We spend the same amount of compute per token. And so, because of that, how can you, like, mitigate this? By generating more tokens, so more thoughts, more compute, because you ha…”
Thomas Scialom Jul 23, 2024 ▶ 51:19 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Disclosure
Scialom: Meta expanded Llama 3's vocabulary size to support multilingual capabilities
“Lama III compared to Lama II is multilingual, has multilingual capabilities. We worked on that. And so, because you have languages that are not just Latin languages like English, there's a lot of different characters you want to include them to represent, like…”
Thomas Scialom Jul 23, 2024 ▶ 55:24 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Disclosure
Scialom: Multimodal Llama 3 will add parameters beyond 405B
“For the text text model only? Yes. A bit of additional parameters for the multimodal version that we come later.”
Thomas Scialom Jul 23, 2024 ▶ 0:36 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Disclosure
Scialom: Llama 1 and 2 flagship size was chosen to reproduce Chinchilla
“Lama two, maybe I would say it's like Lama one. We had a flagship model, which was seven TB. It's also because the project was taking some routes to reproducing a chinchilla, which was a seven TB.”
Thomas Scialom Jul 23, 2024 ▶ 9:23 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Insight
Scialom: Larger tokenizers allow models to see more text per compute unit
“With a bigger vocabulary, for the same text, you have less tokens, right? And so you can train your model on the same amount of knowledge with fewer steps. So for the same compute, you can see more knowledge if you don't epoch.”
Thomas Scialom Jul 23, 2024 ▶ 56:46 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI

Show 1statements(1 left)

The other half of the tape: Thomas Scialom's own voice is left out of every number here. Other people bring the name up 10 times in 5 episodes on Latent Space. every mention, with the transcript →

Who brings them up most Alessio Fanelli 5Shawn Wang 1

Every mention by year

tap a year for its mentions
005210420242025episodesmentions
02420242025episodes it came up in
001.322.5420242025episodesmentions per episode

Appearances (1)

EpisodeDateSpeaking time
Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI Jul 23, 2024 39m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.