Llama 3.1 405B, every mention
12 scenes · ← back to Llama 3.1 405B
tap a year for its mentions
every year anyone Shawn Wang 6Alessio Fanelli 5Alistair Pullen 2Eugene Cheah 1
Verbatim, from the transcripts: the passages where Llama 3.1 405B comes up
The Unreasonable Effectiveness of Reasoning Distillation: using DeepSeek R1 to beat OpenAI o1
- ▶ 18:40 unnamed speaker So, the main point, yeah, the, that's a, the main point I'm getting to is like, for that particular model, it's a seven B model, and the data comes from 70 B, ah, sorry, four oh five B, Lama five B, and we were able to beat the Lama four…
2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
- ▶ 14:11 Shawn Wang LAMA-FORO-FIB does well compared to Gemini and GPT-E-FORO. 5 times in the scene
- ▶ 1:17:43 Shawn Wang Um, and I think what you're starting to see now, uh, in July is the emergence of four O mini and deep sea V two as outliers to the July frontier where July frontier used to be maintained by four O Lama four five.
Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI
- ▶ 32:43 Alessio Fanelli Because I think everybody agrees with open source models eventually will catch up, and I think with four, then with Lama to three, one, four, five B, we close the gap, and then all one just reopened the gap so much, and it's unclear.
Building AGI in Real Time (OpenAI Dev Day 2024)
- ▶ 1:14:09 Alistair Pullen Fine tune four or five B on the same data set, like same context window length, right? 2 times in the scene
The Winds of AI Winter (Q2 Four Wars of the AI Stack Recap)
- ▶ 7:39 Alessio Fanelli And I think if we can apply the same principles to like a model as big as four or five B and bring them into like maybe the seven B form factor, that would be great. 4 times in the scene
[LLM Paper Club] Llama 3.1 Paper: The Llama Family of Models
- ▶ 10:02 unnamed speaker Their scaling laws show that for a four or two B model, you want to train on 16 and a half trillion tokens. 3 times in the scene
- ▶ 18:01 Eugene Cheah And the reason why they probably need it for, for the ATGB is because they're training 400 or five B and yeah, and
- ▶ 24:35 unnamed speaker There's a scale AI benchmark where, um, Sonnet and, uh, four O were compared against the new four or five B model. 4 times in the scene
- ▶ 40:13 unnamed speaker And that's, I think the big part of the license play of this too, where, um, they, they did actually finally change their license to allow people to generate synthetic data, train on outputs of the four or five B I think the four or five B… 2 times in the scene
- ▶ 58:28 unnamed speaker I'm curious if you could share, what does it take to serve a four or five B model on together? 3 times in the scene
- ▶ 1:11:48 unnamed speaker so basically i found like a llama 3.1 like a four oh five b actually might do better in some domains specific question than like a chat gpt 6 times in the scene