Llama 3.1, every mention
24 scenes, the whole family · ← back to Llama 3.1
tap a year for its mentions
every year anyone Alessio Fanelli 8Shawn Wang 6Nathan Lambert 3Alistair Pullen 2Eugene Cheah 1Andrej Karpathy 1
Verbatim, from the transcripts: the passages where Llama 3.1 comes up
The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
- ▶ 2:31 Nathan Lambert I think meta has different priorities and their things for llama 3.1, which is a great set of models at the time. 2 times in the scene
The Unreasonable Effectiveness of Reasoning Distillation: using DeepSeek R1 to beat OpenAI o1
- ▶ 18:40 unnamed speaker So, the main point, yeah, the, that's a, the main point I'm getting to is like, for that particular model, it's a seven B model, and the data comes from 70 B, ah, sorry, four oh five B, Lama five B, and we were able to beat the Lama four…
The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024]
- ▶ 12:38 Nathan Lambert An example that I used in the blog post I wrote today on this is like Lama, 3.1 details their vows for math.
2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
- ▶ 14:11 Shawn Wang LAMA-FORO-FIB does well compared to Gemini and GPT-E-FORO. 5 times in the scene
- ▶ 1:17:43 Shawn Wang Um, and I think what you're starting to see now, uh, in July is the emergence of four O mini and deep sea V two as outliers to the July frontier where July frontier used to be maintained by four O Lama four five.
- ▶ 1:33:13 Alessio Fanelli And then July was Lama three.
Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI
- ▶ 32:43 Alessio Fanelli Because I think everybody agrees with open source models eventually will catch up, and I think with four, then with Lama to three, one, four, five B, we close the gap, and then all one just reopened the gap so much, and it's unclear.
Building AGI in Real Time (OpenAI Dev Day 2024)
- ▶ 1:14:09 Alistair Pullen Fine tune four or five B on the same data set, like same context window length, right? 2 times in the scene
llm.c's Origin and the Future of LLM Compilers - Andrej Karpathy at CUDA MODE
- ▶ 20:28 Andrej Karpathy We actually thought maybe we would have it done by today, but, uh, there's a few more, a few more, a little bit more work to do, but we will have Lama 3.1, um, training in Lama.c very, very soon.
The Winds of AI Winter (Q2 Four Wars of the AI Stack Recap)
- ▶ 7:06 Alessio Fanelli I think it would be amazing if we could replicate the same on the open models to then, because now we can use Lama 3.1 to generate synthetic data for training and fine tuning. 2 times in the scene
- ▶ 7:39 Alessio Fanelli And I think if we can apply the same principles to like a model as big as four or five B and bring them into like maybe the seven B form factor, that would be great. 4 times in the scene
- ▶ 1:12:38 unnamed speaker And that means lifting up the, like, I had this difference chart between Lama, three point zero, eight B Lama, 3.07 TB versus the 3.1 differences.
[LLM Paper Club] Llama 3.1 Paper: The Llama Family of Models
- ▶ 3:07 unnamed speaker So like they dropped three to three, one, three, one was a pretty big update. 2 times in the scene
- ▶ 3:11 unnamed speaker The eight B got a lot better than 70 B got a lot better. 2 times in the scene
- ▶ 3:11 unnamed speaker The eight B got a lot better than 70 B got a lot better. 4 times in the scene
- ▶ 10:02 unnamed speaker Their scaling laws show that for a four or two B model, you want to train on 16 and a half trillion tokens. 3 times in the scene
- ▶ 18:01 Eugene Cheah And the reason why they probably need it for, for the ATGB is because they're training 400 or five B and yeah, and
- ▶ 24:35 unnamed speaker There's a scale AI benchmark where, um, Sonnet and, uh, four O were compared against the new four or five B model. 4 times in the scene
- ▶ 40:13 unnamed speaker And that's, I think the big part of the license play of this too, where, um, they, they did actually finally change their license to allow people to generate synthetic data, train on outputs of the four or five B I think the four or five B… 2 times in the scene
- ▶ 53:03 unnamed speaker Um, but yeah, a lot of, yeah, a lot of, a lot of good stuff, uh, in the five to seven minutes we have left, uh, I wanted to give some time to, I guess, Hassan, if you want to, Hassan's actually built an app that's kind of cool with, uh,…
- ▶ 53:59 unnamed speaker Cause obviously Lama 3.1 larger context, you can fit in a lot of stuff, which is great. 3 times in the scene
- ▶ 56:09 unnamed speaker Um, and so if you do the math, that's like, seventy four million, uh, tokens, and if I go over to our pricing, right now we're at 18 cents per million tokens for eight B 3.1, um, and so that comes out to about 12 dollars, um, from the…
- ▶ 58:28 unnamed speaker I'm curious if you could share, what does it take to serve a four or five B model on together? 3 times in the scene
- ▶ 1:11:48 unnamed speaker so basically i found like a llama 3.1 like a four oh five b actually might do better in some domains specific question than like a chat gpt 6 times in the scene