Llama 3.1 405B, every mention

12 scenes · ← back to Llama 3.1 405B

tap a year for its mentions
0015230420242025episodesmentions
02420242025episodes it came up in
00428420242025episodesmentions per episode

every year anyone Shawn Wang 6Alessio Fanelli 5Alistair Pullen 2Eugene Cheah 1

Verbatim, from the transcripts: the passages where Llama 3.1 405B comes up

loading…

The Unreasonable Effectiveness of Reasoning Distillation: using DeepSeek R1 to beat OpenAI o1 Jan 24, 2025 · 1 mention

  • ▶ 18:40 unnamed speaker So, the main point, yeah, the, that's a, the main point I'm getting to is like, for that particular model, it's a seven B model, and the data comes from 70 B, ah, sorry, four oh five B, Lama five B, and we were able to beat the Lama four…

2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents Jan 1, 2025 · 6 mentions

  • ▶ 14:11 Shawn Wang LAMA-FORO-FIB does well compared to Gemini and GPT-E-FORO. 5 times in the scene
  • ▶ 1:17:43 Shawn Wang Um, and I think what you're starting to see now, uh, in July is the emergence of four O mini and deep sea V two as outliers to the July frontier where July frontier used to be maintained by four O Lama four five.

Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI Nov 25, 2024 · 1 mention

  • ▶ 32:43 Alessio Fanelli Because I think everybody agrees with open source models eventually will catch up, and I think with four, then with Lama to three, one, four, five B, we close the gap, and then all one just reopened the gap so much, and it's unclear.

Building AGI in Real Time (OpenAI Dev Day 2024) Oct 4, 2024 · 2 mentions

  • ▶ 1:14:09 Alistair Pullen Fine tune four or five B on the same data set, like same context window length, right? 2 times in the scene

The Winds of AI Winter (Q2 Four Wars of the AI Stack Recap) Aug 2, 2024 · 4 mentions

  • ▶ 7:39 Alessio Fanelli And I think if we can apply the same principles to like a model as big as four or five B and bring them into like maybe the seven B form factor, that would be great. 4 times in the scene

[LLM Paper Club] Llama 3.1 Paper: The Llama Family of Models Jul 29, 2024 · 19 mentions

  • ▶ 10:02 unnamed speaker Their scaling laws show that for a four or two B model, you want to train on 16 and a half trillion tokens. 3 times in the scene
  • ▶ 18:01 Eugene Cheah And the reason why they probably need it for, for the ATGB is because they're training 400 or five B and yeah, and
  • ▶ 24:35 unnamed speaker There's a scale AI benchmark where, um, Sonnet and, uh, four O were compared against the new four or five B model. 4 times in the scene
  • ▶ 40:13 unnamed speaker And that's, I think the big part of the license play of this too, where, um, they, they did actually finally change their license to allow people to generate synthetic data, train on outputs of the four or five B I think the four or five B 2 times in the scene
  • ▶ 58:28 unnamed speaker I'm curious if you could share, what does it take to serve a four or five B model on together? 3 times in the scene
  • ▶ 1:11:48 unnamed speaker so basically i found like a llama 3.1 like a four oh five b actually might do better in some domains specific question than like a chat gpt 6 times in the scene
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.