Llama 2, every mention

54 scenes, the whole family · ← back to Llama 2

tap a year for its mentions
0040880152023202420252026episodesmentions
08152023202420252026episodes it came up in
0047.58152023202420252026episodesmentions per episode

every year anyone Shawn Wang 17Thomas Scialom 10Nathan Lambert 8Soumith Chintala 7Dylan Patel 6Ben Firshman 5Jerry Liu 4Alessio Fanelli 4Eugene Yan 3RJ Haneke 2

Verbatim, from the transcripts: the passages where Llama 2 comes up

loading…

Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten Aug 3, 2026 · 1 mention

🔬 "The Most Innovative Diffusion Research Is Happening in Drug Discovery, Not Image Generation" Jun 30, 2026 · 3 mentions

  • ▶ 0:55 RJ Haneke I'm joined by my co-host Brandon Anderson, and we're privileged to have Evan Feinberg, founder and CEO of Genesis Molecular AI, and Sergei Yudinov, who led Lama II and Lama III pre-training before he joined Genesis as CTO. 2 times in the scene
  • ▶ 20:36 Evan Feinberg Sergei is being humble, but Sergei led the LLAMA II research team at, at, at Meta when, when, when he was still there.

⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF Oct 20, 2025 · 2 mentions

  • ▶ 8:54 Elie Bakouch And for example, a good, uh, a good way to view that is that, uh, DeepSeq rig three is still using the same Adam parameter than, uh, Lama two.
  • ▶ 9:05 Elie Bakouch I don't know if it's just me, but I feel that like the, the, the, the hyperpenters for Lama two AB, for example, shouldn't be the optimal one for, uh, this, uh, mega DeepSeq model with, uh, with a lot of, uh, of parameter.

Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave) Oct 16, 2025 · 1 mention

  • ▶ 8:09 Kyle Corbitt Yeah, they were really strong models, um, better than the Llama two that they were, you know, effectively replacing.

A Technical History of Generative Media Sep 8, 2025 · 1 mention

  • ▶ 8:39 Gorkem Yurtseven And then obviously after stable diffusion, I think like four or five months later, LAMA-II came out and, um, there was a decision point again.

Better Data is All You Need — Ari Morcos, Datology Aug 29, 2025 · 1 mention

⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo Jul 14, 2025 · 1 mention

2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents Jan 1, 2025 · 1 mention

Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024] Dec 24, 2024 · 1 mention

  • ▶ 17:04 Loubna Ben Allal For example, Lama 3.21 B it matches Lama two 13 B from that was the release last year on the LMSS arena, which is basically the default go to leaderboard for evaluating models using human evaluation.

[LLM Paper Club] Llama 3.1 Paper: The Llama Family of Models Jul 29, 2024 · 5 mentions

  • ▶ 7:40 unnamed speaker They used a lot of, like, they trained classifiers on Lama II outputs.
  • ▶ 23:12 unnamed speaker Uh, they, they, like, uh, this is a big contrast to Lama two, where they were intentionally not training for code, and then they put out code Lama separately.
  • ▶ 31:10 Eugene Yan Like in the first one, you can see they actually use LAMA tool to filter out bad data, right? 3 times in the scene

Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI Jul 23, 2024 · 28 mentions

  • ▶ 0:18 Shawn Wang You've, you've done so much work in a very short amount of time at Meta, but you were most notably leading Llama II.
  • ▶ 4:12 Shawn Wang Maybe we should get right into Lama two, spend a little bit of time there, and then, and then we'll go into Lama three. 4 times in the scene
  • ▶ 8:15 Alessio Fanelli And then Lama two, seven, 13, 70. 2 times in the scene
  • ▶ 12:21 Alessio Fanelli So Llama Two, you have a pretty good model.
  • ▶ 17:38 Shawn Wang Like, uh, I'm not sure when you're releasing the Lama three research paper, but in Lama two, you talked a little bit about, uh, the architecture choices, like in any, 6 times in the scene
  • ▶ 22:59 Shawn Wang We know that that's changed for between Lama two and Lama three. 4 times in the scene
  • ▶ 28:39 Thomas Scialom When we started, uh, Lama Two, I had, like, this budget of annotations in millions of dollars, and, okay, what to do? 5 times in the scene
  • ▶ 40:06 Thomas Scialom I feel it was much easier during LAMA II 3 times in the scene
  • ▶ 46:27 Thomas Scialom So we need to work on that, and we did LAMA-II, and then now LAMA-III.
  • ▶ 55:24 Thomas Scialom The first thing obvious to say is Lama III compared to Lama II is multilingual, has multilingual capabilities.

The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka Jul 5, 2024 · 7 mentions

  • ▶ 1:07:21 unnamed speaker Which Lama-II now. 3 times in the scene
  • ▶ 1:20:12 unnamed speaker Um, you commented on a few things, like, Lama one, two, three glowed up a lot. 3 times in the scene
  • ▶ 1:37:28 Yi Tay And then, people make so much big deal about, about, uh, uh, like, uh, you know, trading past Chinchilla Scaling Law, like, oh, Lamao-Doo's the first mall, like, like, like, T-Five base, right, was one trillion tokens, that was really so…

How to train a Million Context LLM — with Mark Huang of Gradient.ai May 31, 2024 · 2 mentions

  • ▶ 16:05 Mark Huang You know, the final limit that you have, like, 32 K is usually the limit of Lama two was, was that long.
  • ▶ 33:03 Mark Huang Um, which, which was seen, like, we do have historical precedent where CodeLlama was, you know, trained further from the, the, the original CodeLlama was trained further from Lama II, and it just lost,

LLM Asia Paper Club Survey Round May 22, 2024 · 1 mention

  • ▶ 23:28 unnamed speaker In the paper, they treat, they do the experiments with Lama seven B and Gemma, I think Lama two seven B and Gemma seven B as black boxes and vice using the other open source model as the white box for the uncertainty estimation.

Breaking down the OG GPT Paper by Alec Radford Apr 23, 2024 · 3 mentions

  • ▶ 44:16 unnamed speaker I think this was used in one of the Lama II models. 3 times in the scene

Personal AI Meetup - Bee, BasedHardware, LangChain LangFriend, Deepgram EmilyAI Apr 6, 2024 · 1 mention

  • ▶ 13:57 Damien Murphy Um, but some, some, some of our customers will actually run their own, um, like LAMA-II model, uh, super close to the, the GPUs that are running the speech-to-text and text-to-speech, and that just removes all the network latency out of…

Open Source AI is AI we can Trust — with Soumith Chintala of Meta AI Mar 6, 2024 · 7 mentions

  • ▶ 44:11 Soumith Chintala Um, LAMA II, I was more closely involved in, um, I helped them a reasonable amount with, like, their, um, infrastructure needs and stuff. 7 times in the scene

A Brief History of the Open Source AI Hacker - with Ben Firshman of Replicate Feb 28, 2024 · 11 mentions

  • ▶ 39:11 Ben Firshman Then, kind of, the same thing happened, like, middle of last year with language models in Llama II, where the same kind of stable diffusion effect happened with, with Llama. 3 times in the scene
  • ▶ 1:04:17 Ben Firshman So models like Llama II, like Stable Diffusion, we, um, we, you know, we work with Meta and Stability to like maintain those models, and we've done a ton of optimizations to make those really fast.
  • ▶ 1:08:37 Ben Firshman We, actually, for our, so for Llama II and Mistral, I think not Mixtral, I can't remember exactly, we have, you know, similar performance and similar price to some of these other services.
  • ▶ 1:11:02 unnamed speaker When Lama two came out, uh, I've wrote a post about this, about it's like open source and there's open weights, then there's restrictive weights. 6 times in the scene

The State of AI in production — with David Hsu of Retool Feb 7, 2024 · 1 mention

  • ▶ 50:11 David Hsu Now, however, where we are right now is I think GPT-IV is so far ahead of terms of performance that, and I would say, ah, model performance is so important right now because like the average, you know, like, you know, ah, you can argue…

The Four Wars of the AI Stack - Dec 2023 Recap Jan 26, 2024 · 1 mention

  • ▶ 2:44 unnamed speaker I think over the last maybe like four or five months, everybody's so focused on, uh, fine tuning Lama two and like a DPO to improve these models, max trial and all these things.

The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert Jan 11, 2024 · 12 mentions

  • ▶ 4:39 Nathan Lambert I think the quote from the Llama II paper is a great kind of tidbit on 3 times in the scene
  • ▶ 44:56 Nathan Lambert Um, like, LLAMA-II is famous for kind of having this, like, helpfulness and safety reward models.
  • ▶ 46:10 unnamed speaker Uh, I think you may have joined one or, one of our spaces back in, uh, when Lama II was released. 3 times in the scene
  • ▶ 52:05 unnamed speaker Um, just in a domain of, uh, human preference data suppliers, uh, Scali, I very happily will tell you that they, they supplied, uh, all that data for Lama too.
  • ▶ 1:02:35 Nathan Lambert The interesting thing that people are confused about more is rejection sampling, because Meta talked about it in Llama II. 2 times in the scene
  • ▶ 1:20:35 Nathan Lambert Starting pre-training is very hard, so it's like you still want to, you still want to learn from Llama II and Llama III, so that's fun.
  • ▶ 1:30:27 Nathan Lambert We could swap between Llama-II and Mixtrawl and kind of see, like, does RLHF work

The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis Dec 5, 2023 · 6 mentions

  • ▶ 16:54 Dylan Patel Llama-seventy-b was two million batch size, and like, you talk to someone at one of the frontier labs, and they're like, ha, right? 4 times in the scene
  • ▶ 26:06 Dylan Patel Hey, to run Llama's seventy billion requires two terabytes a second of memory bandwidth, 2.1, at reading, human reading speed. 2 times in the scene

Beating GPT-4 with Open Source Models - with Michael Royzen of Phind Nov 3, 2023 · 1 mention

  • ▶ 37:24 Michael Royzen Um, over time, it's, I think, it's become more clear with the release of Llama II and Llama III on the horizon that we will once again see a return to, um, vertical applications running their own models.

Why AI Agents Don't Work (yet) - with Kanjun Qiu of Imbue Oct 21, 2023 · 1 mention

  • ▶ 20:27 Shawn Wang So, uh, one way I'll put this is, um, people estimate that Llama II maybe took about three, four million dollars to compute, but probably 20, twenty-five million dollars worth of labeling data.

The End of Finetuning — with Jeremy Howard of Fast.ai Oct 20, 2023 · 1 mention

  • ▶ 43:26 Jeremy Howard So Code Llama was a, I think it was like a five hundred billion token fine-tuning of Llama II using code.

RAG is a hack - with Jerry Liu of LlamaIndex Oct 12, 2023 · 4 mentions

  • ▶ 57:29 Jerry Liu I think there's a lot of people trying out, like, Llama II. 4 times in the scene

RWKV: Reinventing RNNs for the Transformer Era Aug 31, 2023 · 2 mentions

  • ▶ 1:27:35 unnamed speaker Um, but right now, let's say, you know, uh, Lama II was trained on two trillion tokens. 2 times in the scene

FlashAttention-2: Making Transformers 800% faster AND exact Aug 3, 2023 · 7 mentions

  • ▶ 51:12 unnamed speaker Llama two came out on Tuesday? 7 times in the scene
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.