Llama, every mention
279 scenes, the whole family · ← back to Llama
tap a year for its mentions
every year anyone Shawn Wang 63Alessio Fanelli 35Thomas Scialom 22Nathan Lambert 21Soumith Chintala 15George Hotz 13Yining Zhang 9Lin Qiao 7Pratik Bhavsar 6Mark Huang 6
Verbatim, from the transcripts: the passages where Llama comes up
Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
- ▶ 1:09:39 Philip Kiely Llama, not Llama two, but Llama three.
- ▶ 1:09:39 Philip Kiely Llama, not Llama two, but Llama three. 3 times in the scene
The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
The AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO
- ▶ 22:39 Dan Biderman Um, so the examples we like to give is that, uh, if you take a Lama, a 70 B model, and you load one article from Wikipedia, which is a few tens of kilobytes, and you have the model read this, uh,
The Future of AI Infra: from Kubernetes to Agent Sandboxes — Akshat Bubna, Modal CTO
- ▶ 51:28 Akshat Bubna Swapping out the tokenizer in Lama and whatnot.
🔬 "The Most Innovative Diffusion Research Is Happening in Drug Discovery, Not Image Generation"
- ▶ 0:55 RJ Haneke I'm joined by my co-host Brandon Anderson, and we're privileged to have Evan Feinberg, founder and CEO of Genesis Molecular AI, and Sergei Yudinov, who led Lama II and Lama III pre-training before he joined Genesis as CTO. 2 times in the scene
- ▶ 0:55 RJ Haneke I'm joined by my co-host Brandon Anderson, and we're privileged to have Evan Feinberg, founder and CEO of Genesis Molecular AI, and Sergei Yudinov, who led Lama II and Lama III pre-training before he joined Genesis as CTO. 2 times in the scene
- ▶ 1:42 Sergei Yudinov I later on led LAMA team, um, LAMA two and LAMA three models.
- ▶ 20:36 Evan Feinberg Sergei is being humble, but Sergei led the LLAMA II research team at, at, at Meta when, when, when he was still there.
The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin
- ▶ 1:00:03 Matei Zaharia So we, uh, decided, you know, even though we, we did launch, uh, open source model DBRX, and, you know, we, we went up to, like, sort of above the LAMA-R III scale, we decided that we really want to focus on, there'll be so many people…
AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Why AI Labs With Unlimited GPUs Still Fail — Anjney Midha, AMP
- ▶ 37:25 Anjney Midha Like, you, you are an athlete of the mind, and you perform at the highest levels, and to get there, whether you're, you know, Anastasius or Waylon at Berkeley, or you are Robin, who, with Black Forest and created Stable Diffusion, or if…
Satya Nadella on AI: @NoPriorsPodcast x Latent Space Crossover Special at Microsoft Build 2026
- ▶ 10:14 Satya Nadella Like you can use your Lama harness, whatever, or you can use the, um, uh, you know, any open harness or any harness of yours and train with your tools and multiple models and your context.
Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
- ▶ 20:33 Shawn Wang And Lama as well?
Mistral: Voxtral TTS, Forge, Leanstral, & Mistral 4 — w/ Pavan Kumar Reddy & Guillaume Lample
- ▶ 38:05 Guillaume Lample So, me and Tim were at Meta, we released Lama, and I think what was really nice to see that before this, for most researchers, like universities, it was impossible to, to work on LLNs. 2 times in the scene
Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
- ▶ 33:30 unnamed speaker We're not the best at training MOEs when they're pre-trained, like we saw this with LAMA-III, right?
- ▶ 52:21 Kyle Kranen And previously context, like I think the, the Lama four or five B context of a similar size was like 40 or 80 gigabytes in the same precision.
Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
- ▶ 56:55 Shawn Wang And so Lama had this, like, if you have seven hundred million daily active users, you're not allowed to use our model or you have to talk to us, something like that.
[State of AI Papers 2025] Fixing Research with Social Signals, OCR & Implementation — Team AlphaXiv
- ▶ 3:45 unnamed speaker I think when we started, we were very deliberate about getting authors, like, from LoRa, DPO, Lama, and actually having really useful, cool exchanges.
SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow)
- ▶ 26:56 Nikhila Ravi Um, we also use Llama in our data engine. 2 times in the scene
- ▶ 31:03 unnamed speaker Interestingly, you use Lama four. 2 times in the scene
- ▶ 31:05 unnamed speaker I saw there's a mix of Lama three and Lama four here, but it looks like it does best with Gemini 2.5, which makes sense given this comparable set of MLMs.
- ▶ 39:01 Pengchuan Zhang Source images and it generates kind of non-phases from, for example, kind of NAMA generate caption and we pass the caption to get the non-phases.
- ▶ 41:14 Pengchuan Zhang That is a breakthrough, and then, kind of, we kind of fine-tune our, kind of, for example, NAMA's 3.2 with our, kind of, verification, human annotated verification data.
The Future of Email: Superhuman CTO on Your Inbox As the Real AI Agent (Not ChatGPT) — Loïc Houssier
- ▶ 30:12 Loïc Houssier Uh, but, uh, we use Base-Ten to run some, uh, I would say some LAMA, some BERT model for classification.
⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
- ▶ 8:54 Elie Bakouch And for example, a good, uh, a good way to view that is that, uh, DeepSeq rig three is still using the same Adam parameter than, uh, Lama two.
- ▶ 1:00:25 Alessio Fanelli Like, uh, I have this, like a MCP client, I built, and we use Lama, um, AB for like a conversation naming, you know, I feel like that's like a great use case for like a five hundred million parameter model, but there's no simple API to use…
Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
- ▶ 8:09 Kyle Corbitt Yeah, they were really strong models, um, better than the Llama two that they were, you know, effectively replacing.
Building Jamba 3B: the tiny Hybrid Transformer State Space Reasoning Model - Barak Lenz, CTO of AI21
- ▶ 6:09 Barak Lenz Straight off the bat, compare themselves to, to latest, uh, attention architecture introduced by Lama that, that introduced a lot of corrections that cause things to work. 2 times in the scene
⚡️Raising $1.1b to build the fastest LLM Chips on Earth — Andrew Feldman, Cerebras
- ▶ 5:23 unnamed speaker Like you started out with the Lama number and then you sort of branched out into all the others.
- ▶ 13:45 unnamed speaker Like, uh, you know, I'm looking at these, these charts of QN-III, LAMA, LAMA-IV, OSHGPT, and obviously there's more.
- ▶ 13:45 unnamed speaker Like, uh, you know, I'm looking at these, these charts of QN-III, LAMA, LAMA-IV, OSHGPT, and obviously there's more.
A Technical History of Generative Media
- ▶ 8:39 Gorkem Yurtseven And then obviously after stable diffusion, I think like four or five months later, LAMA-II came out and, um, there was a decision point again.
Better Data is All You Need — Ari Morcos, Datology
- ▶ 14:49 Alessio Fanelli And you were at Meta from 2018 to September 23, which is both during Lama one and Lama two. 3 times in the scene
- ▶ 14:49 Alessio Fanelli And you were at Meta from 2018 to September 23, which is both during Lama one and Lama two.
- ▶ 26:12 Shawn Wang So my conspiracy theory for what happened to Llama four is the lawyers got to it.
- ▶ 27:11 Ari Morcos With 4.5 and Lama four and others.
- ▶ 54:02 Ari Morcos Um, and I think we've also seen evidence for this, like looking at the difference between Lama and Quen with respect to their ability to be post-trained, right? 2 times in the scene
- ▶ 1:03:52 Ari Morcos Like you look at just like the Lama series, you know, if you want to exclude Lama four, do so. 2 times in the scene
- ▶ 1:03:52 Ari Morcos Like you look at just like the Lama series, you know, if you want to exclude Lama four, do so.
The AI Agenda: GPT5 leaks and the business of AI News — Steph Palazzolo, The Information
- ▶ 13:38 Shawn Wang So, so far, Deep Seek, obviously, like, still one of the biggest news of the, the year, what was a big gift to, to, uh, to, to the, the inference providers and the sort of relative decline of Lama and the disappointment Lama four was, uh,…
- ▶ 13:38 Shawn Wang So, so far, Deep Seek, obviously, like, still one of the biggest news of the, the year, what was a big gift to, to, uh, to, to the, the inference providers and the sort of relative decline of Lama and the disappointment Lama four was, uh,…
- ▶ 17:34 Stephanie Palazzolo So I'm like, I feel like we've seen progress in models, and even for Meta, like, maybe the, I mean, as we saw with Llama IV, even the progress wasn't, like, amazing, it seems.
The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
- ▶ 2:23 Nathan Lambert Suite of models from, I think, eight, seven D and four or five B is based on llama at the time.
- ▶ 2:31 Nathan Lambert I think meta has different priorities and their things for llama 3.1, which is a great set of models at the time. 2 times in the scene
- ▶ 1:13:16 Shawn Wang In April, uh, you said LlamaFor, did Meta just push the panic button?
🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
- ▶ 52:53 Scott Wu I'm going to ask Devin to benchmark the performance of Llama and a couple of different API providers.
- ▶ 2:01:23 Varun Mohan I think Lama four, depending on where it goes, it could be materially better.
⚡️Math Olympiad gold medalist explains OpenAI and Google DeepMind IMO Gold Performances
- ▶ 30:57 Dr. Jasper Zhang Like, uh, if you look at, if you want to run like Lama three and, and like now compared to now, it's like three X, two X better.
⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
- ▶ 10:08 Pratik Bhavsar Another, of course, the other surprise was that all the Lama models were not performing well on our benchmark.
- ▶ 10:14 Pratik Bhavsar Uh, 3.3 and even the Lama four all were really performing extremely poor.
- ▶ 10:14 Pratik Bhavsar Uh, 3.3 and even the Lama four all were really performing extremely poor.
- ▶ 10:39 Pratik Bhavsar The Lama one was great.
- ▶ 10:40 Pratik Bhavsar Lama two was great.
- ▶ 10:41 Pratik Bhavsar Three also was great.
Information Theory for Language Models: Jack Morris
- ▶ 40:57 Shawn Wang Uh, so Gemma three N is like a really good candidate right now because it's like a four B model that is like claimed to be better than Lama four and GPT 4.1, uh, according to, you know, certain arenas that shall not be named.
- ▶ 57:36 Jack Morris I think, like, maybe even if we tested this with LALAMA architecture, like, there's sort of like a GPT++ architecture, like, I would guess that can store better data just because the kind of numerical flow is a little bit better, the…
- ▶ 1:05:41 Shawn Wang Lama does it frequently.
The Shape of Compute (Chris Lattner of Modular)
- ▶ 2:23 Chris Lattner And so we need to be state of the art on NVIDIA GPUs meeting and beating NVIDIA's best on things like a Lana three model, which by the way is serving end to end, like very high bar, by the way, this is like
The Utility of Interpretability — Emmanuel Amiesen
- ▶ 2:37 Vibhu (Viboo) What, why should we probe Gemma, Lama? 2 times in the scene
- ▶ 43:10 unnamed speaker He did Neuronpedia and released a bunch of SAEs for, I think, the Llama models and the Gemma models.
- ▶ 48:18 Vibhu (Viboo) Like, we'll train an SAE on one layer of LAMA and probe around, but then people are like, okay, how does this have much impact?
- ▶ 1:27:30 Emmanuel Ameisen There's some of the Lama models.
The Rise and Fall of the Vector DB category: Jo Kristian Bergum (ex-Chief Scientist, Vespa)
- ▶ 19:54 unnamed speaker I was like, um, you know, this is, Lama four is going to reignite the long context versus fact debate, but it will actually resolve the debate, but not in the way that you want.
Claude Plays Pokémon Hackathon: Escape from Mt. Moon!
- ▶ 48:31 Hackathon Organizer I don't know, maybe Llama plays Pokemon today, right?
The Agent Network — Dharmesh Shah, Agent.ai + CTO of HubSpot
- ▶ 1:34:39 Shawn Wang Uh, most of Mustafa, who was part of, so they had image generation in Lama three.
The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
- ▶ 3:54 Shawn Wang Like, the, the Lama four or five Bs, the, uh,
- ▶ 14:41 Rishabh Agarwal Like, that's the cool thing about, that's why they were able to distill from DeepSeq model to Lama or Quen, and you don't have to even think about that tokenizer, and that's why this is common, right?
- ▶ 17:12 Rishabh Agarwal And for a fixed number of tokens, if you have a much smaller model, so let's say lama-seventyb versus lama-seventyb, you can generate 10 times more data in principle, right? 2 times in the scene
npm install Agents — with Sunil Pai and Rita Kozlov (VP AI) of Cloudflare
- ▶ 29:28 Rita Kozlov Basically LamaGuard as a part of our AI gateway.
Outlasting Noam Shazeer, Crowdsourcing Chai AI w/ 1.4m DAU — with William Beauchamp, Chai Research
- ▶ 25:50 William Beauchamp And so I'm very interested in what Lama four is going to look like and if they're able to sort of match what Deep Seek have been able to achieve with this performance per dollar gain.
- ▶ 56:30 William Beauchamp What the money let them do was, if they wanted to fine-tune Alarma-seventyb on eight H-one-hundreds overnight, if you give them money, then they can do it.
- ▶ 1:03:09 William Beauchamp Right, and then you can give that to, like, a Llama-seventyb. 2 times in the scene
The Unreasonable Effectiveness of Reasoning Distillation: using DeepSeek R1 to beat OpenAI o1
- ▶ 6:56 unnamed speaker So part of it is also that the, the open models, like, I think maybe before the discussion was like, is the data that you get from LLAMA three, four or five B that much worse than like four.
DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
- ▶ 4:39 Yining Zhang I think at the base time, something like LAMA-Seventy-B is more common. 3 times in the scene
- ▶ 4:45 Yining Zhang I think LAMA-Seventy-B has released the 400 zero five billion weights, but I think there are just a few users use that. 3 times in the scene
- ▶ 8:28 unnamed speaker So I think a lot of companies as well, like together, they'll also release like quantized versions of the Lama models, right?
- ▶ 10:08 unnamed speaker And so when it comes to the quantization question, uh, where it matters is that we would never quantize the model behind, you know, the user's back and say, look at us, there's a faster Lama, uh, 70 B that has been, you know, uh, somehow…
- ▶ 12:52 Yining Zhang I, I think it's higher than every other open source AIM, even the LAMA 400 zero five billion.
- ▶ 14:14 unnamed speaker Like, I think, well, Lama four or five B is dense.
- ▶ 14:53 Yining Zhang So the reason why Lama, uh, open-sourced, uh, the MOE model, because I, I think they, they tried to train our MOE model, but they failed. 2 times in the scene
Beating Google at Search with Neural PageRank and $5M of H200s — with Will Bryk of Exa.ai
- ▶ 46:34 unnamed speaker So obviously they're, they're spending a lot of time in it, but then you have maybe the GPU poor, which are still working on making llama good.
The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024]
- ▶ 12:38 Nathan Lambert An example that I used in the blog post I wrote today on this is like Lama, 3.1 details their vows for math.
2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
- ▶ 14:58 Shawn Wang So when Lama 3.3 launched, they only launched the 70 B because they use four or five B to distill the 70 B.
- ▶ 16:18 Shawn Wang Um, I mean, Lama four will be reasoning oriented. 2 times in the scene
- ▶ 25:23 Alessio Fanelli It's like, oh, you can fine tune llama.
- ▶ 28:38 Shawn Wang Zero capabilities, and it's sudden emergence of GPT-IV.
- ▶ 30:39 Shawn Wang I, I feel like, um, like, like you didn't last year, you had people like Hau Tien who worked on Lava, uh, which is take Lama and add vision.
- ▶ 1:27:22 Shawn Wang Um, we also had Lama III released, so, you know, I think people always want to 2 times in the scene
- ▶ 1:33:13 Alessio Fanelli And then July was Lama three.
Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
- ▶ 14:53 Graham Neubig This is old, um, and we need to update this, basically, but, um, we evaluated Claude, GPT-FORO, O-ONE-MINI, um,…
Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
- ▶ 11:49 Loubna Ben Allal Uh, for example, here we released the dataset called FineWebEDU, and the way we built it is by taking LAMA-III and asking it to rate the educational content of web pages
- ▶ 20:25 Loubna Ben Allal For example, if you compare how much, how long LAMA was trained compared to LAMA-III, there is a huge increase in the pre-training length. 2 times in the scene
- ▶ 20:25 Loubna Ben Allal For example, if you compare how much, how long LAMA was trained compared to LAMA-III, there is a huge increase in the pre-training length. 2 times in the scene
- ▶ 22:36 Loubna Ben Allal Uh, for example, our 1.7 B model outperforms Lama one B and also .2.
2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
page 1 of 3 · 100 scenes per page · newest episode first next →