Llama, every mention

279 scenes, the whole family · ← back to Llama

tap a year for its mentions
0015025300502023202420252026episodesmentions
025502023202420252026episodes it came up in
004258502023202420252026episodesmentions per episode

every year anyone Shawn Wang 63Alessio Fanelli 35Thomas Scialom 22Nathan Lambert 21Soumith Chintala 15George Hotz 13Yining Zhang 9Lin Qiao 7Pratik Bhavsar 6Mark Huang 6

Verbatim, from the transcripts: the passages where Llama comes up

loading…

Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten Aug 3, 2026 · 4 mentions

The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI Jul 22, 2026 · 3 mentions

  • ▶ 15:57 Eiso Kant And at the time, if I recall, you look at like the early Lama papers and things like that, people were juicing Epsilon like quite a bit. 2 times in the scene
  • ▶ 53:29 Eiso Kant So this wasn't even, there was only, I think, Lama out at the time and that's it.

The AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO Jul 13, 2026 · 1 mention

  • ▶ 22:39 Dan Biderman Um, so the examples we like to give is that, uh, if you take a Lama, a 70 B model, and you load one article from Wikipedia, which is a few tens of kilobytes, and you have the model read this, uh,

The Future of AI Infra: from Kubernetes to Agent Sandboxes — Akshat Bubna, Modal CTO Jul 8, 2026 · 1 mention

🔬 "The Most Innovative Diffusion Research Is Happening in Drug Discovery, Not Image Generation" Jun 30, 2026 · 6 mentions

  • ▶ 0:55 RJ Haneke I'm joined by my co-host Brandon Anderson, and we're privileged to have Evan Feinberg, founder and CEO of Genesis Molecular AI, and Sergei Yudinov, who led Lama II and Lama III pre-training before he joined Genesis as CTO. 2 times in the scene
  • ▶ 0:55 RJ Haneke I'm joined by my co-host Brandon Anderson, and we're privileged to have Evan Feinberg, founder and CEO of Genesis Molecular AI, and Sergei Yudinov, who led Lama II and Lama III pre-training before he joined Genesis as CTO. 2 times in the scene
  • ▶ 1:42 Sergei Yudinov I later on led LAMA team, um, LAMA two and LAMA three models.
  • ▶ 20:36 Evan Feinberg Sergei is being humble, but Sergei led the LLAMA II research team at, at, at Meta when, when, when he was still there.

The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin Jun 24, 2026 · 1 mention

  • ▶ 1:00:03 Matei Zaharia So we, uh, decided, you know, even though we, we did launch, uh, open source model DBRX, and, you know, we, we went up to, like, sort of above the LAMA-R III scale, we decided that we really want to focus on, there'll be so many people…

AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan Jun 22, 2026 · 2 mentions

  • ▶ 27:22 unnamed speaker But it'd be nice if I just had like a llama guard or the, whatever the.
  • ▶ 34:08 unnamed speaker Release their own, like, you know, Llama has one, OpenAI has one, Google has one.

Why AI Labs With Unlimited GPUs Still Fail — Anjney Midha, AMP Jun 18, 2026 · 1 mention

  • ▶ 37:25 Anjney Midha Like, you, you are an athlete of the mind, and you perform at the highest levels, and to get there, whether you're, you know, Anastasius or Waylon at Berkeley, or you are Robin, who, with Black Forest and created Stable Diffusion, or if…

Satya Nadella on AI: @NoPriorsPodcast x Latent Space Crossover Special at Microsoft Build 2026 Jun 3, 2026 · 1 mention

  • ▶ 10:14 Satya Nadella Like you can use your Lama harness, whatever, or you can use the, um, uh, you know, any open harness or any harness of yours and train with your tools and multiple models and your context.

Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He Jun 1, 2026 · 1 mention

⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind May 24, 2026 · 1 mention

Mistral: Voxtral TTS, Forge, Leanstral, & Mistral 4 — w/ Pavan Kumar Reddy & Guillaume Lample Mar 30, 2026 · 2 mentions

  • ▶ 38:05 Guillaume Lample So, me and Tim were at Meta, we released Lama, and I think what was really nice to see that before this, for most researchers, like universities, it was impossible to, to work on LLNs. 2 times in the scene

Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup Mar 8, 2026 · 2 mentions

  • ▶ 33:30 unnamed speaker We're not the best at training MOEs when they're pre-trained, like we saw this with LAMA-III, right?
  • ▶ 52:21 Kyle Kranen And previously context, like I think the, the Lama four or five B context of a similar size was like 40 or 80 gigabytes in the same precision.

Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell Feb 5, 2026 · 1 mention

  • ▶ 46:01 Myra Deng We replicated a lot of these features in, in our llama models as well.

Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith Jan 9, 2026 · 1 mention

  • ▶ 56:55 Shawn Wang And so Lama had this, like, if you have seven hundred million daily active users, you're not allowed to use our model or you have to talk to us, something like that.

[State of AI Papers 2025] Fixing Research with Social Signals, OCR & Implementation — Team AlphaXiv Dec 31, 2025 · 1 mention

  • ▶ 3:45 unnamed speaker I think when we started, we were very deliberate about getting authors, like, from LoRa, DPO, Lama, and actually having really useful, cool exchanges.

SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow) Dec 18, 2025 · 7 mentions

  • ▶ 26:56 Nikhila Ravi Um, we also use Llama in our data engine. 2 times in the scene
  • ▶ 31:03 unnamed speaker Interestingly, you use Lama four. 2 times in the scene
  • ▶ 31:05 unnamed speaker I saw there's a mix of Lama three and Lama four here, but it looks like it does best with Gemini 2.5, which makes sense given this comparable set of MLMs.
  • ▶ 39:01 Pengchuan Zhang Source images and it generates kind of non-phases from, for example, kind of NAMA generate caption and we pass the caption to get the non-phases.
  • ▶ 41:14 Pengchuan Zhang That is a breakthrough, and then, kind of, we kind of fine-tune our, kind of, for example, NAMA's 3.2 with our, kind of, verification, human annotated verification data.

The Future of Email: Superhuman CTO on Your Inbox As the Real AI Agent (Not ChatGPT) — Loïc Houssier Dec 11, 2025 · 1 mention

  • ▶ 30:12 Loïc Houssier Uh, but, uh, we use Base-Ten to run some, uh, I would say some LAMA, some BERT model for classification.

⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF Oct 20, 2025 · 2 mentions

  • ▶ 8:54 Elie Bakouch And for example, a good, uh, a good way to view that is that, uh, DeepSeq rig three is still using the same Adam parameter than, uh, Lama two.
  • ▶ 1:00:25 Alessio Fanelli Like, uh, I have this, like a MCP client, I built, and we use Lama, um, AB for like a conversation naming, you know, I feel like that's like a great use case for like a five hundred million parameter model, but there's no simple API to use…

Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave) Oct 16, 2025 · 1 mention

  • ▶ 8:09 Kyle Corbitt Yeah, they were really strong models, um, better than the Llama two that they were, you know, effectively replacing.

Building Jamba 3B: the tiny Hybrid Transformer State Space Reasoning Model - Barak Lenz, CTO of AI21 Oct 11, 2025 · 2 mentions

  • ▶ 6:09 Barak Lenz Straight off the bat, compare themselves to, to latest, uh, attention architecture introduced by Lama that, that introduced a lot of corrections that cause things to work. 2 times in the scene

⚡️Raising $1.1b to build the fastest LLM Chips on Earth — Andrew Feldman, Cerebras Oct 1, 2025 · 3 mentions

  • ▶ 5:23 unnamed speaker Like you started out with the Lama number and then you sort of branched out into all the others.
  • ▶ 13:45 unnamed speaker Like, uh, you know, I'm looking at these, these charts of QN-III, LAMA, LAMA-IV, OSHGPT, and obviously there's more.
  • ▶ 13:45 unnamed speaker Like, uh, you know, I'm looking at these, these charts of QN-III, LAMA, LAMA-IV, OSHGPT, and obviously there's more.

A Technical History of Generative Media Sep 8, 2025 · 1 mention

  • ▶ 8:39 Gorkem Yurtseven And then obviously after stable diffusion, I think like four or five months later, LAMA-II came out and, um, there was a decision point again.

Better Data is All You Need — Ari Morcos, Datology Aug 29, 2025 · 11 mentions

  • ▶ 14:49 Alessio Fanelli And you were at Meta from 2018 to September 23, which is both during Lama one and Lama two. 3 times in the scene
  • ▶ 14:49 Alessio Fanelli And you were at Meta from 2018 to September 23, which is both during Lama one and Lama two.
  • ▶ 26:12 Shawn Wang So my conspiracy theory for what happened to Llama four is the lawyers got to it.
  • ▶ 27:11 Ari Morcos With 4.5 and Lama four and others.
  • ▶ 54:02 Ari Morcos Um, and I think we've also seen evidence for this, like looking at the difference between Lama and Quen with respect to their ability to be post-trained, right? 2 times in the scene
  • ▶ 1:03:52 Ari Morcos Like you look at just like the Lama series, you know, if you want to exclude Lama four, do so. 2 times in the scene
  • ▶ 1:03:52 Ari Morcos Like you look at just like the Lama series, you know, if you want to exclude Lama four, do so.

The AI Agenda: GPT5 leaks and the business of AI News — Steph Palazzolo, The Information Aug 6, 2025 · 3 mentions

  • ▶ 13:38 Shawn Wang So, so far, Deep Seek, obviously, like, still one of the biggest news of the, the year, what was a big gift to, to, uh, to, to the, the inference providers and the sort of relative decline of Lama and the disappointment Lama four was, uh,…
  • ▶ 13:38 Shawn Wang So, so far, Deep Seek, obviously, like, still one of the biggest news of the, the year, what was a big gift to, to, uh, to, to the, the inference providers and the sort of relative decline of Lama and the disappointment Lama four was, uh,…
  • ▶ 17:34 Stephanie Palazzolo So I'm like, I feel like we've seen progress in models, and even for Meta, like, maybe the, I mean, as we saw with Llama IV, even the progress wasn't, like, amazing, it seems.

The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai) Jul 31, 2025 · 4 mentions

  • ▶ 2:23 Nathan Lambert Suite of models from, I think, eight, seven D and four or five B is based on llama at the time.
  • ▶ 2:31 Nathan Lambert I think meta has different priorities and their things for llama 3.1, which is a great set of models at the time. 2 times in the scene
  • ▶ 1:13:16 Shawn Wang In April, uh, you said LlamaFor, did Meta just push the panic button?

🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R) Jul 28, 2025 · 2 mentions

  • ▶ 52:53 Scott Wu I'm going to ask Devin to benchmark the performance of Llama and a couple of different API providers.
  • ▶ 2:01:23 Varun Mohan I think Lama four, depending on where it goes, it could be materially better.

⚡️Math Olympiad gold medalist explains OpenAI and Google DeepMind IMO Gold Performances Jul 24, 2025 · 1 mention

  • ▶ 30:57 Dr. Jasper Zhang Like, uh, if you look at, if you want to run like Lama three and, and like now compared to now, it's like three X, two X better.

⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo Jul 14, 2025 · 6 mentions

Information Theory for Language Models: Jack Morris Jul 2, 2025 · 3 mentions

  • ▶ 40:57 Shawn Wang Uh, so Gemma three N is like a really good candidate right now because it's like a four B model that is like claimed to be better than Lama four and GPT 4.1, uh, according to, you know, certain arenas that shall not be named.
  • ▶ 57:36 Jack Morris I think, like, maybe even if we tested this with LALAMA architecture, like, there's sort of like a GPT++ architecture, like, I would guess that can store better data just because the kind of numerical flow is a little bit better, the…
  • ▶ 1:05:41 Shawn Wang Lama does it frequently.

The Shape of Compute (Chris Lattner of Modular) Jun 13, 2025 · 1 mention

  • ▶ 2:23 Chris Lattner And so we need to be state of the art on NVIDIA GPUs meeting and beating NVIDIA's best on things like a Lana three model, which by the way is serving end to end, like very high bar, by the way, this is like

The Utility of Interpretability — Emmanuel Amiesen Jun 6, 2025 · 5 mentions

  • ▶ 2:37 Vibhu (Viboo) What, why should we probe Gemma, Lama? 2 times in the scene
  • ▶ 43:10 unnamed speaker He did Neuronpedia and released a bunch of SAEs for, I think, the Llama models and the Gemma models.
  • ▶ 48:18 Vibhu (Viboo) Like, we'll train an SAE on one layer of LAMA and probe around, but then people are like, okay, how does this have much impact?
  • ▶ 1:27:30 Emmanuel Ameisen There's some of the Lama models.

The Rise and Fall of the Vector DB category: Jo Kristian Bergum (ex-Chief Scientist, Vespa) Apr 19, 2025 · 1 mention

  • ▶ 19:54 unnamed speaker I was like, um, you know, this is, Lama four is going to reignite the long context versus fact debate, but it will actually resolve the debate, but not in the way that you want.

Claude Plays Pokémon Hackathon: Escape from Mt. Moon! Apr 5, 2025 · 1 mention

The Agent Network — Dharmesh Shah, Agent.ai + CTO of HubSpot Mar 28, 2025 · 1 mention

  • ▶ 1:34:39 Shawn Wang Uh, most of Mustafa, who was part of, so they had image generation in Lama three.

The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind Mar 23, 2025 · 4 mentions

  • ▶ 3:54 Shawn Wang Like, the, the Lama four or five Bs, the, uh,
  • ▶ 14:41 Rishabh Agarwal Like, that's the cool thing about, that's why they were able to distill from DeepSeq model to Lama or Quen, and you don't have to even think about that tokenizer, and that's why this is common, right?
  • ▶ 17:12 Rishabh Agarwal And for a fixed number of tokens, if you have a much smaller model, so let's say lama-seventyb versus lama-seventyb, you can generate 10 times more data in principle, right? 2 times in the scene

npm install Agents — with Sunil Pai and Rita Kozlov (VP AI) of Cloudflare Mar 19, 2025 · 1 mention

Outlasting Noam Shazeer, Crowdsourcing Chai AI w/ 1.4m DAU — with William Beauchamp, Chai Research Jan 26, 2025 · 4 mentions

  • ▶ 25:50 William Beauchamp And so I'm very interested in what Lama four is going to look like and if they're able to sort of match what Deep Seek have been able to achieve with this performance per dollar gain.
  • ▶ 56:30 William Beauchamp What the money let them do was, if they wanted to fine-tune Alarma-seventyb on eight H-one-hundreds overnight, if you give them money, then they can do it.
  • ▶ 1:03:09 William Beauchamp Right, and then you can give that to, like, a Llama-seventyb. 2 times in the scene

The Unreasonable Effectiveness of Reasoning Distillation: using DeepSeek R1 to beat OpenAI o1 Jan 24, 2025 · 1 mention

  • ▶ 6:56 unnamed speaker So part of it is also that the, the open models, like, I think maybe before the discussion was like, is the data that you get from LLAMA three, four or five B that much worse than like four.

DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing) Jan 19, 2025 · 12 mentions

  • ▶ 4:39 Yining Zhang I think at the base time, something like LAMA-Seventy-B is more common. 3 times in the scene
  • ▶ 4:45 Yining Zhang I think LAMA-Seventy-B has released the 400 zero five billion weights, but I think there are just a few users use that. 3 times in the scene
  • ▶ 8:28 unnamed speaker So I think a lot of companies as well, like together, they'll also release like quantized versions of the Lama models, right?
  • ▶ 10:08 unnamed speaker And so when it comes to the quantization question, uh, where it matters is that we would never quantize the model behind, you know, the user's back and say, look at us, there's a faster Lama, uh, 70 B that has been, you know, uh, somehow…
  • ▶ 12:52 Yining Zhang I, I think it's higher than every other open source AIM, even the LAMA 400 zero five billion.
  • ▶ 14:14 unnamed speaker Like, I think, well, Lama four or five B is dense.
  • ▶ 14:53 Yining Zhang So the reason why Lama, uh, open-sourced, uh, the MOE model, because I, I think they, they tried to train our MOE model, but they failed. 2 times in the scene

Beating Google at Search with Neural PageRank and $5M of H200s — with Will Bryk of Exa.ai Jan 10, 2025 · 1 mention

  • ▶ 46:34 unnamed speaker So obviously they're, they're spending a lot of time in it, but then you have maybe the GPU poor, which are still working on making llama good.

The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024] Jan 2, 2025 · 1 mention

  • ▶ 12:38 Nathan Lambert An example that I used in the blog post I wrote today on this is like Lama, 3.1 details their vows for math.

2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents Jan 1, 2025 · 9 mentions

Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands) Dec 25, 2024 · 1 mention

  • ▶ 14:53 Graham Neubig This is old, um, and we need to update this, basically, but, um, we evaluated Claude, GPT-FORO, O-ONE-MINI, um,…

Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024] Dec 24, 2024 · 6 mentions

  • ▶ 11:49 Loubna Ben Allal Uh, for example, here we released the dataset called FineWebEDU, and the way we built it is by taking LAMA-III and asking it to rate the educational content of web pages
  • ▶ 20:25 Loubna Ben Allal For example, if you compare how much, how long LAMA was trained compared to LAMA-III, there is a huge increase in the pre-training length. 2 times in the scene
  • ▶ 20:25 Loubna Ben Allal For example, if you compare how much, how long LAMA was trained compared to LAMA-III, there is a huge increase in the pre-training length. 2 times in the scene
  • ▶ 22:36 Loubna Ben Allal Uh, for example, our 1.7 B model outperforms Lama one B and also .2.

2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024] Dec 24, 2024 · 1 mention

  • ▶ 4:34 Dan Fu So when I try to do something like upload a whole book to Gemini, what happens beyond the, or maybe not Gemini, because we don't necessarily know what architecture is, but let's say we upload it to Llama, what happens beyond the scenes,…
page 1 of 3 · 100 scenes per page · newest episode first next →
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.