Bornstein: Meta tailored Llama license thresholds to exclude only two companies
“I do remember that the numbers were like specifically chosen at that time that you could go find it was like two companies in the world. That like fit the definition that they had excluded from their license.”
Bubna: Ramp trained custom tokenizers to swap into LLaMA
“Ramp actually early in the day was training their own tokenizer and, like, Swapping out the tokenizer in Lama and whatnot.”
Chamath: Meta fumbled its open-source AI strategy with Llama
“Meta really fumbled this with Llama.”
Ethan He: NVIDIA Cosmos uses a 7B video model with a larger LLM rewriter
“I think in in Cosmos, we use Lama or we use mix, mix through. And the Cosmos video model itself is only seven B, and the model, the language model is a prompt rewriter. It's bigger than that.”
O'Driscoll: Meta's pivot to closed-source AI harms the ecosystem
“Also worth noting that they're talking about being much more closed source, which is a significant thing, because at some point someone's going to need the American version of open source, and Lama was that, and now they're pivoting more to be more closed sour…”
Rory O’Driscoll Apr 16, 2026 ▶ 43:39 SpaceX's Financials Leaked: Is it Worth $2TN | Meta Debuts Muse Spark: Are They Back in the AI Race? · 20VC with Harry Stebbings
Lample: Post-training breakthroughs like DPO were impossible without open LLaMA
“And if you look at many of the techniques that were developed after, for instance, Temma was open source, like all these post-training approaches like even DPOD, like performance optimization, all of this were done by people that had access to this model, and …”
Chamath: BitTensor Subnet 3 trained a distributed 4B parameter Llama model
“Two days ago, you may not have seen this because you were busy on stage, but there was a training run that happened in this crypto project called BitTensor. Subnet three, they managed to train a four billion parameter llama model, totally distributed with a bu…”
DeepSeek 128k context fits in 8GB KV cache versus Llama's 80GB
“For context like the total I think the total context length of DeepSeq is a 128,000 tokens, or it might be 256,000 with rope extension. That entire context, I think it's a 128,000, fits into eight gigabytes. And previously context, like I think the Lama four o…”
Cannon-Brookes: Anthropic is one of Atlassian's largest AI models
“We use Anthropic a lot. It's one of our biggest models. We use, you know, multiple models which I think is what most good SaaS vendors are doing within their customers' choices, right? Like we use a lot of Gemini, a lot of Anthropic. We have a whole bunch of L…”
Goodfire AI: We replicated code error and malicious features in Llama
“We replicated a lot of these features in, in our llama models as well.”
Soldaini: Most open AI models are open weights, not open source
“Majority of models that get release I think the best term to describe them is open weights. Your Quinn, your Gemma, your Lama you know, Kimi it's what gets release is a set of weights that correspond either to the final state of model, that's the most common, …”
Meta AI's User Numbers Are Inflated by Accidental Instagram Searches
“That was complete garbage. That was people searching Instagram who accidentally hit a llama model that made some things happen, and they were like, oh, go away. I actually am just looking for a user.”
Lambert: Meta withholding its leading benchmark model is bad execution
“But to be a model that claims to be open and then not release the model that is your leading claim is just, like, that is, like, bad execution.”
Kantrowitz: Meta failed to convincingly build voice and avatars with Llama
“Meta has been trying extremely hard to build voice and avatars with Lama, and it hasn't been able to do it convincingly.”
Meta Llama 3.3 and Llama 4 perform poorly on agent benchmarks
“Another, of course, the other surprise was that all the Lama models were not performing well on our benchmark. 3.3 and even the Lama four all were really performing extremely poor.”
Depue: Yann LeCun had no involvement with Meta's LLaMA models
“I think it's a shame that they didn't necessarily have someone in the llama org who was as visible as him because he was not involved with llama at all. Like fair fundamental AI research is like a whole other group that they mostly write papers and do very aca…”
Depue: Meta may eventually shift to releasing closed flagship AI models
“I w it wouldn't surprise me if they go kind of the Google route where like they are still, they still do open source. They still have the llama brand, but it isn't like their flagship thing where they start doing, whether it's for internal products or it's, th…”
Patel: Meta treats Llama as a toy rather than approaching AGI correctly
“I think they're treating it as like a sort of like toy within the meta universe. And I don't think that's the correct way to think about AGI.”
Kantrowitz: Attention on Meta's Llama has dropped amid DeepSeek's surge
“As deep sea continues to surge, the conversation around Meta's llama models has certainly fallen off in a big way.”
D'Sa: Ultra-fast voice AI latency yields diminishing returns
“There's also kind of diminishing returns after a while. To give you an example, I once built this Cerebris demo. I used like a Lama seven B or eight B Lama eight B have to remember these numbers on the primary accounts, but Lama eight B hooked up to Cerebris. …”
Patel: Meta's next Llama model will match DeepSeek-V3's cost efficiency
“And Meta's Meta is going to release their new llama soon enough. Right. And that one is going to be, you know, a similar level of cost decrease probably similar areas, deep seek V three.”
Gurley: Meta restricts free Llama usage for apps over 700M users
“Some of the players, most notably meta with Lama, have a usage restriction against the free use of the model at seven hundred million, and that's what you were referring to, and so at least in a tweet Sam suggested they won't have that in theirs.”
Agarwal: Synthetic data distillation bypasses model vocabulary and tokenizer mismatches
“The one nice thing about this kind of distillation is it doesn't matter if you have a vocabulary mismatch, because we're not using the next token distribution or probability labels. You can distill from one model which uses some random tokenizer to another mod…”
Mascorro: Distillations from DeepSeek-R1 Outperformed Direct RL on Smaller Models
“So it turns out in their experiments, they took Lama's EV and some of these are QN models, and they basically apply RL straight the same way they did it with R one on these base models. And it turns out that it improved in some fields, but it was not a signifi…”
Krieger: Distillation Is Unnecessary for Frontier Open-Source AI Progress
“I think the open source models, Like take Llama, for example, like they've been able to do that from their own research and perspective and data ingestion and training. And so I guess I would say distillation does not feel essential in order to unlock those th…”
Mullenweg: Meta's Llama isn't truly open source due to user restrictions
“However, there's a clause in it that says, if you're above a certain threshold of monthly active users, I forget what it is, like, it's big, it's like seven hundred fifty million, so it's pretty high. You need a license from them. And so that does not give you…”
Matt Mullenweg Mar 2, 2025 ▶ 23:57 Matt Mullenweg on the future of open source and why he’s taking a stand
Gerstner: DeepSeek has displaced Meta's Llama in enterprise AI experimentation
“I heard from several inference players that you and I are friends with that all of a sudden DeepSeek rather than Lama is the enterprise open source model of choice that everybody's experimenting with and playing with.”
Morin: DeepSeek validates Meta's open-source strategy by feeding advances back
“This DeepSeek thing is exactly what we were talking about, that by releasing Llama, you know, DeepSeek surely used it, and many others, sounds like, and the fruits of that win come back to Meta, like, overnight, right? And they can plug it back in and make mor…”
Enterprise AI ROI requires targeted micro-models arbitrating between large foundational models
“Can you now devolve or basically create a micro model that's very targeted to your business? That basically was a recipe. So how do I look at what's happening within, you know, within a CHATGP versus Gemini? How do I start looking at Lama differently? And that…”
Coogan: Western AI Labs Will Port DeepSeek's Optimizations Into Their Models
“It's open source. The paper's out there. So that will get ported back to Llama, to OpenAI, to Anthropic, but I haven't seen that.”
Coogan: DeepSeek Open Sourcing Leading-Edge Model Will Disrupt AI Competition
“Zuck has been running this playbook with Llama open sourcing the lagging edge AI model, and it's had a strong effect on the competition. And so deep seek open sourcing, the leading edge model is going to have a similar effect.”
Coogan: Meta's Llama is the most-used LLM due to Instagram integration
“They're, because Llama is the number one LLM, not because people are hosting it independently. Although they are, plenty of companies are hosting Llama independently. They're the number one used LLM Because it's tucked into Instagram, and so they just have a t…”
Zhang: Meta Failed at Training MoE Models for Llama Series
“The reason why Lama open-sourced the MOE model, because I think they tried to train our MOE model, but they failed. So that, that's why they didn't open source MOE mode for Lama series.”
Kuyda: AI product startups should fine-tune open models, not build foundation models
“If you can fine tune a llama based model and focus on the logic, the product, the application layer on what you are, what value you're actually providing to the user versus training your own model that becomes obsolete in three months and requires incredible, …”
Gerstner: Meta charges hyperscalers licensing fees for Llama models
“Where he said they already charge license fees to the large hyperscalers to use Lama and to provide it to them in certain ways, etc. I think it's de minimis revenue in the scheme of Meta, but, you know, he's smart enough in that. Bill Gurley: I don't know. Bra…”
Levie: Companies will blow billions training models instantly supplanted by Meta's Llama
“There'll be lots of companies that just blow billions of dollars doing, you know, training runs of a model that get, you know, supplanted instantly by Lama.”
Ben Allal: LLaMA 3 used 15x more pre-training tokens than original LLaMA
“LAMA was trained on one trillion tokens, but LAMA-III was trained on 15 trillion tokens.”
Ben Allal: Hugging Face SmolLM2-1.7B outperforms Llama 3.2 models
“So it's a series of three models, which are the best in class in each model size. For example, our 1.7 B model outperforms Lama one B and also .2.”
Soldani: Llama and Qwen models fail OSI open source AI definition
“Under this definition, for example, Lama or some of the Quen models are not open source because the license says you can, you can't use this model for this, or it says if you use this model, you have to name the output this way or derivative needs to be named …”
Cerebras WSE-3 runs Llama inference 70x faster than NVIDIA GPUs
“Cerebris came out that the wafer scale engine three can serve llama 70 B at 2.1 thousand sorry, 202,100 tokens per second and serves llama four or five B at nearly 1000 tokens per second. So this, you know, to give you an understanding, like this is about 70 t…”
Fine-Tuning Open Source Models Lacks Levers of Full Vertical Training
“Taking those models and trying to fine tune them It's just, it's not as effective as building it yourself and you have much fewer levers to pull than if you actually have access to the data and you can change the data that goes into that process.”
Angelopoulos: The Chatbot Arena leaderboard is currently not an apples-to-apples comparison
“None of the leaderboard currently is apples to apples, because you have, like, Gemini Flash, you have, you know, all sorts of tiny models, like Llama Like, eight B and four or five B are not apples to apples.”
Meta Open-Sources AI to Commoditize Complements and Lower Long-Term Costs
“When you think about it, it's actually a form of operating leverage, where he's basically saying there's a big fixed cost I am willing to bear in order to bootstrap this ecosystem and commoditize all of these complements, commoditize all of these other closed …”
Drew Houston brings an external GPU on planes to run Llama locally
“When I'm on a plane or something or where, like, you don't have access or the Internet's not reliable, I actually bring a gaming laptop on the plane with me. It's, like, a little, like, blue briefcase-looking thing, and then I, like, literally hook up a GPU, l…”
Bosworth Strongly Advocated Internally For Meta To Open-Source LLaMA
“I was one of the loudest voices internally to channel, encouraging us to open source Llama one, Llama two.”
Bret Taylor: Companies needing standard models should simply fine-tune Llama or Mistral
“You know, I think that in that market probably if you need a model like that, you should download llama. That's the answer. It's like, you don't need much of a cheat sheet on that, you know, and or maybe Mr. All, but pick one of the open source models that are…”
Schmidt: ChatGPT, Llama, and DeepSeek use Nous Research's YaRN context extension
“Bone here is the lead author of a method we developed called YARN, which is a context window extension method that we released and did the research on. It is now used by every, every model you use nowadays, everything, everything Chachipiti, Lama, DeepSeq, all…”
Schmidt: Fewer than ten organizations worldwide can train Llama-scale AI models
“Yeah, I mean, I would, it would probably be in the number of ones on my hand and it probably wouldn't use all my fingers, you know. Yeah, I mean, you basically have, OpenA, Anthropic, Meta, X, Google, and then you have a few Mistral, and then Deep Seek and a c…”
Frosst: Meta's Llama is a corporate giveaway for notoriety, not open source
“Lama, Lama three, like the stuff Meta is doing great models, great engineers, really cool stuff. But it's not open source the way Linux is open source. It's not open source the way, you know, I don't know, like Wikipedia's contributions are open source in that…”
Karpathy: RoPE is the only major transformer architecture change in five years
“The transformer hasn't changed that much. You know, we've added the rope positional and the rope relative positional encodings. That's like the major change. Everything else doesn't really matter too much. It's like plus three percent on a small few things. Bu…”
Mollick: Silicon Valley founders build narrow wrappers despite predicting AGI
“It is very strange from one hand for all of these people in Silicon Valley to be like, yeah, you know, AGI is coming. And then the applications they're building are like these very narrow, like, hey, I slapped something on top of llama. And you know, it's like…”
Sam Lessin: Meta's Lack of a Simple Llama API Is a Strategic Mistake
“And like, the thing about LOM or anything is like, there just doesn't exist. Like someone will package that eventually, but like, they don't prioritize that. I actually think it's a strategic mistake. I think they should, purely for developer mindshare, not as…”
Sam Lessin: Llama Is Positioned Perfectly for Enterprise IT Consultancies
“Well, I think Llama's well set up for this, because then they can say, well, you're going to want, you don't want to use the off the shelf stuff. You want to use a custom instance.”
Yi Tay: Meta's Llama is corporate open weights, not grassroots open source
“To me, Lama Tree is like... Meta has an org that is hypothetically very similar to Gemini or something but they just decide to release the weights It's open weights It's open weights and everything”
Wang: Most serious enterprises will adopt on-premise AI models
“And this is actually why I think there's a very there's a very big sort of opportunity for whether it's open source models or the llama models or the mistral models or whatnot, basically these models that can go on prem and that enterprises can take. And then …”
Conover: Commercial LLMs struggle to generate 5,000 output tokens in one generation
“There is a characteristic output length for these models. Let's say it's about 1200 tokens. Like it is very difficult to get any of the commercial LMs or LLAMA to write 5000 tokens.”
Schroepfer: Meta's early AI lab built the skills that powered Llama
“We built an AI lab at Meta and it was kind of like, huh, this AI thing is going to be pretty big. Let's like start building the skill set. And that's why you see Meta doing Lama now and a bunch of other things, because they, we had all the people, we had all t…”
Horowitz: Only researchers can tell top AI models apart
“I think if you look at the very top models you know, Claude and OpenAI and Mistral and Lama The only people who I feel like really can tell the difference as users amongst those models are the people who study them. You know, like they're getting pretty close.”
Tavel: Frontier AI models will remain closed source unless Meta shifts market
“If you want a model that's on the frontier, that's going to be closed source. Now, look, every, as I mentioned, every month it feels different with what Meta's doing with Llama, that may actually fundamentally change the game. If Meta is willing to make that h…”
Gerstner: Open Source Will Rapidly Commoditize Non-Frontier AI Models
“Apart from the people who are on the very frontier, if you're on the frontier and you have something totally different, it seems to me that that's a place where that is defensible. But if you're not on the frontier, man, it seems that these are going to be rea…”