language models

also referred to as: language model

54 statements across 32 episodes · 22 bullish · 13 bearish · 29 people on the record · first statement Oct 20, 2023 by Jeremy Howard · across every show →

Everything said about language models, oldest first

Oct 20, 2023 negative
Insight
Alignment tax happens because models are fine-tuned instead of continued pre-trained
“ULM fit is the wrong approach and that's why we're seeing a lot of these You know, so-called alignment tax, and this view of like, oh, a model can't both code and do other things. You know, I think it's actually because people are training them wrong.”
Jeremy Howard Oct 20, 2023 ▶ 46:32 The End of Finetuning — with Jeremy Howard of Fast.ai
Oct 20, 2023 positive
Insight
Howard: Accurate next-token prediction forces models to learn world models and causality
“I thought, okay, so if I do this at a much bigger scale, using all of Wikipedia, what would it need to be able to do to finish a sentence in Wikipedia effectively, to do it quite accurately, quite often? I thought, geez, it would actually have to know a lot ab…”
Jeremy Howard Oct 20, 2023 ▶ 12:37 The End of Finetuning — with Jeremy Howard of Fast.ai
Oct 20, 2023 positive
Insight
Howard: There is no fine-tuning, only continued pre-training
“To me, the right way to do this is to fine, fine-tune language models, is to actually throw away the idea of fine-tuning. There's no such thing. There's only continued pre-training.”
Jeremy Howard Oct 20, 2023 ▶ 45:27 The End of Finetuning — with Jeremy Howard of Fast.ai
Jan 11, 2024 positive
Insight
Lambert: A 100x smaller language model filters output better than RLHF
“You could use like a hundred times smaller language model and do much better at filtering than RLHF”
Nathan Lambert Jan 11, 2024 ▶ 45:20 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024
Assertion Supported
Lambert: Training LLM reward models on 0-to-10 ratings failed
“People tried that with language models, which is if you have a prompt and a completion and you just have someone rate it from zero to 10, could you then train a reward model on all of these completions and zero to 10 ratings and see if you could actually chang…”
Nathan Lambert Jan 11, 2024 ▶ 37:43 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Feb 28, 2024 positive
Insight
Firshman: CLIs are a more natural fit for LLMs than GUIs
“It's almost more natural to a CLI than it is in a graphical user interface, because it feels like there's back and forth with the computer. Yeah. Almost funnily like a language model. So I think there's some interesting intersection of like CLIs and language m…”
Ben Firshman Feb 28, 2024 ▶ 6:26 A Brief History of the Open Source AI Hacker - with Ben Firshman of Replicate
Apr 11, 2024 positive
Assertion Not checkable as stated
Byun: LLM self-reported uncertainty is reasonably well-calibrated in production
“We found it to be pretty calibrated. There varies on the model.”
Jungwon Byun Apr 11, 2024 ▶ 44:29 Supervise the Process of AI Research — with Jungwon Byun and Andreas Stuhlmüller of Elicit
Apr 24, 2024 positive
Insight
Liu: LLMs excel at identifying nodes and edges for knowledge graphs
“One of the things we found out about these language models is that not only can you define nodes, it's really good at figuring out what are nodes and what are edges.”
Jason Liu Apr 24, 2024 ▶ 24:32 High Agency Pydantic over VC Backed Frameworks — with Jason Liu of Instructor
Apr 24, 2024 bearish
Disclosure
Liu: Admitted being bearish on LLMs for four years before recent breakthroughs
“I mean, the biggest one really was the fact that, like, I think for just four years I was so bearish on language models. And just NLP in general, I was just like, ah, like, none of this really works. Like, why would I spend time focusing on this? I gotta go do…”
Jason Liu Apr 24, 2024 ▶ 6:47 High Agency Pydantic over VC Backed Frameworks — with Jason Liu of Instructor
Apr 27, 2024 bullish
Opinion
Haisfield: WebSim gives language models far greater expressive tools
“That's like one of the things that's magic about WebSim generally is that it gives language models much far greater tools for expression, right?”
Rob Haisfield Apr 27, 2024 ▶ 30:19 This World Does Not Exist — Joscha Bach, Karan Malhotra, Rob Haisfield (WorldSim, WebSim, Liquid AI)
Jun 11, 2024 bullish
Insight
Conover: Programming LLMs as zero-marginal-cost machine learning systems is underappreciated
“I think that it is underappreciated how powerful, there's the generative capabilities of language models, but there's also the ability to program them to function as arbitrary machine learning systems, basically for marginally zero cost.”
Mike Conover Jun 11, 2024 ▶ 38:17 How AI is Eating Finance - with Mike Conover of Brightwave
Jun 21, 2024 negative
Assertion Not checkable as stated
Brady: Elicit routinely sees 10x p90 latency variation when prompting LLMs
“We do often normally, in fact, see a 10 X variation in P-ninety latency over the course of half an hour. When we're prompting these models, which is way higher than if you're working with a, you know, a more, more kind of conventional conventionally backed API…”
James Brady Jun 21, 2024 ▶ 8:08 How To Hire AI Engineers (ft. James Brady and Adam Wiggins of Elicit)
Jun 25, 2024 neutral
Insight
Albrecht: LLM emergence is an artifact of non-linear evaluation metrics
“This emergent behavior that you're seeing, Is not really emergent behavior, but is really a function of the evaluation metrics that we're using.”
Josh Albrecht Jun 25, 2024 ▶ 55:56 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Aug 7, 2024 neutral
Insight
Ravi: Video segmentation requires far less context than language models
“A difference to language models is probably the amount of context that you need is significantly less than maintaining a long multi-term conversation. And so, you know, coupling this short-term spatial memory with this, like, longer-term object pointers we fou…”
Nikhila Ravi Aug 7, 2024 ▶ 44:41 Segment Anything 2: Memory + Vision = Object Permanence — with Nikhila Ravi and Joseph Nelson
Aug 28, 2024
Insight
Carlini: If LLMs always give desired answers, questions aren't hard enough
“When you're using these models, if you're getting the answer you want, always, it means you're not asking them hard enough questions.”
Nicholas Carlini Aug 28, 2024 ▶ 15:30 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Aug 28, 2024 negative
Insight
Carlini: Language models should not be trusted in adversarial situations
“My research says is entirely on this. Like you probably shouldn't trust these models to do the things in adversarial situations.”
Nicholas Carlini Aug 28, 2024 ▶ 23:21 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Aug 28, 2024 neutral
Insight
Carlini: Always qualify claims about AI with 'for current models'
“Whenever someone says X is true about language models, you should always append the suffix for current models, because I'll be the first to admit I was one of the people who was very much on the opinion that these language models are fun toys and are going to …”
Nicholas Carlini Aug 28, 2024 ▶ 24:36 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Aug 28, 2024 negative
Opinion
Carlini: Most AI commentators spin arguments based on ideology rather than reality
“I feel like most people who write about language models being good or bad, some underlying message of like, you know, they have their camp and their camp is like, AI is bad or AI is good or whatever. And they like, they spin whatever they're gonna say accordin…”
Nicholas Carlini Aug 28, 2024 ▶ 5:40 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Sep 19, 2024 neutral
Assertion Supported
Jamil: LLM Attention Allocates Most Weight to Initial Tokens
“So we have seen with the paper called sync attentions that actually the language model allocates a lot of a lot of, because when you do the attention mechanism, you are doing a weighted sum over the tokens, and each token is given a weight, and we see that mos…”
Umar Jamil Sep 19, 2024 ▶ 50:38 [Paper Club] Writing in the Margins: Chunked Prefill KV Caching for Long Context Retrieval
Sep 27, 2024 negative
Insight
Shunyu Yao argues modern LLMs make prompt engineering tricks obsolete
“I feel like in some sense, I feel like prompt engineering, even it's like a slightly negative word at the time, because it refers to all those kind of weird tricks that you have to apply. But I think we don't have to do that anymore. Like given today's progres…”
Shunyu Yao Sep 27, 2024 ▶ 29:31 Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
Nov 28, 2024 neutral
Insight
Schluntz: Language models prefer small diffs over major refactors
“Language models frequently will produce like a smaller diff when possible, rather than trying to do a big refactor.”
Erik Schluntz Nov 28, 2024 ▶ 11:13 The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic
Jan 2, 2025 neutral
Insight
Lambert: Autoregressive LLMs lack explicit internal structures for intermediate state
“Language models have no ability to do this. They are. Kind of per token computation devices where each token is outputted after doing this forward pass and within that there's no explicit structure to hold these intermediate states.”
Nathan Lambert Jan 2, 2025 ▶ 3:37 The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024]
Mar 7, 2025 neutral
Prediction Not checkable as stated
The Fundamental Language of AI Will Likely Remain Python
“The things that language models understand best tend to be the kind of piece of code that are represented on the internet. You know, maybe I don't think a lot of people would be happy with this, that maybe the kind of fundamental language of AI becomes Python …”
Misha Laskin Mar 7, 2025 ▶ 12:53 Solve coding, solve AGI [Reflection.ai launch w/ CEO Misha Laskin]
Mar 7, 2025 positive
Insight
LLMs Have Strong Priors for Code, But None for Mouse Movement
“A language model, for example, has no prior for a mouse movement. It, you know, it really never seen that on the internet, but it has a really strong prior for code. So out of the two categories, say web browsing and coding, coding is the only one that is actu…”
Misha Laskin Mar 7, 2025 ▶ 7:19 Solve coding, solve AGI [Reflection.ai launch w/ CEO Misha Laskin]
Mar 7, 2025 negative
Prediction Not checkable as stated
Superintelligence Cannot Be Trained Entirely From Scratch
“In the era of language models, I don't think you'll be able to train superintelligence from scratch.”
Misha Laskin Mar 7, 2025 ▶ 5:54 Solve coding, solve AGI [Reflection.ai launch w/ CEO Misha Laskin]
Mar 7, 2025 bullish
Prediction Not checkable as stated
Future UIs Will Be Built as Programmatic Interfaces for AI Models
“Over the coming years, there'll be more kind of AI friendly or language model friendly UIs. And what's friendly to a language model is, is code. So the way a model will be doing work, not just for coding and software engineering, Is by basically making functio…”
Misha Laskin Mar 7, 2025 ▶ 9:37 Solve coding, solve AGI [Reflection.ai launch w/ CEO Misha Laskin]
Apr 29, 2025
Insight
Jin: Language models map directly to reinforcement learning policies
“So the states are, like, the text prefixes, so the initial states, like, the prompt the actions are the next tokens that means, like, a language model is, like, exactly what a policy is, right? A policy maps a state to a probability distribution of our next ac…”
Roger Jin Apr 29, 2025 ▶ 4:03 What is an RL environment? w/ Nous Research's Roger Jin
Jun 6, 2025
Insight
Ameisen: LLMs plan future tokens rather than operating purely myopically
“Language models are next token predictors is like a fact. Like that is what they do. They are trained to predict the next token. However, that does not mean that they myopically only consider the next token When they choose the next token, you can work on brea…”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:13:16 The Utility of Interpretability — Emmanuel Amiesen
Jun 6, 2025 neutral
Insight
Ameisen: Superposition is more severe in language models than in vision models
“That means that like language models pack a lot more in less space than Vision models. So maybe like a kind of like really hand wavy analogy, right? It's like, well, if you want curve detectors, like you don't need that many curve detectors. You know, if each …”
Emmanuel Ameisen Jun 6, 2025 ▶ 35:01 The Utility of Interpretability — Emmanuel Amiesen
Jun 6, 2025 neutral
Assertion Supported
Ameisen: Language model neurons are far less directly interpretable than vision neurons
“If you look at just the neurons of a lot of vision models, you can See neurons that are curve detectors or that are edge detectors or that are high, low frequency detectors. And so you can sort of like make sense of the neurons mostly. But if you look at neuro…”
Emmanuel Ameisen Jun 6, 2025 ▶ 34:31 The Utility of Interpretability — Emmanuel Amiesen
Jun 11, 2025 positive
Insight
Duffy: Playable AI game benchmarks teach people how LLMs operate
“If we make this playable, you know, then it kind of can teach people how to use AI, like language models just by playing. Cause you'll like understand how they work. You have to negotiate against them. You see their responses.”
Alex Duffy Jun 11, 2025 ▶ 6:36 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
Jun 11, 2025 positive
Insight
Duffy: LLM context should contain enough information for a human to play
“If you're looking at the context that's being sent to the language model, could you play the game?”
Alex Duffy Jun 11, 2025 ▶ 17:24 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
Jul 2, 2025 bearish
Opinion
Morris: No Evidence We Can Build Pure Reasoning Models Without World Knowledge
“I don't think we have a lot of evidence that we can build a system like this that like is really, really good at reasoning, but really dumb about the world. Like, I don't know if we have the tools.”
Jack Morris Jul 2, 2025 ▶ 41:31 Information Theory for Language Models: Jack Morris
Jul 2, 2025 neutral
Assertion Supported
Morris: Language models hit a hard memorization plateau regardless of dataset scaling
“Like, no matter how you scale the training size, you hit this like perfect, perfect ish plateau in auto memorization, which we call the model capacity.”
Jack Morris Jul 2, 2025 ▶ 38:32 Information Theory for Language Models: Jack Morris
Jul 16, 2025 bullish
Opinion
Rizwan: Programming currently yields the highest economic ROI for LLMs
“In terms of economic value, programming is definitely the highest cost of benefit for language models right now.”
Saoud (Saud) Rizwan Jul 16, 2025 ▶ 12:34 Cline: The Collaborative AI Coder
Jul 31, 2025 bearish
Assertion Not checkable as stated
Lambert: Current Language Models Cannot Prioritize Experiments for Multi-Week Research Plans
“So it's like, how do you come up with a research plan in 10 weeks? Like there's a lot of, how do you prioritize which experiments to do? It's like, there's a lot of inductive biases that go into that, that I don't like a language model would not do well at tha…”
Nathan Lambert Jul 31, 2025 ▶ 46:31 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Oct 2, 2025 negative
Insight
Howard: Correcting LLM errors in chat history degrades subsequent model answers
“The autoregressive nature of language models means that if they make a mistake, and you correct it, and then say, no, that was a mistake, please do it this way instead. The more often you do that, the worse the dialogue answers get. Because it's in the trainin…”
Jeremy Howard Oct 2, 2025 ▶ 16:17 The antidote to AI fatigue — Answer.ai Solveit
Nov 2, 2025 negative
Insight
Rumbelow: LLMs Are Ill-Suited for Large Numeric Datasets
“Language models are just that, right? They're models of language. They are not particularly well suited for understanding arbitrary, you know, like big numeric data sets.”
Jessica Rumbelow Nov 2, 2025 ▶ 12:40 ⚡️Automating Scientific Discovery - Jessica Rumbelow, Leap Labs
Nov 6, 2025 bullish
Prediction Open · timeframe Nov 2035
Zuckerberg: Specialized virtual cell models will merge into a biological Omni model
“I would imagine you're taking these different types of virtual cell models and eventually merging them into the equivalent of like a biological Omni model, kind of like how on the language model side, you had people that did language and then, you know, people…”
Mark Zuckerberg Nov 6, 2025 ▶ 35:09 Priscilla Chan and Mark Zuckerberg: Frontier AI + Virtual Biology To Solve All Diseases
Dec 31, 2025 negative
Opinion
Merullo: Current machine unlearning techniques merely suppress data rather than removing it
“I would describe it more as not unlearning, but maybe suppression. I think there's, like, really, like, I guess, guarantees that you've fully removed information from a model is, is, I don't think it's been convincingly showed anywhere yet”
Jack Merullo Dec 31, 2025 ▶ 6:44 [State of MechInterp] SAEs in Production, Circuit Tracing, AI4Science, "Pragmatic" Interp — Goodfire
Dec 31, 2025 neutral
Insight
Merullo: LLM memorization spans a gradient from reasoning to rote recall
“You can actually see, like the way that we, like, disentangle memorization, you can kind of see this like, gradient of memorization in between both mechanistically and behaviorally with, like, logical reasoning tasks being quite distinct from rote memorization…”
Jack Merullo Dec 31, 2025 ▶ 6:02 [State of MechInterp] SAEs in Production, Circuit Tracing, AI4Science, "Pragmatic" Interp — Goodfire
Dec 31, 2025 bullish
Insight
LLMs are commoditizing like raw compute, shifting value to abstraction layers
“Language models themselves are more like compute or GPU a generation ago, where what can we build at the layer above? And in software systems, we've traditionally thought of VMware being a great example. You have the operating system and the underlying archite…”
Andy Konwinski Dec 31, 2025 ▶ 7:04 [State of Research Funding] Beyond NSF, Slingshots, Open Frontiers — Andy Konwinski, Laude Institute
Feb 5, 2026 positive
Insight
Deng: Visual interpretability yields faster feedback cycles than language models
“With language models, when you get features, you still have to do auto interpret and things like that to actually get an understanding of what this concept is. But in image and video and world, it's like extremely easy to grok what the concept is because you c…”
Myra Deng Feb 5, 2026 ▶ 53:36 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Feb 5, 2026 positive
Assertion Supported
Deng: Models internally represent uncertainty preceding hallucinatory behavior
“We've seen that models internally have some awareness of like uncertainty or some sort of like user pleasing behavior that leads to hallucinatory behavior.”
Myra Deng Feb 5, 2026 ▶ 27:50 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Feb 12, 2026
Assertion Supported
AlphaFold Models Have Fewer Parameters but Higher Compute Costs Than LLMs
“They, in terms of parameters, are actually not very big. They are definitely below a billion parameters. You know, if you're here these days in LLM space, you know, a model with less than a billion parameters, you'd think can't do anything. But on the other ha…”
Gabriele Corso Feb 12, 2026 ▶ 35:06 🔬Generating Molecules, Not Just Models
Jun 1, 2026 positive
Insight
Ethan He: Language models prompt AI models better than humans
“Most of the people were actually not very good at prompting. Actually, language models have a better sense of how to prompt AI models. AI models know AI models better.”
Ethan He Jun 1, 2026 ▶ 1:29:09 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 1, 2026 bullish
Prediction Not checkable as stated
Ethan He: LLMs will soon become context-aware and manage context
“I think one thing pretty, pretty interesting. I think might be happening soon is the language models will be like context aware and manage its own context.”
Ethan He Jun 1, 2026 ▶ 1:35:33 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 1, 2026 neutral
Insight
Ethan He: Training video models costs roughly the same as medium-scale LLMs
“So surprisingly video models is like the cost is very, is comparable to language models. And obviously the largest scale is language model. Maybe like a medium scale language models.”
Ethan He Jun 1, 2026 ▶ 34:15 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 1, 2026 bullish
Prediction Not checkable as stated
Ethan He: LLM Video Agents Will Orchestrate Diffusion Models and Editing Tools
“Video agents, mostly language models, they'll call these generative model, either it's a separate model or a diffusion head or whatever as tool. So this model can iteratively Refine the results or even like you generate longer content through a very long trend…”
Ethan He Jun 1, 2026 ▶ 1:21:56 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 1, 2026 positive
Insight
Ethan He: Visual intelligence in video generation models stems primarily from language models
“The visual intelligence are actually mostly coming from language. Like, these video models, especially from now, since the diffusion model technology is more mature, the, like, every time you see there, there's some improvement on these models, I would say mos…”
Ethan He Jun 1, 2026 ▶ 1:14:55 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 1, 2026 neutral
Assertion Partly supported
Shawn Wang: Fixed-capability LLM inference costs drop 100x to 1000x annually
“In language models, it is roughly 100 to a thousand times every 12 to 18 months for the same given level of LMSYS ELO.”
Shawn Wang Jun 1, 2026 ▶ 28:15 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Aug 26, 2026 neutral
Insight
Anandkumar: Scientific AI Bottleneck Is Real-World Testing, Not Hypothesis Generation
“Yes, you can do a lot of hypothesis generation. You can have ideas, but ideas are not enough, right? So you can have a lot of ideas. The bottleneck is going, testing, and verifying that they work in the real world.”
Anima Anandkumar Aug 26, 2026 ▶ 4:20 🔬 Why Transformers Hit a Wall the Moment Physics Shows Up — Anima Anandkumar, Caltech
Aug 26, 2026 negative
Opinion
Regulating AI for science like large language models creates serious problems
“A lot of regulatory frameworks equate AI with language models and Yes, language models can, you know, manipulate people, can have all these kinds of harmful impacts that we should think about controlling, but AI for science is different. So I think this one si…”
Anima Anandkumar Aug 26, 2026 ▶ 1:20:19 🔬 Why Transformers Hit a Wall the Moment Physics Shows Up — Anima Anandkumar, Caltech
Sep 4, 2026 bullish
Insight
Anandkumar: Dense physics feedback enables better AI self-improvement than sparse LLMs
“And the difference there is compared to language where self-improvement needs something like human feedback or other reward signals that are very sparse. They just tell you yes or no, thumbs up or down. We have dense feedback because the physics laws, there's …”
Anima Anandkumar Sep 4, 2026 ▶ 18:00 Faster Chips That Don't Melt — Anima Anandkumar & Benedikt Jenik, Accelerated Understanding
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.