Anandkumar: Dense physics feedback enables better AI self-improvement than sparse LLMs
“And the difference there is compared to language where self-improvement needs something like human feedback or other reward signals that are very sparse. They just tell you yes or no, thumbs up or down. We have dense feedback because the physics laws, there's …”
Tandon: AI and LLMs will transform healthcare labor by late 2020s
“The story of the late 20 twenties in healthcare is definitely one of language models and AI replacing standard labor or augmenting that labor to make it much more efficient.”
Tandon: LLMs make international healthcare software localization 10x easier
“And so adapting our tools to the needs of the region It was an order of magnitude easier than we initially thought, and language models have made localization and workflow optimization for a region much easier.”
Anandkumar: Scientific AI Bottleneck Is Real-World Testing, Not Hypothesis Generation
“Yes, you can do a lot of hypothesis generation. You can have ideas, but ideas are not enough, right? So you can have a lot of ideas. The bottleneck is going, testing, and verifying that they work in the real world.”
Regulating AI for science like large language models creates serious problems
“A lot of regulatory frameworks equate AI with language models and Yes, language models can, you know, manipulate people, can have all these kinds of harmful impacts that we should think about controlling, but AI for science is different. So I think this one si…”
Hurst: Internet data does not exist for physical robot control
“These language models are trained off of the entire data on the internet. And that data does not exist for robot control.”
Schmidhuber: Proprietary AI Labs Have No Moat Against Open Source
“None of the companies have has a mode, you know, because whenever there's a new benchmark breaking record or something, benchmark, record breaking language model that does this or this or whatever. A few months later, there's the same thing in open source, you…”
Roberts: AI models improve performance by generating running thought tokens in language
“The natural way it thinks is in language. It's a language model, and so that's sort of this key insight that, that you can cause it to do better just by producing a thought process in, in, in token space, in, in language.”
Shawn Wang: Fixed-capability LLM inference costs drop 100x to 1000x annually
“In language models, it is roughly 100 to a thousand times every 12 to 18 months for the same given level of LMSYS ELO.”
Ethan He: Training video models costs roughly the same as medium-scale LLMs
“So surprisingly video models is like the cost is very, is comparable to language models. And obviously the largest scale is language model. Maybe like a medium scale language models.”
Ethan He: Visual intelligence in video generation models stems primarily from language models
“The visual intelligence are actually mostly coming from language. Like, these video models, especially from now, since the diffusion model technology is more mature, the, like, every time you see there, there's some improvement on these models, I would say mos…”
Ethan He: LLM Video Agents Will Orchestrate Diffusion Models and Editing Tools
“Video agents, mostly language models, they'll call these generative model, either it's a separate model or a diffusion head or whatever as tool. So this model can iteratively Refine the results or even like you generate longer content through a very long trend…”
Ethan He: Language models prompt AI models better than humans
“Most of the people were actually not very good at prompting. Actually, language models have a better sense of how to prompt AI models. AI models know AI models better.”
Ethan He: LLMs will soon become context-aware and manage context
“I think one thing pretty, pretty interesting. I think might be happening soon is the language models will be like context aware and manage its own context.”
Tandon: Healthcare LLMs will improve outcomes enough to lower malpractice premiums
“That yes, occasionally you have, like, weird performance, but on the averages, and like in the 99% of cases, it improves outcomes so much that the malpractice implications are actually very net positive, and I think because of that, you will, you should see a …”
McDermott: Replicating a ServiceNow app with LLMs costs 10x more
“We've actually done the math on this. And so for a simple application on our platform, it would be 10 times greater in cost to try to replicate it with a language model.”
Vuong: SayCan showed language models can reduce robot-specific training data
“I think the first is Seikan, which to me was the first demonstration of language model and how you can bring all of the common sense knowledge in language model into robotics, and therefore that significantly kind of reduces the need to collect robot-specific …”
Rieseberg: AI models are grown rather than built, making capabilities unpredictable
“We often say that models are more grown than built out of the nature of how these language models are being made. So you don't always know ahead of time necessarily what are they going to be very good at, what are they maybe going to be bad at.”
Levine: General Language Models Proved Easier Than Narrow NLP Systems
“Again, in much the same way that for language models, it turned out to be Easier in some ways to solve natural language tasks in their full generality than to narrowly target like machine translation or sentiment analysis or whatever.”
Levine: Weakly Labeled Web Data Builds Foundational AI World Understanding
“When you can leverage Weekly labeled data, like data that you like, you know, in the case of language models that you just mine from the web, you actually learn more about the world. So you establish, like, foundation of world understanding, and then on top of…”
Levine: LLMs Show True Compositional Generalization Through IPA Paragraphs
“But if you ask a good language model, it will write paragraphs in IPA for you. And that is compositional generalization. It means that you have never seen this particular language, this particular alphabet, used to write paragraphs, but you understand paragrap…”
Poetiq's agentic harness outperforms new base models without code changes
“With poetic what we end up giving you is a you know, people are calling these things harnesses now, but you know, or agentic system or whatever you want to call it, that sits on top of one or more language models, and it just performs better than them. And whe…”
AlphaFold Models Have Fewer Parameters but Higher Compute Costs Than LLMs
“They, in terms of parameters, are actually not very big. They are definitely below a billion parameters. You know, if you're here these days in LLM space, you know, a model with less than a billion parameters, you'd think can't do anything. But on the other ha…”
Deng: Models internally represent uncertainty preceding hallucinatory behavior
“We've seen that models internally have some awareness of like uncertainty or some sort of like user pleasing behavior that leads to hallucinatory behavior.”
Deng: Visual interpretability yields faster feedback cycles than language models
“With language models, when you get features, you still have to do auto interpret and things like that to actually get an understanding of what this concept is. But in image and video and world, it's like extremely easy to grok what the concept is because you c…”
LLMs are commoditizing like raw compute, shifting value to abstraction layers
“Language models themselves are more like compute or GPU a generation ago, where what can we build at the layer above? And in software systems, we've traditionally thought of VMware being a great example. You have the operating system and the underlying archite…”
Merullo: LLM memorization spans a gradient from reasoning to rote recall
“You can actually see, like the way that we, like, disentangle memorization, you can kind of see this like, gradient of memorization in between both mechanistically and behaviorally with, like, logical reasoning tasks being quite distinct from rote memorization…”
Merullo: Current machine unlearning techniques merely suppress data rather than removing it
“I would describe it more as not unlearning, but maybe suppression. I think there's, like, really, like, I guess, guarantees that you've fully removed information from a model is, is, I don't think it's been convincingly showed anywhere yet”
Isenberg: AI Domain Terminology Triggers Sophisticated Reasoning in LLMs
“Using advanced prompting terms can trigger more sophisticated modes of operation. So models are trained on a vast amount of text about AI itself. So using terms from the field activates specific powerful behaviors.”
Soldaini: Training LLMs on longer sequences causes quadratic compute slowdown
“It's because the longer the input that a model is trained on, the slower it is. The rate at which it gets slower, it's higher than the length of a context. It's a quadratic slowdown.”
Lambert: Big tech will realize 95-98% of LLM potential by 2030
“I think that how I describe it is that big tech has all collectively realized that these language models plus scaffolding is going to unlock absolutely incredible value. And I have very high probability, barring extreme geopolitical situations, that big tech E…”
Li: Robotics trails LLMs because training data lacks 3D physical actions
“In in language, because they have this perfect setup where their training data are in words, eventually tokens, and then they produce a model that outputs words. So you have this perfect alignment between what you hope to get, which we call objective function,…”
Li: Spatial intelligence and world models are as important as LLMs
“We believe that spatial intelligence and world modeling is As important, if not more, to language models and complementary to language models.”
Zuckerberg: Specialized virtual cell models will merge into a biological Omni model
“I would imagine you're taking these different types of virtual cell models and eventually merging them into the equivalent of like a biological Omni model, kind of like how on the language model side, you had people that did language and then, you know, people…”
Kuyda: OpenAI temporarily abandoned language models for video game agents
“And very quickly they stopped working on language models, and we were very upset because we really wanted to continue going there, but they didn't want to talk about any language models because no one was really working on them and that may have made us feel v…”
Rumbelow: LLMs Are Ill-Suited for Large Numeric Datasets
“Language models are just that, right? They're models of language. They are not particularly well suited for understanding arbitrary, you know, like big numeric data sets.”
Andreessen: Major technological breakthroughs require at least 40 years of prior work
“If you look at the history of Technology, it's almost always the case that the big breakthroughs are the result of, you know, usually at least 40 years of sort of work ahead of time, you know, four decades. Right, in fact, language models themselves are the cu…”
Schrittwieser: Task duration dictates how much work can be delegated to AI
“The reason I think why task length specifically is interesting is because that's What allows you to delegate more and more work to language models, to agents. Now, even if you have a very clever model, but if it needs feedback or the interaction with you very …”
Schrittwieser: Language models possess an implicit world model
“So I think, yes, I would say that language models have an, not an explicit world model, but they do have an implicit model of the world.”
Schrittwieser: Adding model reasoning improves RL training stability and scaling
“One direction of scaling RL and making it more stable is by improving this by, for example, putting more reasoning into your language model to generate much more high quality training data. That can then give us training that is much more stable, and then we c…”
Tworek: Calling LLMs strictly next-token predictors is inaccurate in RL era
“Language models do on their own, like fundamental level is they are often called as next token prediction machines. And that's not completely accurate in the age of reinforcement learning, but they still operate on mostly on tokens that are mostly text.”
Misra: Pure language processing is insufficient for human-level intelligence
“Language is great, but language is not the answer. You know, when I'm looking at catching a ball that is coming to me, I'm mentally doing that simulation in my head. I'm not translating it to language to figure out where it'll land.”
Zelikman: Language models can be trained to simulate students for test design
“Like, even back in my PhD, I think one of my, I guess, less well-known works was actually about, we showed that you can train language models to simulate different kinds of students. For tests. Yeah, yeah. And by simulating students, you can actually design be…”
Howard: Correcting LLM errors in chat history degrades subsequent model answers
“The autoregressive nature of language models means that if they make a mistake, and you correct it, and then say, no, that was a mistake, please do it this way instead. The more often you do that, the worse the dialogue answers get. Because it's in the trainin…”
Douglas: Simple RL methods work better on language models than complex strategies
“One of the craziest things about RL on language models in the, in, like, the RL from verified rewards regime, is it's almost the simplest possible thing. It's, like, almost too simple to work. And this is, again, comes back to that question of taste, where Rea…”
Douglas: AI reasoning strategies emerge naturally with enough compute and RL feedback
“Give it math questions, tell it whether it got them right or wrong, and the model will learn. This is, it comes down to a bit of lesson in scale and search, is just allow the model to search, have enough compute to run the experiments, and the model actually e…”
Joseph: Training purely on raw LLM generations cannot produce a better model
“Theoretically, I shouldn't be able to train a better model than that. Like, I'm just going to get the same thing out. So I think that's-”
Frosst: Prompt engineering will be replaced by understanding how LLMs actually work
“So I think the idea of saying like, oh yeah, you got to learn how to prompt is going to go away. I think the idea of saying you need to learn how language models work and you need to know what they can and can't do in the same way you had to learn how a comput…”
Frosst: AI language models are critical national infrastructure like power plants
“I think it's a good idea for countries to have infrastructure within their countries. Like, I think it's a good idea for people to have power plants in the country. You know, I like that Canada has several nuclear power plants and has several water power plant…”
Lambert: Current Language Models Cannot Prioritize Experiments for Multi-Week Research Plans
“So it's like, how do you come up with a research plan in 10 weeks? Like there's a lot of, how do you prioritize which experiments to do? It's like, there's a lot of inductive biases that go into that, that I don't like a language model would not do well at tha…”
Finn: LLMs can generate synthetic prompts to relabel robot data
“We can use language models to relabel and generate hypothetical human prompts for the scenarios that the robots are in.”
Mann: Language Models Understand Human Values in a Core Way
“And since then, my estimation of how hard the problem would be has gone down significantly actually because things like language models actually do really understand human values in a core way. The problem is definitely not solved, but I'm more hopeful than I …”
Laskin: AI models will interact with enterprise software primarily via APIs
“And so the way these language models are going to interact with any piece of software, not just Software engineering software, like Salesforce and other CRMs and creative tools and so forth. The majority of those interactions are going to be through function c…”
Rizwan: Programming currently yields the highest economic ROI for LLMs
“In terms of economic value, programming is definitely the highest cost of benefit for language models right now.”
Srinivas: Perplexity bet LLMs would handle reasoning over unstructured data
“So we bet on the fact that language models can do all the reasoning and parsing and like structuring later, but the more important thing is to start with something more unstructured, and that ended up becoming perplexity.”
Morris: Language models hit a hard memorization plateau regardless of dataset scaling
“Like, no matter how you scale the training size, you hit this like perfect, perfect ish plateau in auto memorization, which we call the model capacity.”
Morris: No Evidence We Can Build Pure Reasoning Models Without World Knowledge
“I don't think we have a lot of evidence that we can build a system like this that like is really, really good at reasoning, but really dumb about the world. Like, I don't know if we have the tools.”
Acharya: LLMs are averaging machines, but great art requires the edge
“These language models are these averaging machines, and you don't, with art, you almost definitely don't want the average of all the novels or all the writing or all the authors. You want something that's at the edge.”
Zach Cohen: LLMs like ChatGPT are already good enough to teach
“Language models are good enough to teach. Like, they really are, right? Like, if you want to learn something, ChatGPT is a great place to go learn something.”
Duffy: Playable AI game benchmarks teach people how LLMs operate
“If we make this playable, you know, then it kind of can teach people how to use AI, like language models just by playing. Cause you'll like understand how they work. You have to negotiate against them. You see their responses.”