why aren't all 109 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 4 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Prediction Open · timeframe Jul 2028
Lambert: LMSYS is probably setting up a deep research arena
“I mean, they're probably setting up a deep research arena, because that's the data that, I mean, if I was open AI working on deep research, that's the data that I want, and there are competitors, and LMSYS is the entity that has the market placement to set it …”
Insight
Lambert: North star of reasoning models is dynamic token budget calibration
“I think that has to be the north star for most people working on reasoning, which is the model will just Spend the right amount of tokens on it.”
Assertion Partly supported
Lambert: SimpleQA benchmark scores drop across reasoning models tested without tools
“You look at all the evals from reasoning models, and one of the trends is that, like simple QA numbers all drop. It's like DeepSeq R-one to the new R-one, it goes down. It's like all the new, like, QN-II to QN-III, simple QA goes down, at least when you're eva…”
Prediction Open · timeframe Jul 2030
Lambert: All major frontier AI labs will build their own search indexes
“I think they'll all do end up doing their own index and it should, it's one of those things that's like Google should have an advantage again, but who knows if they do.”
Insight
Lambert: Scaling RL long enough requires a curriculum of increasing difficulty
“If you scale RL long enough,
You're going to need a curriculum of things getting harder. And like, that's pretty obvious.”
Prediction Open · timeframe Jul 2028
Lambert: Academics cannot match industry compute on Humanity's Last Exam
“I just think it's kind of unlikely that we're going to win as a academic and a state of the art number because they're going to start spending millions of tokens per query. And it's just a lot of, it's a lot of compute burn. Like the getting, beating that on t…”
Insight
Lambert: Reasoning models solved basic skills; planning is the next frontier
“So I came up with four and the foundational one was skills, which is What I would say that we have already done with O-one and R-one, which is you do a lot of RL, you show the inference time scaling works and you get really high benchmark numbers. And then the…”
Assertion Not checkable as stated
Lambert: Current Language Models Cannot Prioritize Experiments for Multi-Week Research Plans
“So it's like, how do you come up with a research plan in 10 weeks? Like there's a lot of, how do you prioritize which experiments to do? It's like, there's a lot of inductive biases that go into that, that I don't like a language model would not do well at tha…”
Insight
Lambert: Parallel compute provides robustness rather than low-probability search
“Well, I don't think we're using parallel compute in a way to search over like low probability tokens. We're using it to get robustness. If you use like O-one pro is, it was so nice because it Just had a very predictable depth to it, even on niche topics where …”
Insight
Lambert: Code maintainability is a human preference problem in RL
“The software stuff is not easy because it's almost like maintainability almost feels like a human preference type issue again, where somebody could look at it and be like, yeah, that's not as good, but adding the heuristic and trading seems very messy.”
Insight
Lambert: RLVR is harder to over-optimize on math than code
“For math, it's a bit harder to over optimize, I think. Unless you have tools and the model learns to search and cheat instead of learning math, which I'm sure somebody could see that out in the world, which is like, oh, I'll just find the, you're training. It'…”
Insight
Lambert: Custom personality fine-tuning is open source AI's winning turf
“If open models are to win, part of it could be just, like, everybody can have exactly the model they want. We're serving GPT-IV. It's kind of its thing. You can prompt it, but if fine-tuning is more effective than prompting, everybody can have the model. That …”
Prediction Not checkable as stated
Lambert: OpenAI's open model will be best-in-class in its size category
“I expected. It'll be best in class for some size Category in some subset of tasks. That's like, OpenAI only does things like that.”
Prediction Open · timeframe Jul 2028
Lambert: Jony Ive and OpenAI hardware will run in the cloud
“I think that thing will run on the cloud. I don't think that'll run local anyways.”
Assertion Not checkable as stated
Lambert: Early DeepSeek and Qwen reasoning models are substantially narrower than o1
“And I think that these models are really substantially narrower than these full O-one models from OpenAI. So OpenAI is, if you use O-one, you can do it for a lot more tasks. If you use, like I was using the DeepSeq model, and it's supposed to be for math or co…”
Prediction Not checkable as stated
Lambert: Open community will eventually match OpenAI's large-scale RL infrastructure
“And this is something that these early relative models are not going to be doing because we don't like, no one has this infrastructure like open AI does. It'll take a while to do that, but people will make it.”
Opinion
Lambert: Reinforcement fine-tuning will succeed where answer correctness matters over style
“It is just a new paradigm for fine tuning, and I have seen some of this work, and I'm pretty Optimistic that it'll work for kind of kind of really specific capabilities where answers matter rather than features in your style of text mattering.”
Assertion Supported
Lambert: Reinforcement fine-tuning requires only dozens of labeled samples
“This reinforcement fine tuning does many passes over the data, which is why they can say you only need dozens of labeled samples to actually learn from it, which is very different than. Previous training regimes”
Prediction Not checkable as stated
Lambert: Open judge models will become core open RL infrastructure
“We already have a bunch of open models that are doing like judge of models and Prometheus and other things that are designed specifically for LM as a judge. And I see that continuing to just become part of this kind of open RL infrastructure.”
Disclosure
Lambert: AI2 received early industry tip on reinforcement fine-tuning
“We got a tip from a industry lab member to do this a few months early. So we got a head start”
Opinion
Lambert: Major tech companies now realize they need dedicated RLHF teams
“I think any major company that wasn't doing RLHF is now realizing they have to have a team around this.”
Assertion Not checkable as stated
Lambert: Open-source and research communities lack large-scale RLHF capabilities
“At the same time, we don't have a lot of that in the like open and research communities at the same scale.”
Opinion
Lambert: Anthropic is considered the master of RLHF techniques
“Everyone knows Anthropik is kind of the masters of this”
Prediction Not checkable as stated
Nathan Lambert: RL will remain a distinct field from language modeling
“I think in the long run it will still settle out, or RL will still be a field that people work on just because of these kind of fundamental things that I talked about, that it's just viewing the whole problem formulation different than predicting text, really,…”
Insight
Lambert: RLHF operates as a contextual bandit problem, not sequential RL
“What happens is the state is a prompt, and then you do a completion, and then you throw it away, and you grab a new prompt. Where really in like RL, you, as an RL researcher, you would think of this as being like, you take a state, you complete, Get some compl…”
Prediction Not checkable as stated
Lambert: AI community will clarify if chain-of-thought maps to RL within a year
“I think in the next year that'll probably get kind of made more concrete by the community on, like, if you can easily draw out, like, if chain of thought reasoning is more like RL.”
Insight
Lambert: RLHF provides a richer signal per unit of compute than pre-training
“As reinforcement learning is so much less compute, like it, Is a richer signal in terms of its impact, because if they could do what RLHF is doing at pre-training, they would, but they don't know how to have that effect in like a stable manner. Otherwise, ever…”
Insight
Lambert: Instruction tuning is more important than RLHF for most practitioners
“I think for most people, instruction tuning is probably still more important in their day-to-day life. I think instruction tuning works very well. You can write samples by hand that make sense. You can get the model to learn from them. You could do this with v…”
Assertion Supported
Lambert: DPO benchmark gains rely largely on the UltraFeedback dataset
“Everyone's using this ultra feedback data set and it boosts AlpacaVal, MTBench, TruthfulQA, and like the qualitative model a bit. We don't really know why.”
Insight
Lambert: Scaling from 7B to 70B parameters fixes nuance and repetition
“I think the things that people see now is like the small models don't really handle nuance as well, and they could be more repetitive if, even if they have really good instruction tuning, but if you take that kind of seven to seventy billion parameter jump, li…”
Opinion
Lambert: The vast majority of instruction tuning data remains simple Q&A
“There's much more, like there's surely kind of more tricky things that people do, but I still think the vast majority of it is question and answer. It's like, please explain this topic to me, generate this thing for me. That hasn't changed that much this year.…”
Insight
Lambert: RLHF fails completely without a KL divergence constraint
“The KL constraint prevents that. There's not that much documented work on that, but there's a lot of people that know if you take that away, it just doesn't work at all.”
Insight
Lambert: RLHF avoids controversial preferences, focusing on correctness and style
“The reason this really is done on a deep level is that you're not actually trying to model any, like, contestable preference in this. Like, you're not trying to go into things that are controversial or anything. It's really, like, the notion of preference is t…”
Opinion
Lambert: OpenAI trains its RLHF models primarily on good-vs-good answer comparisons
“I think open AIs of the world are all in good answer, and have learned to eliminate everything else.”
Assertion Supported
Lambert: Anthropic, ChatGPT, and Bard Use Post-Generation Moderation Classifiers
“Anthropic and ChatGPT and Bard almost surely have a classifier after, which is like, is this text good? Is this text bad?”
Assertion Not checkable as stated
Lambert: Meta spent roughly $6M to $8M on Llama 2 preference data
“So I would say, still say, like, six to eight million is safe to say that they're spending, if not more, they're probably also buying other types of data and or throwing out data that they don't like.”
Insight
Lambert: Published AI training compute costs understate total experimentation budgets
“The compute costs listed in the paper always are way lower, because all they have to say is, like, what is one run cost, but they're running tens or hundreds of runs, so it's like, okay, like, it's kind of a meaningless number, yeah.”
Opinion
Lambert: Only 20% to 40% of Meta's Llama RLHF data is useful
“I do think that if we had all the llama data, we wouldn't know what to do with all of it. Like, probably, like, 20 to 40% would be pretty useful for people, but not the whole data set. Like, a lot of it's probably kind of gibberish, because they had a lot of d…”
Prediction Held up
Lambert: Practitioners Will Adopt Constitutional AI for Preferences in 2024
“I think in twenty-twenty-four at some point people will start doing things like constitutional AI for preferences.”
Assertion Not checkable as stated
Lambert: Zephyr was the first open model to succeed with public RLHF
“I think Zephyr was the first model that showed success with RLHF in the public, but that's a long time from everyone knowing that it was something that people are interested in to having any, like, check mark.”
Insight
Lambert: RLHF Performance Depends on Data and Systems Over RL Details
“It really ends up being kind of, like, gibberish that I think is less important now, because it's more about data and infrastructure than RL details, than, like, value functions and everything. A lot of the papers have different terms in the equations. I think…”
Opinion
Lambert: Scaling PPO in RLHF is a Nightmare
“PPO is kind of a nightmare to scale.”
Assertion Supported
Lambert: Meta Used Rejection Sampling to Bootstrapping Llama 2 RLHF
“Llama started their RLHF process with this to get some signal out of preference data. That preference data went into a reward model, and then the reward model did a good enough ranking that it was, like, essentially superpowered instruction tuning based on rew…”
Prediction Not checkable as stated
Lambert: Offline RL could take off due to simpler training pipelines
“There's a few papers that people have published
Not a lot of traction.
I think it could take off.
Some people that I know in the RLHF area really think a lot of people are doing this in industry just because it makes the kind of training process simpler and th…”
Insight
Lambert: Anthropic Constitutional AI and OpenAI Superalignment share intellectual roots
“The constitutional AI and the super alignment is, like, very conceptually linked. It's like a group of people that has, like, a very similar intellectual upbringing, and they work together for a long time, like, coming to the same conclusions in different ways…”
Insight
Lambert: DPO is closer to RLHF than RLHF is to RL
“I think
DPO is closer to RLHF than RLHF is to RL.”
Prediction Held up
Lambert: More DPO models will emerge than any other method
“I expect to see more DPO models than anything else in the next six months.”
Assertion Supported
Lambert: DPO Has Become the Standard Release Expectation for Open-Source LLMs
“I think DPO releases are kind of becoming expected because Mistral released a DPO model as well. I think the slide after this is just like, there's a ton. It's like Intel releases DPO models, Stability releases DPO models. At some point, you just have to accep…”