why aren't all 109 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 4 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Supported
Lambert: Tulu 3 matches or beats Meta Llama 3.1 on core evals
“On, like, core evals for our Suite of models from, I think, eight, seven D and four or five B is based on llama at the time. It's like it matches or beats meta on these core valves.”
Opinion
Lambert: Deep Research relies on modular RL tasks rather than end-to-end outcomes
“I think the deep research blog post kind of hints that they do a bunch of small scale RL and then poof, the system works. Which I think is much more of what's happening is people train on a bunch of small things and they do some prompting and they see that whe…”
Insight
Lambert: RLHF can never be permanently solved
“In the same way that chatbot arena can never be saturated. RLHF can never be solved.”
Insight
Lambert: The RL algorithm is not the most important component in reasoning models
“I definitely don't think the algorithm tends to be the most important thing.”
Prediction Not checkable as stated
Lambert: Hybrid reasoners may be phased out except for niche uses
“I think in plenty of ways, like hybrid reasoners might just be aged out except for niche applications because quality is so much more important than having a hundred X less inference tokens. It's like you just pay for it and compute and that'll get better.”
Insight
Lambert: SFT cannot teach emergent tool use; models must learn via RL environments
“It's very easy to get the model to do tools if you prompt it to, but it's very hard to get the like RL model to learn that the tool is useful. And that's why it's to go through these things where it's like 80 failed tool uses and it still gets it or like it st…”
Insight
Lambert: OpenAI's Model Spec is more useful than Anthropic's Constitution
“The model spec is much more useful than a constitution because the constitution is like an intermediate training artifact that you give to the training algorithm in order to get the model that you want. It is not necessarily like what model did we, like we don…”
Insight
Lambert: Top AI talent is dramatically cheaper than GPU clusters
“Talent is cheaper than GPUs by a dramatic margin, and At the end of the day, it's like, okay, if we're spending this much, they go to the room and they stare in the mirror and you're like, wait, it might not actually be that ridiculous to spend this money on t…”
Assertion Partly supported
Lambert: OLMo 32B roughly matches original GPT-4 level while fully open
“Like Olmo-Thirty-Tube is if you squint like original GPT-IV level and fully open.”
Insight
Lambert: OpenAI o1 uses large-scale RL on verifiable outcomes, not MCTS
“You should take open AI at their face value, which they are doing very large scale RL on the verifiable outcomes is what I've added, especially in context of the RL API that they've released, which I'll talk about more, but most of the reasons to believe in mo…”
Prediction Not checkable as stated
Major Foundation Model Companies Will Train on AI2's Vision Data
“The things that this model is good at are things that all the foundation companies, like they're just going to take our data and train on it.”
Insight
Lambert: Startups should avoid RLHF unless it offers niche advantage
“I don't really recommend most startups to do it unless it's like going to provide them a clear competitive advantage in their kind of niche. Yeah. Because you're not going to make your model ChatGPT like better than OpenAI or anything like that.”
Insight
Lambert: A 100x smaller language model filters output better than RLHF
“You could use like a hundred times smaller language model and do much better at filtering than RLHF”
Opinion
Lambert: Chatbot Arena is the best available evaluation benchmark for LLMs
“I have, if we make it to evaluation, I'd pretty much say that Chat Arena is the best limited evaluation that people have to learn how to use language models, and like, It's very valuable data”
Insight
Lambert: Reinforcement Learning Does Far Less for Alignment Than People Think
“It's like, now you're getting to the point where you don't even really need this to get a good model, so that's why it's like, okay, the RL is such a small part of the actual, like, doing RLHF. Like, RLHF is a metaphor for, like, all language model adaptation,…”
Assertion Supported
Lambert: RLHF has not been shown to improve underlying model benchmark capabilities
“RLHF is not that shown to improve capabilities yet. I think one of the fun ones is from the GPT-IV technical report. They essentially listed their kind of bogus evaluations, because it's a hilarious table, because it's like LSAT AP exams, and then like AMC-X a…”
Insight
Lambert: RLVR is broader than ground truth because code is verifiable
“The verifiable rewards is actually a more general notion because only like math questions have a ground truth where code is verifiable, precise instruction following is verifiable.”
Insight
Lambert: Inference scaling plots misleadingly suggest search is an easy control knob
“The core of that article is just, they're taking points from within training, or there's a natural variance, and then you line them up. And if you line them up, then you get this nice inference time-scaling behavior, which is, and now people, a lot of people h…”
Opinion
Lambert: AI benchmarks like ARC-AGI should prioritize testing without harnesses
“Harnesses are cool, but they're gonna, they're,
They're a handicap that's changing the learning dynamics substantially. So it's good. It's good demos, but I feel like the core thrust has to be no harnesses.”
Insight
Lambert: AI academics must build datasets and evals rather than papers
“If you're trying to have impact in AI right now, it's as an academic, you have to like level up out of papers to artifacts, which is models, datasets, evals. Datasets and evals are easier for people to have impact on.”
Opinion
Lambert: Labs trade code usability for massive RL performance gains
“That's just like the labs are trading off massive gains in performance or small detriments in usability. And it's like, do you ship that model? Yeah. Like you just ship it and deal with it later, but I'm sure they could, I'm sure that's a fixable thing.”
Insight
Lambert: RLVR on math does not degrade knowledge benchmark performance
“I think part of the intuition of RLVR is that the model is good at knowing which prompt area it is, which is why the models don't get worse on knowledge benchmarks if you're trading on like just math or precise instruction following. So the model just kind of …”
Opinion
Lambert: The local model community is much smaller than assumed
“Like the local modeling community, I think is much smaller than people give it credit for, because most of the use for open models is still in APIs.”
Opinion
Lambert: Meta withholding its leading benchmark model is bad execution
“But to be a model that claims to be open and then not release the model that is your leading claim is just, like, that is, like, bad execution.”
Opinion
Lambert: AI is developing new reasoning modes that look less human
“I think a big Trend of the year is that we're seeing new types of language model reasoning that look less human and that Can be good for kind of separating the discourse for expecting a really narrow type of behaviors.”
Opinion
Lambert: OpenAI's o1 uses token streams as intermediate state compute
“Why oh, one is exciting is because it's a new type of language models that are going to maximize on this view of reasoning, which is that chain of thought in kind of a forward stream of tokens can actually do a lot to achieve better outcomes when you're doing …”
Assertion Supported
Molmo Reads Clocks but Fails to Generalize to Dials
“The model didn't work on clocks and then the lead was really on clocks and no models work on clocks. So they're like, we've got to make it work on clocks. One of the interesting things is that it doesn't work on dials, even though it works on clocks.”
Opinion
Lambert: OpenAI's rumored Q* was likely just a moderate benchmark bump
“They probably just got like a moderate bump on one of their benchmarks, and then everyone lost their minds, so it doesn't really matter.”
Insight
Lambert: AI field hasn't learned much more, just articulates ignorance better
“I think it's, I try to do it every six or 12 months is my current, is my estimated cadence, just to refine the ways that I say things, and people will see That we don't know that much more, but we have a bit of better way of saying what we don't know.”
Opinion
Lambert: Reinforcement learning in language models is contrived and not real RL
“And the view of RL in language models is pretty contrived already. So it's not like we're doing real RL.”
Insight
Lambert: RLHF reward models lack inductive biases for true preferences
“In RLHF calling things a preference model is a little annoying because There's no inductive bias of what a preference is.”
Insight
Lambert: Claude's constitution dictates output priorities, not model beliefs
“If you look at Claude's constitution, like, that doesn't mean the model believes these things. It's just trying Trained and to prioritize these things.”
Assertion Not checkable as stated
Lambert: Frontier AI labs do not recruit economics or social choice academics
“The RLHF techniques that people use were built in, like, labs like OpenAI and DeepMind, where there are some of these people, they have, they, these places do a pretty good job of trying to get these people in the door when you compare them to, like, startups,…”
Opinion
Lambert: Open-source claims of ChatGPT-level performance are overblown
“I think the claims of ChatGPT level are long overblown in most of the things in open source.”
Opinion
Lambert: Annotator disagreement data in RLHF is not used for anything
“I don't really think this disagreement data is used for anything, but it's good to know, like, what the distribution of prompts is, who's doing it, how many samples you have, controlling the workforce.”
Prediction Not checkable as stated
Lambert: Open source will learn to train models on arbitrary preference data
“I really think people in open source and academics are going to figure out how to use any preference data on any model just because they're scrappy.”
Assertion Supported
Lambert: GPT-4 achieves 80% preference labeling agreement versus 70% for humans
“Essentially, people also think that synthetic data is, like, GPT-IV is more accurate than humans at labeling preferences, so if you look at these diagrams, like, humans are about 60 to 70% agreement, or, like, that's what the models get to, and if humans are a…”
Prediction Not checkable as stated
Lambert: OpenAI Will Not Aggressively Ban Synthetic Training Scraping
“I don't expect OpenAI to go too crazy on this, because they're just gonna, there's gonna be so much backlash against them.”
Assertion Supported
Lambert: RLHF reward models achieve only 65% to 75% validation agreement
“If you look at a test set, you'll have a chosen and rejected, and you can take the reward model you're training, pass in those completions, And you see if the chosen predicted reward, so the scalar number is higher than the rejected predicted reward, and this …”
Assertion Not checkable as stated
Lambert: AI2 trained 70B TÜLU 2 on the first run without ablations
“Let's just try the Zephyr recipe on seventy billion parameters, and it's literally, like, the first run. It's like, we did no ablations, didn't change any parameters, we just copied them all over. And like, that's the model that people have been working with”
Disclosure
AI2 Plans to Release Fully Open Pre-Trained LLMs With Data and Code
“The Allen Institute is training, pre-training language models, or pre-training, like, open language models, where we'll be able to share, like, data, code, everything, the kind of horn that everyone likes to get annoyed about these days, it's like, well, I'm n…”
Assertion Contradicted
Lambert: GPT-4 Turbo Gap Over Original GPT-4 Exceeds TÜLU 2 to GPT-4 Gap
“So it's like the difference from these, the GPT-IV Turbo to like the GPT-IV that was first released is bigger than the difference from Tulu-II to GPT-IV.”
Opinion
Lambert: Frontier Labs Lack Visibility into Cross-Model RLHF Sensitivity
“I think big labs are so over-indexed, are indexed on their own base models, so they don't know, like, what's swapping between CloudBase or GPT-IV-Base, how that would change any notion of preference or what you do with RLHF.”
Assertion Supported
Lambert: OpenAI retrains reward models with curated and user prompt mixtures
“And this is like a sort of outer loop optimization that no one in the open is even remotely qualified to talk about, but OpenAI does monitor and they'll like rerun RLHF and train a new reward model with a mixture of their curated data and user prompts to try t…”
Opinion
Lambert: Scale AI has historically struggled to retain technical ML talent
“I think they've historically had trouble keeping, like, technical ML talent, but they've started a new research lab, so that should help.”
Assertion Not checkable as stated
Lambert: Academia relied on UltraFeedback for open preference tuning for a year
“The academic community had been using this one data set since like all the way back in the hugging face models of like Zephyr beta is when this ultra feedback data set got popular. And still a year later is like this state of the art data set for open preferen…”
Insight
Lambert: Context compression is crucial for long-horizon AI agents
“Compressing context, like that's not I don't think that's really a verifiable thing, but that being messed up, like that's a super crucial skill for long context actions and long longer tasks is just compressing well, and that's going to take some training nov…”
Assertion Supported
Lambert: Frontier AI labs still rely on human preference data
“Every time I check in with people at frontier labs, they're like, yeah, we still use human preference data.”