Jan 2, 2025 · 16m · latent-space

The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024]

Nathan Lambert · 14m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this NeurIPS 2024 presentation, Nathan Lambert explores the state of artificial intelligence reasoning, examining how chain-of-thought prompting, verifiable post-training reinforcement learning, and reinforcement fine-tuning drive frontier problem-solving capabilities.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 0.0 Guest teaching 0.0 Guest disagreement 1.8 The hosts pushing back 0.0
05100:0010:000:57–3:08 · The hosts as informed peer 0/10 Defining Reasoning and Challenging Human-Centric Benchmarks Lambert delivers a solo talk challenging standard human-centric definitions of reasoning, criticizing public discourse around LLM limitations as ridiculous. The format is a monologue with zero host participation.3:10–5:16 · The hosts as informed peer 0/10 Chain of Thought as Intermediate Compute Lambert breaks down how chain-of-thought acts as intermediate variable compute in forward-pass token streams. The monologue continues without any host dialogue.5:19–8:45 · The hosts as informed peer 0/10 Decoding OpenAI's o1 and Open Replications Lambert dispels speculative community theories about OpenAI o1 utilizing PRMs or MCTS, arguing it is fundamentally massive RL on verifiable outcomes. As a solo conference presentation, host scores remain zero.8:50–11:16 · The hosts as informed peer 0/10 The Emergence of Reinforcement Fine-Tuning Lambert explains the mechanics of reinforcement fine-tuning (RFT), highlighting how small amounts of RL on top of strong base models avoid performance degradation. No host is present.11:20–13:39 · The hosts as informed peer 0/10 Data Formats and Grader Models in RL Lambert explains prompt-answer data formatting and the role of grader models and LLM judges in reward shaping. The segment is purely instructional.13:41–16:08 · The hosts as informed peer 0/10 Empirical Evaluation Results and Concluding Remarks Lambert shares empirical evaluation graphs from AI2 research and invites the audience to ask questions. There is no host involvement.0:57–3:08 · Guest teaching 0/10 Defining Reasoning and Challenging Human-Centric Benchmarks Lambert delivers a solo talk challenging standard human-centric definitions of reasoning, criticizing public discourse around LLM limitations as ridiculous. The format is a monologue with zero host participation.3:10–5:16 · Guest teaching 0/10 Chain of Thought as Intermediate Compute Lambert breaks down how chain-of-thought acts as intermediate variable compute in forward-pass token streams. The monologue continues without any host dialogue.5:19–8:45 · Guest teaching 0/10 Decoding OpenAI's o1 and Open Replications Lambert dispels speculative community theories about OpenAI o1 utilizing PRMs or MCTS, arguing it is fundamentally massive RL on verifiable outcomes. As a solo conference presentation, host scores remain zero.8:50–11:16 · Guest teaching 0/10 The Emergence of Reinforcement Fine-Tuning Lambert explains the mechanics of reinforcement fine-tuning (RFT), highlighting how small amounts of RL on top of strong base models avoid performance degradation. No host is present.11:20–13:39 · Guest teaching 0/10 Data Formats and Grader Models in RL Lambert explains prompt-answer data formatting and the role of grader models and LLM judges in reward shaping. The segment is purely instructional.13:41–16:08 · Guest teaching 0/10 Empirical Evaluation Results and Concluding Remarks Lambert shares empirical evaluation graphs from AI2 research and invites the audience to ask questions. There is no host involvement.0:57–3:08 · Guest disagreement 3/10 Defining Reasoning and Challenging Human-Centric Benchmarks Lambert delivers a solo talk challenging standard human-centric definitions of reasoning, criticizing public discourse around LLM limitations as ridiculous. The format is a monologue with zero host participation.3:10–5:16 · Guest disagreement 2/10 Chain of Thought as Intermediate Compute Lambert breaks down how chain-of-thought acts as intermediate variable compute in forward-pass token streams. The monologue continues without any host dialogue.5:19–8:45 · Guest disagreement 3/10 Decoding OpenAI's o1 and Open Replications Lambert dispels speculative community theories about OpenAI o1 utilizing PRMs or MCTS, arguing it is fundamentally massive RL on verifiable outcomes. As a solo conference presentation, host scores remain zero.8:50–11:16 · Guest disagreement 1/10 The Emergence of Reinforcement Fine-Tuning Lambert explains the mechanics of reinforcement fine-tuning (RFT), highlighting how small amounts of RL on top of strong base models avoid performance degradation. No host is present.11:20–13:39 · Guest disagreement 1/10 Data Formats and Grader Models in RL Lambert explains prompt-answer data formatting and the role of grader models and LLM judges in reward shaping. The segment is purely instructional.13:41–16:08 · Guest disagreement 1/10 Empirical Evaluation Results and Concluding Remarks Lambert shares empirical evaluation graphs from AI2 research and invites the audience to ask questions. There is no host involvement.0:57–3:08 · The hosts pushing back 0/10 Defining Reasoning and Challenging Human-Centric Benchmarks Lambert delivers a solo talk challenging standard human-centric definitions of reasoning, criticizing public discourse around LLM limitations as ridiculous. The format is a monologue with zero host participation.3:10–5:16 · The hosts pushing back 0/10 Chain of Thought as Intermediate Compute Lambert breaks down how chain-of-thought acts as intermediate variable compute in forward-pass token streams. The monologue continues without any host dialogue.5:19–8:45 · The hosts pushing back 0/10 Decoding OpenAI's o1 and Open Replications Lambert dispels speculative community theories about OpenAI o1 utilizing PRMs or MCTS, arguing it is fundamentally massive RL on verifiable outcomes. As a solo conference presentation, host scores remain zero.8:50–11:16 · The hosts pushing back 0/10 The Emergence of Reinforcement Fine-Tuning Lambert explains the mechanics of reinforcement fine-tuning (RFT), highlighting how small amounts of RL on top of strong base models avoid performance degradation. No host is present.11:20–13:39 · The hosts pushing back 0/10 Data Formats and Grader Models in RL Lambert explains prompt-answer data formatting and the role of grader models and LLM judges in reward shaping. The segment is purely instructional.13:41–16:08 · The hosts pushing back 0/10 Empirical Evaluation Results and Concluding Remarks Lambert shares empirical evaluation graphs from AI2 research and invites the audience to ask questions. There is no host involvement.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 1:40 Dismissing reasoning skepticism as ridiculous

Lambert forcefully rejects skepticism claiming language models cannot reason, calling debates over human-like definitions unnecessary and ridiculous.

Hardest push from the hosts ▶ 0:07 No host pushback (Monologue)

The episode consists entirely of a solo talk given by Lambert at NeurIPS, resulting in zero host pushback across the recording.

Biggest teaching moment ▶ 5:30 Demystifying o1 architecture against community hype

Lambert educates the room by debunking widespread speculation regarding PRMs and MCTS, explaining that o1 relies on large-scale RL with verifiable outcomes.

The host holds their own ▶ 0:07 No host participation (Monologue)

Because the audio is a standalone presentation without an active interviewer, there are no host counterarguments or demonstrations of expertise.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Defining Reasoning and Challenging Human-Centric Benchmarks 0030 Lambert delivers a solo talk challenging standard human-centric definitions of reasoning, criticizing public discourse around LLM limitations as ridiculous. The format is a monologue with zero host participation.
Chain of Thought as Intermediate Compute 0020 Lambert breaks down how chain-of-thought acts as intermediate variable compute in forward-pass token streams. The monologue continues without any host dialogue.
Decoding OpenAI's o1 and Open Replications 0030 Lambert dispels speculative community theories about OpenAI o1 utilizing PRMs or MCTS, arguing it is fundamentally massive RL on verifiable outcomes. As a solo conference presentation, host scores remain zero.
The Emergence of Reinforcement Fine-Tuning 0010 Lambert explains the mechanics of reinforcement fine-tuning (RFT), highlighting how small amounts of RL on top of strong base models avoid performance degradation. No host is present.
Data Formats and Grader Models in RL 0010 Lambert explains prompt-answer data formatting and the role of grader models and LLM judges in reward shaping. The segment is purely instructional.
Empirical Evaluation Results and Concluding Remarks 0010 Lambert shares empirical evaluation graphs from AI2 research and invites the audience to ask questions. There is no host involvement.

Statements from this episode (11)

Opinion
Lambert: AI is developing new reasoning modes that look less human
“I think a big Trend of the year is that we're seeing new types of language model reasoning that look less human and that Can be good for kind of separating the discourse for expecting a really narrow type of behaviors.”
Nathan Lambert Jan 2, 2025 ▶ 2:58
Insight
Lambert: Autoregressive LLMs lack explicit internal structures for intermediate state
“Language models have no ability to do this. They are. Kind of per token computation devices where each token is outputted after doing this forward pass and within that there's no explicit structure to hold these intermediate states.”
Nathan Lambert Jan 2, 2025 ▶ 3:37
Opinion
Lambert: OpenAI's o1 uses token streams as intermediate state compute
“Why oh, one is exciting is because it's a new type of language models that are going to maximize on this view of reasoning, which is that chain of thought in kind of a forward stream of tokens can actually do a lot to achieve better outcomes when you're doing …”
Nathan Lambert Jan 2, 2025 ▶ 4:42
Insight
Lambert: OpenAI o1 uses large-scale RL on verifiable outcomes, not MCTS
“You should take open AI at their face value, which they are doing very large scale RL on the verifiable outcomes is what I've added, especially in context of the RL API that they've released, which I'll talk about more, but most of the reasons to believe in mo…”
Nathan Lambert Jan 2, 2025 ▶ 5:23
Assertion Not checkable as stated
Lambert: Early DeepSeek and Qwen reasoning models are substantially narrower than o1
“And I think that these models are really substantially narrower than these full O-one models from OpenAI. So OpenAI is, if you use O-one, you can do it for a lot more tasks. If you use, like I was using the DeepSeq model, and it's supposed to be for math or co…”
Nathan Lambert Jan 2, 2025 ▶ 6:26
Prediction Not checkable as stated
Lambert: Open community will eventually match OpenAI's large-scale RL infrastructure
“And this is something that these early relative models are not going to be doing because we don't like, no one has this infrastructure like open AI does. It'll take a while to do that, but people will make it.”
Nathan Lambert Jan 2, 2025 ▶ 8:34
Opinion
Lambert: Reinforcement fine-tuning will succeed where answer correctness matters over style
“It is just a new paradigm for fine tuning, and I have seen some of this work, and I'm pretty Optimistic that it'll work for kind of kind of really specific capabilities where answers matter rather than features in your style of text mattering.”
Nathan Lambert Jan 2, 2025 ▶ 9:18
Assertion Supported
Lambert: Reinforcement fine-tuning requires only dozens of labeled samples
“This reinforcement fine tuning does many passes over the data, which is why they can say you only need dozens of labeled samples to actually learn from it, which is very different than. Previous training regimes”
Nathan Lambert Jan 2, 2025 ▶ 9:57
Assertion Supported
Lambert: Llama 3.1 math evals rely on SymPy and LLM judges
“Lama, 3.1 details their vows for math. They use both SIM by a Python process or Python package for extraction and it's a judge to extract their answers for math.”
Nathan Lambert Jan 2, 2025 ▶ 12:38
Prediction Not checkable as stated
Lambert: Open judge models will become core open RL infrastructure
“We already have a bunch of open models that are doing like judge of models and Prometheus and other things that are designed specifically for LM as a judge. And I see that continuing to just become part of this kind of open RL infrastructure.”
Nathan Lambert Jan 2, 2025 ▶ 13:25
Disclosure
Lambert: AI2 received early industry tip on reinforcement fine-tuning
“We got a tip from a industry lab member to do this a few months early. So we got a head start”
Nathan Lambert Jan 2, 2025 ▶ 15:35
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.