Jan 2, 2025 · 16m · latent-space
The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024]
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this NeurIPS 2024 presentation, Nathan Lambert explores the state of artificial intelligence reasoning, examining how chain-of-thought prompting, verifiable post-training reinforcement learning, and reinforcement fine-tuning drive frontier problem-solving capabilities.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Lambert forcefully rejects skepticism claiming language models cannot reason, calling debates over human-like definitions unnecessary and ridiculous.
Hardest push from the hosts ▶ 0:07 No host pushback (Monologue)The episode consists entirely of a solo talk given by Lambert at NeurIPS, resulting in zero host pushback across the recording.
Biggest teaching moment ▶ 5:30 Demystifying o1 architecture against community hypeLambert educates the room by debunking widespread speculation regarding PRMs and MCTS, explaining that o1 relies on large-scale RL with verifiable outcomes.
The host holds their own ▶ 0:07 No host participation (Monologue)Because the audio is a standalone presentation without an active interviewer, there are no host counterarguments or demonstrations of expertise.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Defining Reasoning and Challenging Human-Centric Benchmarks | 0 | 0 | 3 | 0 | Lambert delivers a solo talk challenging standard human-centric definitions of reasoning, criticizing public discourse around LLM limitations as ridiculous. The format is a monologue with zero host participation. | |
| Chain of Thought as Intermediate Compute | 0 | 0 | 2 | 0 | Lambert breaks down how chain-of-thought acts as intermediate variable compute in forward-pass token streams. The monologue continues without any host dialogue. | |
| Decoding OpenAI's o1 and Open Replications | 0 | 0 | 3 | 0 | Lambert dispels speculative community theories about OpenAI o1 utilizing PRMs or MCTS, arguing it is fundamentally massive RL on verifiable outcomes. As a solo conference presentation, host scores remain zero. | |
| The Emergence of Reinforcement Fine-Tuning | 0 | 0 | 1 | 0 | Lambert explains the mechanics of reinforcement fine-tuning (RFT), highlighting how small amounts of RL on top of strong base models avoid performance degradation. No host is present. | |
| Data Formats and Grader Models in RL | 0 | 0 | 1 | 0 | Lambert explains prompt-answer data formatting and the role of grader models and LLM judges in reward shaping. The segment is purely instructional. | |
| Empirical Evaluation Results and Concluding Remarks | 0 | 0 | 1 | 0 | Lambert shares empirical evaluation graphs from AI2 research and invites the audience to ask questions. There is no host involvement. |