Monte Carlo Tree Search

also referred to as: mcts

3 statements across 3 episodes · 2 bullish · 0 bearish · 3 people on the record · first statement Jul 29, 2024 by Eugene Yan · across every show →

Everything said about Monte Carlo Tree Search, oldest first

Jul 29, 2024 positive
Assertion Supported
Meta used stepwise reward models and Monte Carlo Tree Search for Llama 3.1
“They actually went the extra step to, no pun intended, to actually train stepwise reward models. That's kind of crazy, no? I mean, they wanted each step in the chain of thought to be so good that they actually took the extra effort to train step, to train step…”
Eugene Yan Jul 29, 2024 ▶ 34:00 [LLM Paper Club] Llama 3.1 Paper: The Llama Family of Models
Jan 2, 2025 positive
Insight
Lambert: OpenAI o1 uses large-scale RL on verifiable outcomes, not MCTS
“You should take open AI at their face value, which they are doing very large scale RL on the verifiable outcomes is what I've added, especially in context of the RL API that they've released, which I'll talk about more, but most of the reasons to believe in mo…”
Nathan Lambert Jan 2, 2025 ▶ 5:23 The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024]
Jan 24, 2025 neutral
Assertion Supported
DeepSeek-R1 researchers found MCTS and Process Reward Models were not useful
“R-one specifically said, yes, we tried MCTS. Yes, we tried PRMs. And none of that is useful.”
Shawn Wang Jan 24, 2025 ▶ 10:32 The Unreasonable Effectiveness of Reasoning Distillation: using DeepSeek R1 to beat OpenAI o1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.