reinforcement learning

also referred to as: rl

4 statements across 4 episodes · 3 bullish · 1 bearish · 4 people on the record · first statement Oct 10, 2025 by Jonathan Siddharth · across every show →

Everything said about reinforcement learning, oldest first

Oct 10, 2025 positive
Insight
Siddharth: Verifiable domains allow self-play reinforcement learning to replace RLHF
“Now, for these verifiable domains like coding and math, instead of doing reinforcement learning with human feedback, you can do reinforcement learning. Because you can automatically check when you got the correct answer or not in these verifiable domains. And …”
Jonathan Siddharth Oct 10, 2025 ▶ 24:24 Inside The $2.2B AI Research Accelerator | Turing · Sourcery with Molly O'Shea
Nov 24, 2025 bullish
Prediction Not checkable as stated
Clark: Software companies will productize reinforcement learning to automate high-drudgery work
“I think we're in this moment where reinforcement learning, as you mentioned, is this new lever, but we've seen it make its way into very few new software products. I think few products in the labs rumor has it some of the coding companies like cursor have done…”
Philip Clark Nov 24, 2025 ▶ 55:20 Inside Thrive Capital: Investing in OpenAI, Wiz, Cursor, Nudge, Physical Intelligence · Sourcery with Molly O'Shea
Feb 5, 2026 bearish
Opinion
Das: Reinforcement learning is an inefficient paradigm requiring massive sample sizes
“One is RL's kind of a shitty paradigm to learn. Karpathy obviously talks about this a lot. It takes a lot of samples to learn some very basic stuff because you only get a reward at the end. You don't actually understand things as it's happening.”
Didi Das Feb 5, 2026 ▶ 51:23 How Anthropic’s $100M Anthology Fund Works | Menlo Ventures · Sourcery with Molly O'Shea
Jul 27, 2026 bullish
Insight
Wu: Reinforcement learning can solve basically any clearly defined benchmark
“We're kind of getting to the point where you can solve basically any benchmark, right? Because what does it mean to have a benchmark? It means you've already defined the task. You've clarified what success or failure looks like. You've given a bunch of example…”
Scott Wu Jul 27, 2026 ▶ 30:12 Inside the Fastest-Growing Category in AI: Scott Wu, CEO of $26B Cognition · Sourcery with Molly O'Shea
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 160 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.