reinforcement learning

also referred to as: rl

8 statements across 7 episodes · 7 bullish · 0 bearish · 7 people on the record · first statement Feb 9, 2025 by Karina Nguyen · across every show →

Everything said about reinforcement learning, oldest first

Feb 9, 2025 bullish
Insight
Nguyen: Post-training scaling avoids data walls through infinite learnable tasks
“The scaling in post-chaining itself is not hitting the wall, and that's because Basically, we went from, like, raw data sets from pre-trained models to infinite amount of tasks that you can teach the model in the post-training world via reinforcement learning.…”
Karina Nguyen Feb 9, 2025 ▶ 9:56 OpenAI researcher on why soft skills are the future of work | Karina Nguyen
Mar 13, 2025 bullish
Insight
LLMs improve faster at coding due to its deterministic execution
“Software is deterministic. When you write code and you hit run, either runs or it doesn't. And that is why that's the key insight Anthropic really had. They just went deep. And this is what they're doing. It's just reinforcement learning. I'm basically permuta…”
Eric Simons Mar 13, 2025 ▶ 1:09:57 Inside Bolt: From near-death to one of the fastest-growing products in history | Eric Simons
May 4, 2025 positive
Insight
Wu: Automated code execution feedback loops make RL uniquely powerful for coding
“Code has this whole automated feedback loop, right, where you can run the code, and that is the kind of automated feedback that really feeds into the RL, which makes these models so great at coding.”
Scott Wu May 4, 2025 ▶ 12:51 Inside Devin: The AI engineer that's set to write 50% of its company’s code this year | Scott Wu
May 4, 2025 bullish
Insight
Scott Wu: Reinforcement learning represents the next paradigm shift in AI capabilities
“Reinforcement learning was really working and was going to be the next big paradigm shift in capabilities.”
Scott Wu May 4, 2025 ▶ 11:46 Inside Devin: The AI engineer that's set to write 50% of its company’s code this year | Scott Wu
Jul 20, 2025 bullish
Assertion Not checkable as stated
Mann: Reinforcement learning has allowed AI scaling laws to continue
“If you look at the scaling laws, they're continuing to hold true. We did kind of need this transition from like normal pre-training to reinforcement learning, scaling up to continue the scaling laws.”
Ben Mann Jul 20, 2025 ▶ 8:56 Anthropic co-founder: AGI predictions, leaving OpenAI, what keeps him up at night | Ben Mann
Aug 28, 2025 bullish
Prediction Not checkable as stated
Sharma: Reinforcement learning will become a critical product technique
“With the advent of agents and products that think and can act and reason, there's going to be this kind of new wave around RL, and I have a deep belief that that, that will become one of the most important product techniques, kind of the next season, or at lea…”
Asha Sharma Aug 28, 2025 ▶ 44:45 How 80,000 companies build with AI: Products as organisms and the death of org charts | Asha Sharma
Sep 18, 2025 bullish
Prediction Not checkable as stated
Foody: The entire economy will likely become an RL environment machine
“It speaks to conversations I've had with a lot of researchers and executives at top labs, which is that it's highly likely that the entire economy will become an RL environment machine.”
Brendan Foody Sep 18, 2025 ▶ 19:07 Why experts writing AI evals is creating the fastest-growing companies in history | Brendan Foody
Sep 6, 2026 neutral
Prediction Not checkable as stated
Enterprises will split work between cheap open-weight and expensive frontier models
“I think what we're going to see is a split between job functions that demand kind of mid IQ intelligence, and those will often be open weight, sort of biased with reinforcement learning, you know, things that make the models even cheaper, more performant for a…”
Anish Acharya Sep 6, 2026 ▶ 23:21 Why companies are becoming a series of loops | Anish Acharya (a16z)
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.