reinforcement learning

3 statements across 1 episodes · 2 bullish · 0 bearish · 1 people on the record · first statement Jun 18, 2026 by Jonathan Siddharth · across every show →

Everything said about reinforcement learning, oldest first

Jun 18, 2026 positive
Assertion Supported
Siddharth: AI compute is shifting significantly toward post-training reinforcement learning
“In the past, it was a lot of the compute went into pre-training. Now a lot of compute goes into reinforcement learning in post-training as well. Especially after O-one came out and DeepSeek came out.”
Jonathan Siddharth Jun 18, 2026 ▶ 52:36 Why Coding is the Fastest Path to AGI | Turing CEO Jonathan Siddharth
Jun 18, 2026
Insight
Siddharth: Non-binary knowledge work requires rubric-based AI evaluation
“Like with code or with math, it's relatively more binary, easy to verify. But how do you verify the quality of a board deck? Yeah. It's a, you have to be, you have to have like a good rubric based evaluator.”
Jonathan Siddharth Jun 18, 2026 ▶ 16:27 Why Coding is the Fastest Path to AGI | Turing CEO Jonathan Siddharth
Jun 18, 2026 positive
Insight
Siddharth: Verifiability makes coding ideal for reinforcement learning improvements
“Coding is one of those areas where, because it's verifiable, I think that there is a good path to using reinforcement learning to improve coding models quickly.”
Jonathan Siddharth Jun 18, 2026 ▶ 21:44 Why Coding is the Fastest Path to AGI | Turing CEO Jonathan Siddharth
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.