Siddharth: Verifiable domains allow self-play reinforcement learning to replace RLHF
Jonathan Siddharth · Inside The $2.2B AI Research Accelerator | Turing · Sourcery with Molly O'Shea · Oct 10, 2025 · at 24:24
Turing CEO Jonathan Siddharth explains the shift in frontier AI model post-training toward automated verifiers and experiential self-play in domains like math and coding.
“Now, for these verifiable domains like coding and math, instead of doing reinforcement learning with human feedback, you can do reinforcement learning. Because you can automatically check when you got the correct answer or not in these verifiable domains. And with reinforcement learning, the w the way you teach the model is not through imitation learning, where you're learning from humans, but it's through experiential learning, where you create this environment with prompts and verifiers, and you create basically a mini world model, where as long as the agent gets the right result, ah, that particular trajectory that the agent took to come to the right answer gets reinforced, and when the agent got to a result that didn't get the right answer that doesn't get reinforced. So the model kind of learns from its own experience. It's a form of self-play, which I think is really, really cool.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →