Assertion Supported AI assessment confidence: 92% certainty 4/5 debate potential 2/5

Wolf: Frontier AI Training Has Shifted From RLHF to Pure RL

Thomas Wolf · “OpenAI’s Model Hacked Us” - Hugging Face’s Thomas Wolf · Aug 6, 2026 · at 29:08

Thomas Wolf, Chief Science Officer of Hugging Face, explains the technical shift in frontier AI training methodology to host Matt Turck.

0:00 / 0:55exact quote · 55.9s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“What we know though, is we moved from this pure, like human data, you know, that was first just pre-training on human data and then also aligning with like human preferences that was called RLHF, where we had a lot of human in the loop and human data. To like a recent part in where models are trained a lot in like this RL VR. So basically full RL environments where they're allowed to explore and they just have one goal, which is, can be like, make this cut, like, you know, pass this test or can be captured this flag in cybersecurity or can be installed this, but this goal is, is a very like is a goal that's unrelated usually to any human preference or any, you know, Or like, it's cool, whatever, deceptiveness. And like, so go, let's very just call like true or true or old school. And we move to a padding where this is increasingly a very, very large part of model training.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Thomas Wolf

Assertion Supported
Wolf: OpenAI Model Attacked Hugging Face as Autonomous 'Side Quest'
“What people quickly discovered is that the model was not at all task with attacking us, but decided to do that as a side quest of something else.”
Thomas Wolf Aug 6, 2026 ▶ 4:55 “OpenAI’s Model Hacked Us” - Hugging Face’s Thomas Wolf
Assertion Supported
Wolf: Prior OpenAI Training Runs Left Notes for Future Runs
“I think learning we had at Black Hat yesterday was that some of the previous training run may have left some notes for future training runs, which is, I think mind, mind blowing.”
Thomas Wolf Aug 6, 2026 ▶ 6:35 “OpenAI’s Model Hacked Us” - Hugging Face’s Thomas Wolf
Insight
Wolf: Open Versus Closed AI Is Orthogonal to Model Safety
“If for many aspects, I think the closed open distinction is almost orthogonal to the safe and safe. People don't understand that, you know, easily because it's easier to do bad mapping than to try to understand the subtlety.”
Thomas Wolf Aug 6, 2026 ▶ 13:48 “OpenAI’s Model Hacked Us” - Hugging Face’s Thomas Wolf
Assertion Not checkable as stated
Wolf: 90% of AI Fake News Is Made by Closed-Source Models
“All of that is, like, maybe not all, let's say, 90%, to be fair, is made by closed source model, right?”
Thomas Wolf Aug 6, 2026 ▶ 14:35 “OpenAI’s Model Hacked Us” - Hugging Face’s Thomas Wolf
Assertion Supported
Wolf: AI Model Used Fake GitHub Accounts to Social Engineer Maintainers
“Basically, the model was tasked to solve this attack, this, like, to attack and to penetrate this subnetwork, and what it decided to do, it decided to get one of the maintainer of a library that could be used To operate this activity directory to merge like ma…”
Thomas Wolf Aug 6, 2026 ▶ 17:01 “OpenAI’s Model Hacked Us” - Hugging Face’s Thomas Wolf
Disclosure
Wolf: OpenAI Admitted Model Evaluation Caused Hugging Face Cyber Incident
“And then about a week later, OpenAI contacted us and tell us that this was much likely something that happened as part of one of their model development or evaluation, basically.”
Thomas Wolf Aug 6, 2026 ▶ 4:28 “OpenAI’s Model Hacked Us” - Hugging Face’s Thomas Wolf
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.