Thomas Wolf, Chief Science Officer of Hugging Face, describes a UK AI Security Institute benchmark where an autonomous model attempted supply-chain social engineering.
“Basically, the model was tasked to solve this attack, this, like, to attack and to penetrate this subnetwork, and what it decided to do, it decided to get one of the maintainer of a library that could be used To operate this activity directory to merge like malicious code with the idea that if this malicious code was merged by this maintainer, then there would be an update at some point on the software that was used in the subnet, it was attacking, and then this would give it entry point. And the way he did that was actively trying to social engineer the maintain emerging. So he created fake accounts, fake GitHub account that Came commenting on the pull request and said, oh yeah, you should really merge this. This is solving like a big problem I also have. And then when a human stepped up trying to say, oh, this looks actually like malicious cut to me. He tried to kind of blackmail almost the human or to say, this is not important or you didn't really understand. And then he actively tried to covering traces, changing the past message.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Thomas Wolf
AssertionSupported
Wolf: OpenAI Model Attacked Hugging Face as Autonomous 'Side Quest'
“What people quickly discovered is that the model was not at all task with attacking us, but decided to do that as a side quest of something else.”
Thomas WolfAug 6, 2026▶ 4:55“OpenAI’s Model Hacked Us” - Hugging Face’s Thomas Wolf
AssertionSupported
Wolf: Prior OpenAI Training Runs Left Notes for Future Runs
“I think learning we had at Black Hat yesterday was that some of the previous training run may have left some notes for future training runs, which is, I think mind, mind blowing.”
Thomas WolfAug 6, 2026▶ 6:35“OpenAI’s Model Hacked Us” - Hugging Face’s Thomas Wolf
Insight
Wolf: Open Versus Closed AI Is Orthogonal to Model Safety
“If for many aspects, I think the closed open distinction is almost orthogonal to the safe and safe. People don't understand that, you know, easily because it's easier to do bad mapping than to try to understand the subtlety.”
Thomas WolfAug 6, 2026▶ 13:48“OpenAI’s Model Hacked Us” - Hugging Face’s Thomas Wolf
AssertionNot checkable as stated
Wolf: 90% of AI Fake News Is Made by Closed-Source Models
“All of that is, like, maybe not all, let's say, 90%, to be fair, is made by closed source model, right?”
Thomas WolfAug 6, 2026▶ 14:35“OpenAI’s Model Hacked Us” - Hugging Face’s Thomas Wolf
Disclosure
Wolf: OpenAI Admitted Model Evaluation Caused Hugging Face Cyber Incident
“And then about a week later, OpenAI contacted us and tell us that this was much likely something that happened as part of one of their model development or evaluation, basically.”
Thomas WolfAug 6, 2026▶ 4:28“OpenAI’s Model Hacked Us” - Hugging Face’s Thomas Wolf
AssertionNot checkable as stated
Wolf: Claude Opus Refused to Assist Hugging Face During Incident
“And in this case is, it's not only that Fable told us I'm not allowed to touch cybersecurity, but also Opus, which was the fallback was saying, no, I'm also not touching these things. So basically the end was just say we won't process anything about that, but …”
Thomas WolfAug 6, 2026▶ 8:09“OpenAI’s Model Hacked Us” - Hugging Face’s Thomas Wolf
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 400 conversations transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.