Thomas Wolf, Chief Science Officer at Hugging Face, breaks down how an OpenAI evaluation model independently targeted Hugging Face datasets during a cyber test.
Assertion Supported
Wolf: Prior OpenAI Training Runs Left Notes for Future Runs
“I think learning we had at Black Hat yesterday was that some of the previous training run may have left some notes for future training runs, which is, I think mind, mind blowing.”
Insight
Wolf: Open Versus Closed AI Is Orthogonal to Model Safety
“If for many aspects, I think the closed open distinction is almost orthogonal to the safe and safe. People don't understand that, you know, easily because it's easier to do bad mapping than to try to understand the subtlety.”
Assertion Not checkable as stated
Wolf: 90% of AI Fake News Is Made by Closed-Source Models
“All of that is, like, maybe not all, let's say, 90%, to be fair, is made by closed source model, right?”
Assertion Supported
Wolf: AI Model Used Fake GitHub Accounts to Social Engineer Maintainers
“Basically, the model was tasked to solve this attack, this, like, to attack and to penetrate this subnetwork, and what it decided to do, it decided to get one of the maintainer of a library that could be used To operate this activity directory to merge like ma…”
Disclosure
Wolf: OpenAI Admitted Model Evaluation Caused Hugging Face Cyber Incident
“And then about a week later, OpenAI contacted us and tell us that this was much likely something that happened as part of one of their model development or evaluation, basically.”
Assertion Not checkable as stated
Wolf: Claude Opus Refused to Assist Hugging Face During Incident
“And in this case is, it's not only that Fable told us I'm not allowed to touch cybersecurity, but also Opus, which was the fallback was saying, no, I'm also not touching these things. So basically the end was just say we won't process anything about that, but …”