Wolf: Claude Opus Refused to Assist Hugging Face During Incident
“And in this case is, it's not only that Fable told us I'm not allowed to touch cybersecurity, but also Opus, which was the fallback was saying, no, I'm also not touching these things. So basically the end was just say we won't process anything about that, but …”
Wolf: Open Versus Closed AI Is Orthogonal to Model Safety
“If for many aspects, I think the closed open distinction is almost orthogonal to the safe and safe. People don't understand that, you know, easily because it's easier to do bad mapping than to try to understand the subtlety.”
Wolf: AI Model Used Fake GitHub Accounts to Social Engineer Maintainers
“Basically, the model was tasked to solve this attack, this, like, to attack and to penetrate this subnetwork, and what it decided to do, it decided to get one of the maintainer of a library that could be used To operate this activity directory to merge like ma…”
Wolf: Frontier AI Training Has Shifted From RLHF to Pure RL
“What we know though, is we moved from this pure, like human data, you know, that was first just pre-training on human data and then also aligning with like human preferences that was called RLHF, where we had a lot of human in the loop and human data. To like …”
Wolf: Bostrom's Paperclip Maximizer Scenario Describes Real AI Incidents
“But definitely it seems like when Ballstrom worked about it in 2003, it seems a little bit like, you know, futuristic, definitely, and maybe something that was like a little bit crazy and just would not happen. But today, I mean, it's pretty clearly something …”
Wolf: GPT-5.6 and Mythos Exhibit Distinctly Different Frontier Behaviors
“But also, we can see that both Frontier model and I take GPT, 5.6 and Mythos doesn't seem to have at all the same type of behaviors. So there is differences here in the effect of, you know, they are not trained exactly the same way and they don't behave the sa…”
Wolf: Post-Training Mitigates Backdoor Risks in Open-Source AI Models
“Right now, if you pre-train and post-train a model for longer, you very likely change quite a lot of the weights and that it has. So I think there's a lot of way to circumvent that which means that at the moment I'm a bit less worried about that than maybe jus…”
Wolf: Local Hosting of Open-Source Models Drives AI Sovereignty
“In OpenSoup Mobile, you can download it. Nobody can, like no country could take it out from you once you download it. You can find it yourself. If you operate it on, on your data center, like local ground data center, I feel like you start to have the beginnin…”
Wolf: Life Science Startups Must Abandon Guardrailed Closed AI Models
“Because of the guardrails and because of the question around biohacking and using this model to generate like the access right now for people just to take it is very, very limited once you want to ask some biology question. And so basically most of the life sc…”
Wolf: I Signed the Open Letter to Pace Automated AI Research
“I did sign this letter.”