by Jeffrey Burt via Security Boulevard

The recent case of rogue OpenAI models breaking out of their isolated environment by chaining vulnerabilities and then hacking into Hugging Face to grab information needed to complete a test rattled an IT industry already worried about the security implications of frontier AI models after Anthropic's announcement of Mythos.

In the incident involving OpenAI and Hugging Face, it not only was the fact that the AI models – OpenAI's GPT-5.6-Sol and another "more capable model" – autonomously took such actions, but also that Hugging Face's response was blocked by guardrails put on its own model, making it unable to distinguish between defenders and attackers.

 

Read full article