by Kyle Alspach, via CRN

Autonomous AI And Cyber Risk

Anthropic has disclosed that Claude models gained unintended access to "real-world" systems of three organizations as part of cybersecurity testing. The evaluations were supposed to take place inside isolated environments but the models managed to escape, the company said in a post.

Anthropic's discovery of the hacking incidents by Claude models was prompted by OpenAI's disclosure, last week, of compromises autonomously executed by its own models. Anthropic said it began reviewing the transcripts of its cybersecurity evaluations on July 23, two days after OpenAI's initial disclosure.

The incidents are continuing to raise questions about whether the isolation for frontier AI model testing is strong enough, as well as for organizations that are seeking to protect against autonomous attacks while effectively governing and securing their own LLMs. 

 

Read full article