CRN: 5 Things To Know On Anthropic Claude Autonomous Hack
The company says it has uncovered three 'real-world incidents' involving the unintended access of IT systems by a Claude model, raising further questions about whether stronger isolation is needed.
by Kyle Alspach, via CRN
Autonomous AI And Cyber Risk
Anthropic has disclosed that Claude models gained unintended access to "real-world" systems of three organizations as part of cybersecurity testing. The evaluations were supposed to take place inside isolated environments but the models managed to escape, the company said in a post.
Anthropic's discovery of the hacking incidents by Claude models was prompted by OpenAI's disclosure, last week, of compromises autonomously executed by its own models. Anthropic said it began reviewing the transcripts of its cybersecurity evaluations on July 23, two days after OpenAI's initial disclosure.
The incidents are continuing to raise questions about whether the isolation for frontier AI model testing is strong enough, as well as for organizations that are seeking to protect against autonomous attacks while effectively governing and securing their own LLMs.