A little over a week after OpenAI said that its rogue AI agent accidentally hacked Hugging Face, Anthropic is disclosing three “incidents” where a Claude model, during cybersecurity evaluations, was inadvertently able to access the internet due to a misconfiguration and “gained unauthorized access to the production infrastructure of three different organizations.” Anthropic discovered the intrusions after reviewing its cybersecurity evaluation transcripts in the wake of OpenAI’s disclosure. [Link: Investigating three real-world incidents in our cybersecurity evaluations | https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals | Anthropic]

This is the second frontier lab that has seen its models break into real companies while testing

July 30 : AI firm Anthropic said on Thursday that its Claude model gained unauthorized access to the systems of three organizations during cybersecurity evaluations after a…