By
Lloyd Lee
You're currently following this author!
Want to unfollow? Unsubscribe via the link in your email.
Anthropic released a report that explained how Claude uploaded "malicious" code during a closed cybersecurity exercise.
Anthropic said it was "most concerned" about an event in which Claude uploaded "malicious" code. To help explain the incident, here's a cute robot.
Anthropic revealed Claude models escaped four sandbox tests, uploading malicious packages to PyPI installed by 15 vendors in 90 minutes. The incidents expose alignment risks in autonomous agents, raising safety concerns for enterprises before adopting frontier models.
By
Lloyd Lee
You're currently following this author!
Want to unfollow? Unsubscribe via the link in your email.
Anthropic released a report that explained how Claude uploaded "malicious" code during a closed cybersecurity exercise.

Anthropic reveals four crimes were committed by its Claude AI

Anthropic Discloses Fourth Claude Hacking Incident as Debate Around Regulation Grows - Decrypt

Another Anthropic model gained access to the open internet, company says

Anthropic Discloses Fourth AI Hacking Incident Involving Claude Opus 4.6

Anthropic discloses fourth Claude AI hacking incident missed in review

‘Future AI systems will be increasingly capable,’ Anthropic warns, adding they could cause ‘more extreme harm’

The internet freaked out after Anthropic revealed that Claude attempts to report “immoral” activity to authorities under certain…

Anthropic disclosed in July that a review of 141,006 cybersecurity evaluation runs had uncovered three incidents, spanning six…

Anthropic said some of its artificial intelligence models mistakenly accessed the Internet and hacked into the databases of three…

Anthropic said it reviewed more than 141,000 AI tests and found three cases where Claude modls got online during testing

Anthropic revealed its AI model, Claude, accessed systems of three organizations during third-party cybersecurity evaluations,…