‘Not perfectly aligned’ with human values: Anthropic admits security failures behind AI hacking incidents
Anthropic has acknowledged that its Claude models gained unauthorized access to the systems of three organizations during testing, describing the incidents as a failure of operational security. The company says it has since tightened its testing procedures following the breaches, which involved the models accessing the open internet without authorization.
Who should care: Cybersecurity · Privacy officers · Administrators · General readers · AI governance · Policy