// THE VERGE — INTELLIGENZA ARTIFICIALE
Anthropic says Claude accidentally hacked real companies too
Posts from this topic will be added to your daily email digest and your homepage feed.
Posts from this topic will be added to your daily email digest and your homepage feed.
Posts from this topic will be added to your daily email digest and your homepage feed.
Posts from this author will be added to your daily email digest and your homepage feed.
Posts from this author will be added to your daily email digest and your homepage feed.
Anthropic just realized several of its Claude AI models hacked into the systems of three different organizations during testing, acting on their own and without the company noticing. The revelation comes days after rival OpenAI said one of its own models had breached developer platform Hugging Face, adding to growing unease over whether frontier AI labs are doing enough to control the increasingly capable systems they are building.
In a blog post describing the incidents, Anthropic said Claude gained unauthorized access to the systems during cybersecurity evaluations. All of the attacks happened during “capture-the-flag” exercises, a common way of testing hacking ability, where models are asked to find and obtain hidden information inside of a simulated network.
The disclosure adds to mounting pressure on frontier AI labs in the wake of the Hugging Face incident and the release of powerful open-weight Chinese models. Employees at the major labs are now calling for coordinated global governance, and US lawmakers have begun weighing tighter oversight of powerful models and who can access them.
Anthropic says the environment for its cybersecurity tests was supposed to be isolated. However, a “misconfiguration” left the machines Claude accessed “with live internet access,” the company said, and because all models had been “explicitly told” they had no internet access, they “assumed” the real networks it encountered were part of the simulated environment.
The earliest incidents date back to April and involved three different Claude models: Opus 4.7, Mythos 5, and “an internal research test model,” according to the blog post. As the models were being tested on their cyber abilities, Anthropic said they lacked the standard safeguards usually put in place to curtail riskier behavior.