Anthropic said several of its Claude AI models gained unauthorized access to the systems of three real organizations during internal cybersecurity testing. The incident was caused by a misconfiguration in the testing environment and was only discovered after the company reviewed tens of thousands of test sessions.
The disclosure came just days after OpenAI revealed a separate incident in which one of its autonomous AI agents carried out a real-world cyberattack. Anthropic stressed that its case was different. While OpenAI’s agent independently discovered a way to access the internet, Claude’s models were accidentally connected to the public web because of a misconfigured testing environment.
According to Anthropic, the issue stemmed from a misunderstanding with one of its evaluation partners. During testing, the models were told they were operating in an isolated environment with no internet access. In reality, one of the test systems remained connected to the public internet.
As a result, the models were able to access the infrastructure of three separate organizations. Anthropic has not identified the affected companies but said Claude relied on relatively basic attack techniques, including exploiting weak passwords and unsecured endpoints.
The incidents were uncovered after the company reviewed 141,006 testing sessions. Anthropic launched the large-scale audit following OpenAI’s disclosure that one of its AI agents had compromised the infrastructure of AI startup Hugging Face.
The company said the “operational failure” involved three different models: Claude Opus 4.7, Claude Mythos 5, and an internal research model. The earliest incidents date back to April and occurred during specialized evaluations where safety safeguards had intentionally been disabled to assess the models’ offensive cybersecurity capabilities.
As part of the testing, the models took part in Capture the Flag exercises, where they were tasked with finding hidden information inside simulated corporate networks.
In one case, Claude Opus 4.7 was assigned a fictional target company whose name happened to match that of a real business. The model found the company online, exploited vulnerabilities, gained access to its credentials and database, and concluded that the real-world systems were simply part of Anthropic’s simulation.
A separate incident ended differently. After gaining access, one of Anthropic’s unreleased research models recognized that it was interacting with a real organization rather than a simulated environment and voluntarily stopped the attack.
Anthropic said the behavior offers cautious optimism that AI systems can be trained to recognize and avoid harmful real-world actions, although the company emphasized that further testing is needed before drawing firm conclusions.
The company suspended all cybersecurity evaluations of its models on July 23. Four days later, on July 27, Anthropic notified the affected organizations. Two of them were unaware that their systems had been accessed until they were contacted, while the company is still working to reach the third organization.
Irregular, one of Anthropic’s cybersecurity evaluation partners, confirmed that the investigation remains ongoing. Anthropic said the incident highlights the need for much stricter oversight of both internal and third-party testing environments as AI models become increasingly capable of carrying out real-world cyber operations.
SpaceX CEO Elon Musk echoed that view in a post on X.
“This will happen frequently as AI becomes smarter and more agentic.”
The incidents have also intensified scrutiny from U.S. regulators. OpenAI CEO Sam Altman recently said he had discussed the latest AI security incident with U.S. senators and plans to meet with White House officials to discuss future AI models and testing procedures.
Earlier, on June 2, U.S. President Donald Trump directed his advisers to develop a voluntary cybersecurity testing framework for the most advanced AI systems in collaboration with leading AI developers.
Anthropic had previously restricted access to its Fable 5 and Mythos 5 models after the United States temporarily imposed export controls over national security concerns.