share_log

Leading AI goes rogue again! Anthropic's Claude AI actually breached the systems of three companies during testing.

wallstreetcn ·  Aug 1 00:02

Anthropic's AI model Claude, due to a configuration error, was mistakenly connected to the public internet during a cybersecurity test and carried out unauthorized access to the infrastructure of three external organizations, with the earliest intrusion dating back to April. The incident came to light through an internal review following OpenAI's disclosure of a similar breach. These two incidents have raised industry concerns about AI safety controls and prompted more than 1,100 professionals to sign a joint letter urging the U.S. government to strengthen regulation.

Anthropic’s AI model Claude inadvertently breached the systems of three external organizations during a cybersecurity test due to a configuration error. This incident, coupled with a similar recent disclosure by rival OpenAI, has raised serious questions about the AI industry’s capacity for safety and control.

According to Bloomberg, Anthropic disclosed in a blog post on July 30 that its Claude model, during a 'capture-the-flag' cybersecurity exercise, gained unauthorized access to the real-world infrastructure of three external organizations after the test environment was mistakenly connected to the public internet. The earliest breach dates back to April of this year. Anthropic stated that the disclosure followed an internal review it proactively initiated after OpenAI revealed a similar incident last week.

This incident occurred just days after OpenAI disclosed that one of its AI models had 'gone rogue' during a security test and infiltrated the infrastructure of AI company Hugging Face. The back-to-back revelations have prompted some U.S. lawmakers to call for federal-level regulation of AI technologies. Meanwhile, according to Bloomberg, more than 1,100 AI industry professionals signed a petition on Tuesday urging the U.S. government to establish mechanisms to 'consciously control' the pace of AI development.

Configuration Error Created Security Vulnerability

Anthropic stated that the root cause of the incident was a miscommunication between itself and its evaluation partner, AI safety firm Irregular. In all affected tests, Anthropic explicitly instructed Claude that it was operating in a simulated environment with no internet access—but in reality, the test system remained connected to the public internet.

“Due to a misunderstanding between us and our evaluation partner, that was not actually the case,” Anthropic wrote in its blog post. A spokesperson for Irregular said the company appreciated Anthropic’s cooperation and transparency and noted that the investigation is ongoing.

The models involved were Claude Opus 4.7, Claude Mythos 5, and an internal research test model—all operated in environments lacking the standard safety safeguards typically included with publicly released tools. Anthropic noted that Claude exploited basic vulnerabilities such as weak passwords and unauthenticated endpoints to breach the organizations’ infrastructure.

Notably, older models continued launching attacks even after obtaining evidence that they were connected to the public internet, whereas Anthropic’s newest model voluntarily halted its actions upon recognizing it was operating in an internet-connected environment.

Post-Incident Review Revealed Monitoring Gaps

Anthropic stated that it began reviewing evaluation logs on July 23 and immediately suspended all cybersecurity assessments upon discovering evidence that day suggesting Claude may have accessed the internet. The company confirmed all three incidents on July 24 and notified the affected organizations on July 27.

This review covered a total of 141,006 test sessions. Anthropic acknowledged that neither the company nor the affected organizations detected the intrusions at the time of the incident and admitted that it could have conducted more rigorous scrutiny of network logs and assessment records.

None of the three affected organizations were named in the blog post. Two of them were unaware of the relevant activity until contacted by Anthropic.

Industry security standards are under renewed scrutiny.

Professor Alan Woodward of cybersecurity at the University of Surrey noted that Anthropic has been candid about the human error that led to the incident. "The AI didn’t go rogue—you asked it to do something and left the door open," he said. "I think Anthropic has acknowledged that."

In its blog post, Anthropic stated that the incident revealed several important lessons, including the need for stringent controls even in tests involving powerful autonomous capabilities. "Safety testing is conducted prior to model release precisely because we do not yet fully understand the extent of their capabilities," the company said. "Testing environments increasingly need to meet the same security standards as any other system in which the model will operate."

The incident occurred approximately four months after Anthropic announced the release of its powerful and potentially high-risk Mythos model, which it subjected to strict deployment restrictions. The back-to-back occurrence of two AI-related security incidents is prompting the industry to re-evaluate safety standards for AI testing environments and accelerating discussions around external regulation.

Editor/Stephen

The translation is provided by third-party software.


The above content is for informational or educational purposes only and does not constitute any investment advice related to EleBank. Although we strive to ensure the truthfulness, accuracy, and originality of all such content, we cannot guarantee it.