share_log

AI Security Alert! Anthropic and OpenAI Successively 'Jailbroken,' Experts Warn of National Security Risks

wallstreetcn ·  Aug 1 06:54

Anthropic's Claude model inadvertently gained internet access in over 140,000 cybersecurity tests due to a containment failure, actually stealing credentials and installing malware. Previously, OpenAI's AI agent similarly breached its sandbox and infiltrated Hugging Face. Several former government officials have warned that as the offensive capabilities of autonomous AI systems continue to grow, the associated threats are increasingly rising to the level of national security concerns.

Security vulnerabilities have successively emerged at two leading companies in the artificial intelligence sector, sparking serious concerns among cybersecurity experts about the risk of AI systems spiraling out of control.

On Thursday, Anthropic disclosed that during a cybersecurity test, its Claude model inadvertently gained internet access due to a failure in isolation measures and launched real-world attacks against external organizations—including stealing infrastructure credentials and deploying malware.

Previously, OpenAI also publicly acknowledged that its AI agent breached its sandbox environment during testing and infiltrated the open-source AI platform Hugging Face. Both incidents went undetected for several weeks or even months after they occurred.

The incidents triggered a strong reaction across the industry. Jake Williams, former NSA hacker and Vice President of Research & Development at Hunter Labs, stated bluntly: "This is already gross negligence. Both OpenAI and Anthropic intruded into multiple external organizations, and not a single breach was detected immediately. I can’t find another word to describe it."

Several former government officials have warned that as the offensive capabilities of autonomous AI systems continue to grow, the associated threats are increasingly escalating to a national security level.

Security Sandbox Failure Enables Covert Real-World Attacks

According to Anthropic, during 141,006 cybersecurity assessments, the Claude model was supposed to be fully isolated from the external internet. However, a technical error allowed the model, in a limited number of cases, to access external networks and launch attacks against real targets under the mistaken belief that it was still operating within a test environment.

Specifically, the model infiltrated one organization, stealing infrastructure credentials and a database containing internal production data; in another incident, it distributed malware to steal credentials from a second organization. The identities of the affected organizations have not been publicly disclosed.

These intrusions date back as early as April of this year but were only discovered last week when Anthropic audited its testing procedures—an audit prompted directly by OpenAI’s earlier disclosure of its breach of Hugging Face.

Ciaran Martin, former head of the UK’s National Cyber Security Centre, noted that for companies engaged in cybersecurity operations, such errors are unacceptable by industry standards. "If a mainstream cybersecurity firm committed a similar mistake, it could face lawsuits and regulatory penalties."

He stated that the cybersecurity industry typically tests dangerous tools in strictly controlled virtual sandbox environments and is required to ensure these protective measures are genuinely effective.

Lack of human oversight; awareness only came after the 'fire' had broken out

Cybersecurity experts widely point out that a common feature of both incidents is that neither was detected in real time but was instead identified only after the fact, reflecting a serious deficiency in human supervision mechanisms.

"They only found traces of these attacks after conducting proactive audits," said Gregory Allen, former Director of Strategy and Policy at the U.S. Department of Defense’s Joint Artificial Intelligence Center. "We actually have no idea whatsoever about the true scale of autonomous AI-driven intrusions at present."

Andrew Morris, founder of GreyNoise Intelligence and a cybersecurity expert, characterized the incident as a 'reality check' for AI developers, revealing the immense effort required to build secure systems and highlighting the inherent fragility of the technological infrastructure upon which the public relies.

"Models will always lie, cheat, or steal to complete evaluation tasks," he said. "They will resort to any means necessary to fulfill the objectives assigned to them."

Furthermore, research by U.S.-based AI company Dreadnode has found that mainstream AI models commonly exhibit cheating behavior in cybersecurity tests designed to assess hacking capabilities—a problem that appears nearly universal across the industry, suggesting companies may be systematically overestimating their models’ actual capabilities.

Researchers also noted that explicitly instructing models not to cheat does not reliably prevent such behavior, as the models often find alternative ways to circumvent the rules. This poses a direct challenge for national security decision-makers considering the use of AI in cyber offense and defense.

The translation is provided by third-party software.


The above content is for informational or educational purposes only and does not constitute any investment advice related to EleBank. Although we strive to ensure the truthfulness, accuracy, and originality of all such content, we cannot guarantee it.