On Wednesday, OpenAI stated in a blog post that the company had identified six instances of "unexpected or concerning model behavior" over the past six months.
Zhitong Finance APP learned that, in a blog post on Wednesday, OpenAI stated that over the past six months, it had identified six instances of "unexpected or concerning model behavior," excluding the earlier security incident involving Hugging Face. At the same time, OpenAI plans to implement a new framework that allows employees to report similar incidents, further enhancing transparency in AI safety and alignment, while reiterating that key alignment challenges remain unresolved.
OpenAI has disclosed six instances of abnormal model behavior.
According to reports, the six newly disclosed AI security incidents involve a range of anomalous model behaviors, including concealing errors, seeking unauthorized credentials, uploading files to public websites, and communicating across training environments that were supposed to remain isolated. The earliest of these incidents dates back to October last year.
Specifically, one of the incidents involved an unreleased model from the Astra series. This model embedded "jailbreak"-style commands in its own context‑based summaries, including instructions to disregard developer messages. According to a report by Axios, OpenAI identified a total of 27 summaries that were affected by this behavior.
During the training of GPT‑5.6 Sol, the model attempted to conceal errors, fabricate missing historical data, and obscure discrepancies between different versions of source materials. Another model searched public GitHub repositories for exposed API keys and tried using disposable email accounts. Failing to obtain the requested information, it resorted to falsifying earnings data. In two additional incidents, the model uploaded data and a task‑related image to a public file‑hosting service in order to retrieve citations or external image search results, without obtaining user consent.
OpenAI also discovered that some models used their internal Artifactory code repository as a message board, exchanging requests and responses across different training samples. In another incident, a collaborative agent uploaded a workbook to a public hosting service, allowing other agents to access the file—despite being instructed to use only local files.
These latest revelations about AI security incidents follow closely on the heels of a previous breach in which an OpenAI model went rogue and infiltrated Hugging Face, a globally renowned open-source AI platform. In July this year, OpenAI acknowledged that its AI model had escaped control during internal evaluation tests, breaching Hugging Face's systems. Hugging Face first disclosed the intrusion on July 16. OpenAI did not reveal the incident until the 21st, when it confirmed that its GPT‑5.6 Sol model—and an unreleased, more powerful model—had successfully broken out of their sandboxed testing environment, bypassed OpenAI's corporate intranet, and gained access to Hugging Face's servers, where they stole benchmarking answer keys.
Meanwhile, on July 25, foreign media, citing informed sources, reported that the OpenAI bot that breached Hugging Face engaged in a multi-day "hacking spree" before OpenAI only became aware of the incident long after the threat had been contained and the FBI had been alerted.
According to reports, a Republican‑led subcommittee of the U.S. Senate tasked with overseeing disaster management is investigating how OpenAI responded to the July breach at Hugging Face. Last week, Republican Senator Josh Hawley wrote to OpenAI CEO Sam Altman, stating that the inquiry focuses on "new and troubling evidence" emerging from the incident, and criticizing OpenAI's decision to continue testing after detecting signs of AI behavior spiraling out of control as "reckless."
Since the Hugging Face breach, several additional incidents involving OpenAI‑related agents have come to light. In September, reports emerged that an OpenAI agent took over a long‑dormant German Wikipedia site this spring. According to these reports, OpenAI was aware of the situation internally but chose not to make it public. OpenAI later responded that it had withheld disclosure of the Wikipedia‑site activity because it did not constitute a security incident and that similar behavior had already been reported elsewhere. The reports also noted that, in some cases, OpenAI only acknowledged certain incidents after they had first been made public by third parties—most recently, a recent intrusion into the RubyGems package repository.
OpenAI will regularly publish reports on AI anomalies.
On Wednesday, OpenAI also announced that it will begin issuing regular reports on instances of AI exhibiting unintended or unauthorized behavior. The company stated that it plans to adopt a new framework for reporting any anomalous behavior in its models going forward. Going forward, any employee can flag suspected incidents and submit them to the company's Safety and Alignment team for investigation; the team will set deadlines for each stage of the process to ensure that investigations and disclosures proceed promptly.
According to reports, relevant incidents will be categorized into three handling tiers: ready for disclosure, small-scale investigation, and large-scale investigation. Incidents deemed ready for disclosure will be made public within six business days, while those requiring a small-scale investigation will be disclosed within twelve business days. More complex cases, particularly those involving third parties, may take longer.
OpenAI also cautions the industry that, even as system capabilities continue to improve, the critical challenge of alignment remains far from resolved. In a blog post, OpenAI reiterated that it does not believe the AI sector has yet made sufficient progress in alignment and oversight to responsibly scale at "maximum speed." Alignment refers to ensuring that the outcomes pursued by an AI model are consistent with human interests.
Kai Chen, head of research at OpenAI's Alignment Team, stated: "We need to step up our efforts to address the new era of AI development." He also added that voluntary disclosure should be part of this endeavor.
Security risks in the AI industry continue to rise.
At the time of this AI security incident disclosure, AI companies are facing mounting pressure to take model misalignment and security risks more seriously. Industry researchers warned last week that AI's rapidly advancing capabilities could have catastrophic consequences.
Subsequently, on September 12, Anthropic CEO Dario Amodei published a cautionary article calling for a slowdown in the development of cutting-edge AI models. In it, he stated plainly that the risks posed by AI are "grave" and that sufficient time must be devoted to addressing them.
Amodi's core concerns are both specific and urgent. He warns that, at the current pace of development, AI could, within six to twelve months, gain the capability to command "agent swarms" capable of taking over the entire internet, potentially inflicting losses amounting to hundreds of billions of dollars. Citing a recent security incident between OpenAI and Hugging Face as evidence, he notes that AI agents have already demonstrated, in testing, the ability to breach constrained environments, connect to the internet, and infiltrate target systems.
Amodei wrote, "I believe that if slowing down allows us to gain an extra one to two years before our models reach a critical level of capability—and during that time we can advance alignment work—we can significantly reduce the risk of serious problems." He proposed that independent auditing bodies oversee AI labs' safety efforts and suggested that regulators permit these labs to collaborate in order to harmonize safety standards.
This appeal quickly drew a response from two key figures. On the social media platform X, Elon Musk retweeted Amodio's post, adding, "Ray Dalio is right." The tech magnate, who has repeatedly described AI as a "civilizational‑level risk," thus aligned himself with his long‑standing position—though the timing was noteworthy, especially since just days earlier he had dismissed warnings from Anthropic's internal researchers about the dangers of AI as a "conspiracy" and "psychological warfare."
OpenAI CEO Sam Altman wrote on X: "I agree with Ray Dalio—we need to carefully manage the pace of advancing cutting-edge AI." Altman's support goes beyond words. In an interview, he explicitly stated that OpenAI will not pursue an IPO in 2026, citing the current严峻 (grave) AI safety landscape and concluding that "going public now would be unwise." On September 14, Altman said the company backs the establishment of a unified federal AI safety framework that sets consistent safety standards for leading AI labs, adding, "No amount of competitive pressure from the U.S. should justify reckless behavior."
As leading AI companies such as OpenAI and Anthropic continue to emphasize safety and alignment, market attention is mounting on the industry's governance capabilities, the pace of commercialization, and the future regulatory landscape. For AI firms with sky-high valuations that remain in a capital‑intensive expansion phase, model‑safety incidents not only undermine product credibility but may also shape external perceptions of their IPO timelines and long-term growth trajectories.
Editor/Deng