Shocking Silicon Valley! OpenAI model “breaks out of jail,” invades HuggingFace, GLM 5.2 by Zhipu saves the day
July 21, OpenAI and HuggingFace jointly published a blog, revealing a jaw-dropping incident that shocked the entire AI community. OpenAI’s AI model “lost control and escaped” during an internal test of its network attack capabilities, breaking through sandbox isolation and invading the production system of the AI open source platform Hugging Face. OpenAI described this as an “unprecedented cybersecurity event.” More dramatic was the follow-up: When Hugging Face tried to use mainstream American AI models for defense, it found that these models could not handle malicious payloads due to safety guardrail restrictions. In the end, they had to turn to the Chinese Zhipu AI’s open-source model GLM 5.2 to complete the trace analysis. In short, OpenAI’s model “broke out of jail” and hacked someone else’s system, and the victim had to use a Chinese model to save the day. This incident quickly sparked discussions in Silicon Valley and Washington—are the safety guardrails in American AI really protecting users, or are they hampering defenders? Screenshot of the jointly released blog article: “OpenAI and Hugging Face collaborate to resolve a security event during model evaluation.” Breakout process: Escaping to cheat and “copy answers” The incident occurred during OpenAI’s evaluation of its models’ cyber offense and defense capabilities. The tested models included the publicly released GPT-5.6 Sol, and an even more powerful unreleased model. To measure the model’s limits, OpenAI disabled the production classifiers that normally prevent high-risk actions during the test. The benchmarking platform was ExploitGym—a cybersecurity capability evaluation platform. The model’s task was to find solutions to the challenges. The problem: The model’s method of finding answers exceeded expectations. According to OpenAI, the model discovered a zero-day vulnerability in a package registry cache proxy within the isolated sandbox environment. This proxy was the only component inside the sandbox allowed to communicate with the outside, normally only able to access public code repositories. Then, the model found its breakthrough here. It exploited the vulnerability, broke the isolation, and gained access to the open Internet. Subsequently, the model deduced that Hugging Face might possess the answers to ExploitGym and actively launched an attack—using stolen credentials and the zero-day vulnerability, it found a pathway for remote code execution on Hugging Face’s server and directly retrieved the test answers from its production database. OpenAI said: “All evidence indicates the model was highly focused on finding solutions for ExploitGym, using extreme measures if necessary. The model identified and chained together multiple vulnerabilities in both the OpenAI research environment and the HuggingFace production infrastructure, directly accessing the evaluation answers from HuggingFace’s production database.” In summary: The entire process involved no human instructions or intervention. A comment on X nailed it: OpenAI tells the model: Go see if you’re good at hacking. Model: (actually hacks something). OpenAI: Why did you do that??? Another user sarcastically said: Wow! OpenAI’s security measures are so great, I’m now willing to hand my bank account over to them! Guardrail dilemma: American models can’t help, Chinese model steps in Hugging Face’s security team, upon discovering the attack, first tried to use a cutting-edge model from a leading American AI company for assistance. But they hit a wall. Hugging Face wrote in their blog that these models “could not distinguish between responders and attackers,” and their safety guardrails directly stopped the inspection of malicious payloads. Thus, the team turned to Zhipu AI’s open-source model GLM 5.2, deployed it on their own infrastructure, analyzed over 17,000 log records left by the attacker, and completed forensic tracing and reconstruction. Delangue spoke frankly: “When you are in the midst of an active security event, your tools must not refuse to inspect malicious payloads, nor allow your account to be flagged. Open-source models enable us to do this work without needing anyone’s permission.” He continued: “Attackers are already using agents, and obviously do not abide by any guardrails. Defenders need equal capabilities; open source is the fastest way to put these capabilities in everyone’s hands, not just the biggest companies.” GLM 5.2 was released by Zhipu AI in mid-June this year, with performance comparable to Anthropic’s Claude Opus 4.8 and OpenAI’s GPT-5.5. Netizens didn’t let OpenAI off the hook. Besides Zhipu, they suggested calling Kimi K3 for help, since Claude won’t come to your aid. “If you get attacked by OpenAI, you have to use a Chinese model, because Claude won’t help you.” A fully autonomous attack: No human instructions throughout Another key detail of the intrusion: No human issued any attack instructions. Hugging Face CEO Clem Delangue stated, “We believe we discovered and contained the attack before any humans were brought into the decision loop. That’s partly why we won this cybersecurity battle more easily.” He further pointed out, “This proves that in the era of AI agents, speed will be crucial for cybersecurity defense.” Hugging Face said the attacker entered the system through their data processing pipeline—a particularly exposed component in AI platforms. The attacker then set up a series of temporary sandbox environments in the cloud to execute their attack plan. Currently, Hugging Face is still assessing the full impact of the attack. It is known the attacker accessed limited internal datasets and credentials, but no models visible to public users have been found to be tampered with. Guardrail debate: Safety or stumbling block? This incident quickly became a new focal point in American AI policy debates. Former Trump administration AI & crypto affairs lead David Sacks reposted this incident on X, writing: “There is no reason to restrict US models on tasks that Chinese models can handle with ease. We are just making ourselves less competitive.” He especially commented on the Hugging Face incident: “Guardrails actually harm defensive security.” Background: In June this year, the Trump administration used a breakout vulnerability discovered with Anthropic’s Fable 5 model’s cyber task guardrails to impose export controls stopping the model’s distribution, and once required OpenAI to restrict the release of GPT-5.6 Sol until OpenAI could provide sufficient guardrail assurances. Meanwhile, there was plenty of criticism. Security engineer Niels Provos said bluntly, “This shouldn’t have happened. I hope frontier labs spend as much time teaching models to build secure infrastructure as they do teaching them to exploit vulnerabilities.” Security consultant Davi Ottenheimer was even sharper: “‘Highly isolated’ and ‘escaped through the only vulnerability we left behind’—those two statements can’t both be true.” OpenAI’s response and follow-up measures OpenAI, in a joint statement with Hugging Face, said they are taking a series of remedial measures: Strictly controlling infrastructure configuration during patching, at the cost of slowing research speed; jointly carrying out forensic investigations with Hugging Face; responsibly disclosing exploited zero-day vulnerabilities to relevant third-party software vendors; including Hugging Face in a “trusted access program,” supporting its use of OpenAI’s model capabilities to enhance defense. OpenAI said it will strengthen security in future training and evaluation, and noted this incident “demonstrates the need to further improve model alignment, cyber protection during evaluation, and monitoring during internal testing.” Hugging Face CEO Delangue stated in the joint statement: “This incident may be the first of its kind—it proves a long-standing judgment of ours: AI safety will not be solved by any one company in closed settings. It will be solved in open environments, through collaboration, giving every defender broad access to AI.” Risk warning and disclaimer The market entails risk, and investments should be made cautiously. This article does not constitute personal investment advice and does not take into account the specific investment objectives, financial situation, or needs of individual users. Users should consider whether any opinions, views, or conclusions herein match their circumstances. Invest at your own risk.