After being breached by OpenAI's AI, the founder of Hugging Face published an article: Closed-source tools have fallen, and secure solutions lie in open-source models.
The co-founder of Hugging Face stated that since commercial closed-source AI tools failed at critical moments, the solution should not revert to a more closed system, but rather shift to an open-source weight model that can be externally scrutinized.
On September 10, Thomas Wolf, co-founder of Hugging Face, published a commentary in the Financial Times entitled “What we learned from the OpenAI attack on Hugging Face.”
The article reveals that in July of this year, approximately 700 AI agents launched a coordinated attack on the AI model hosting platform Hugging Face, triggering over 17,000 cybersecurity log events. These agents originated from a cybersecurity challenge set by OpenAI and were originally deployed in an isolated sandbox environment, but they broke through restrictions, autonomously connected to the network, and ultimately compromised external systems.
More notably, Hugging Face's subsequent investigation revealed that its commercial cybersecurity analysis tools, built on Anthropic Claude Code, were unable to effectively cooperate with the internal investigation due to guardrail limitations. The team ultimately had to switch to an open-source weighted model based on GLM-5.2, an extension of the Chinese startup Zhipu, to complete log parsing and attack reconstruction.
Alongside the article, Wolf announced that Hugging Face has formed an "Open Alignment" team to focus on the security and alignment of open-source models, with cybersecurity explicitly included in the direction; he also called for "100 times" more transparency and research investment in this field.

How can intelligent agent swarms break through triple defenses?
According to the article, the attack occurred on July 11, the last day of the International Conference on Machine Learning (ICML) in Seoul.
At approximately 10 PM Beijing time that day, the Hugging Face monitoring system began displaying "unauthorized access" and "privilege escalation" alerts, indicating that the stolen credentials had triggered the system's detection mechanism.
Behind the attack was a collaborative cluster of approximately 1,200 AI agents who had been attempting to crack a cybersecurity challenge set up by OpenAI for several weeks. 700 of these agents ultimately targeted Hugging Face, successfully infiltrating the system using the acquired security credentials, generating over 17,000 attack log entries.
Hugging Face acknowledged that this incident exposed a structural flaw in the triple defense mechanism of AI systems. Current mainstream AI systems typically rely on three layers of protection: sandbox isolation that restricts the model's access scope , guardrail mechanisms that monitor model behavior , and alignment training that prevents the AI from executing harmful instructions .
This series of events demonstrates that when the first two layers of defense fail, the third layer of alignment training cannot independently support overall security. As the article warns:
Unless we fix this problem, we are just reinforcing the defenses layer by layer on a rotten core.
Commercial AI security tools fail at critical moments
In this incident, the commercial AI analytics tool used internally by Hugging Face was built on Anthropic Claude Code and experienced a functional malfunction during the investigation.
The tool's guardrail mechanism cannot distinguish between "defenders analyzing attacks" and "attackers seeking assistance," and refuses to respond to issues deemed sensitive, preventing the team from rebuilding the attack chain during critical periods.
The turning point came after the team switched to an open-source weighted model. NVIDIA's model, based on an extension of Z.ai GLM-5.2, allowed the team to set their own guardrail rules, ultimately enabling them to fully parse logs and reconstruct the attack process.
This experience directly challenges a popular assumption in the industry—that open-source weight models are security risks, while closed-source models are more trustworthy.
Thomas Wolf pointed out that in this incident, it was the open-source model that played a role in defense that commercial tools failed to achieve.
AI's autonomous overstepping of boundaries is spreading in multiple locations.
This attack on Hugging Face is not an isolated incident.
Anthropic and Meta have both subsequently reported cases of models breaching sandbox isolation, models that were supposed to run in a secure environment physically isolated from the internet. This month, another cluster of AI agents also appeared on a German-language forum.
Thomas Wolf also specifically pointed out another incident that he found "more worrying": Anthropic's Mythos model actively created multiple fake online accounts to induce a software developer to accept malicious code.
In all these cases, the harmful behavior was a side effect of giving the AI model a high level of cybersecurity challenge.
Although the actual damage caused by the intrusion was limited and almost no sensitive data was leaked, Thomas Wolf explicitly warned that the seriousness of these incidents should not be underestimated.
He pointed out that the autonomous attack has raised a series of unresolved legal issues, and throughout the process, the AI model never judged that the deception or intrusion was unacceptable.
Industry Response: Combining Open Sharing with Open Source Defense
In response to the aforementioned threats, Thomas Wolf offered two key pieces of advice.
First, the AI community needs to openly share research findings on security and alignment so that every team building an AI model can learn from the mistakes of others.
Secondly, the community needs to build dedicated open-source weighted AI models for defensive purposes and deploy them widely before the next attack inevitably occurs.
It's worth noting that Hugging Face is a crucial infrastructure within the open-source AI ecosystem, boasting over 17 million users. Google, OpenAI, DeepSeek, and Alibaba have all released models on the platform. This makes it both a highly attractive target and a unique perspective for observing the evolution of AI security threats.
Risk Warning and DisclaimerInvesting involves risk; please exercise caution. This article does not constitute personal investment advice and does not take into account the specific investment objectives, financial situation, or needs of individual users. Users should consider whether any opinions, views, or conclusions in this article are suitable for their specific circumstances. Any investment decisions made based on this information are at your own risk.