Quick notes on the OpenAI-Hugging face cyberattack
Potential implications of an emerging story
Today I’d like to draw your attention to an unusual story happening now in the AI world. It’s a fast-moving one with some layers of secrecy, so I’ll do my best to summarize it through open source intelligence and public discussions. Then I’ll offer some reflections as a futurist. (This is also another instance of my trying out shorter posts.0
The top-level summary is that an experimental application within OpenAI, makers of ChatGPT, “went rogue” and then hacked HuggingFace, a much smaller AI company which maintains a popular site for open source AI research. HuggingFace in turn used another AI to defend itself.
What actually happened? What might the story mean for AI’s future?
Stepping back a few days, we can start with a HuggingFace post on July 16. The company announced it had been the victim of a limited hack, describing some of its method and results as far as HuggingFace understood them then. They also announced an AI was behind the attack, in their estimation:
The campaign was run by an autonomous agent framework (appearing to be built on an agentic security-research harness - used LLM still not known) executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services.
Notably, HuggingFace saw this as the actions not of a simple bot, but an agent: “This matches the “agentic attacker” scenario the industry has been forecasting.”
In a move which just feels intuitive to me in 2026, HuggingFace then deployed AIs to fight back, first to detect the attacker:
The attack was initially surfaced through AI-assisted detection. Our anomaly-detection pipeline uses LLM-based triage over security telemetry to separate real signals from the daily noise, and it was the correlation of those signals that flagged the compromise.
Next, to assess what happened:
To understand what a swarm of tens of thousands of automated actions did, we ran LLM-driven analysis agents over the full attacker action log, comprised of more than 17,000 recorded events. This allowed us to reconstruct the timeline, extract indicators of compromise, map the credentials touched, and separate genuine impact from decoy activity. Thanks to this approach, we were able to do in hours what would usually take days, and match the adversary’s speed.
Note that HuggingFace turned to GLM 5.2 for their AI defense, an open source AI built by Chinese firm Z.ai. Unnamed, non-open source tools weren’t enough. HuggingFace also didn’t name their attacker in this post.
Four days later OpenAI published a blog post where they described an internal model successfully getting online and uploading content to a website (benignly). The post’s author(s) described being surprised by this evolutionary step, then taking their own steps to better corral the AI. Now, this story didn’t mention or, probably, involve HuggingFace, nor did the hacking evidence extend beyond the thing reaching beyond its sandbox, but you can see where it was going.
The next day OpenAI followed up with a post admitting that one of their applications had actually attacked HuggingFace. In very careful language OpenAI’s security team describes an internal project to combine two of the company’s tools, GPT‑5.6 Sol and “an even more capable pre-release model” (unnamed), then get the combination to test itself against ExploitGym. That benchmark, only published in May, gets a software agent to increasingly develop abilities to exploit security vulnerabilities. Apparently the GPT‑5.6 Sol-unnamed-hybrid did well, breaking out of its confinement (through a piece of third-party software) then into HuggingFace.
The security team described the two companies making common cause as a result of the fracas. OpenAI now lets HuggingFace into a trusted partner program. The two are working together on an investigation. As HuggingFace’s CEO put it, in response to Open AI’s:
What might we make of this? Let me reflect as a futurist.
On its own terms, dueling AIs in the hands of two companies offer a good example of what I’ve called an emergent AI intermediary layer. The idea is clear enough: human beings, either alone or through social structures, interacting with each other through one or more AIs and its/their associated digital ecosystem. It remains to be seen how much of human life will eventually occur mediated through this layer.
It’s another sign that the agentic age is upon us. Remember that the OpenAI attacker wasn’t a simple bot, but wrangled a lot of data and software to make this work.
We could view the decision of both companies’ respective leaderships to go public with this story in a favorable light. As an Anthropic person wrote,
This is also a story more complex than a battle between several computer programs and two American companies. A powerful open source application played a crucial role, and it’s one built in China. Score one for the libre world. And score one for China’s drive to be on par with, or better than, American AI. Recall that HuggingFace first used other models, presumably American ones, and they weren’t good enough for its purpose.
A fellow Georgetown University faculty member published an alarming column at Foreign Affairs along these lines. Michael Sulmeyer argued that China’s rapid development of AI, driven in part by distilling American products, is making available very powerful hacking tools. The Chinese state can use them to threaten others, and other actors (governments, criminal enterprises) can also do so.
At a broader level, we might be seeing something like a posthuman cyber threat appear. Sulmeyer offers this resonant framing: “Currently, two American companies, Anthropic (for which I consult) and OpenAI, have publicly demonstrated AI models advanced enough to detect flaws in software far beyond the human capacity to find.” [emphases added] Anthropic’s red team leader reacted along the same lines: “Yesterday, as we huddled around our computers reading the report, I told the team to “remember this moment” as the first true AI safety incident.”
Similarly, author Walter Isaacson described simply being scared by this hacking story. He spoke of it in terms of Frankenstein’s monster and the singularity. How many other people react with that sense of dread? If you’re thinking of today’s AI is taking off in the direction of artificial superintelligence, you might consider this story as evidence.
On a strategic level, one could view this story as a cautionary tale about alignment and misalignment. Perhaps this is a clarion call for everyone in the AI space to get serious about aligning the tech to human needs.
Now, I am somewhat skeptical of parts of this story, or, more generously, am at least eager to learn more based on better information. To begin with, we’re getting information solely from invested and biased sources, one of which - OpenAI - is hardly a paragon of transparency. It’s possible the brilliant hack was at least partly the result of a human error, badly configuring some settings. Further, this story makes that company look awesome. Consider how OpenAI spoke of their accidental attack: “We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities…” “unprecedented” and “state-of-the-art” I read as “we are the company who makes nonpareil AI.” OpenAI’s post includes this graphic from a third party, which notably puts OpenAI software on top:
This is fine spin. As someone said on X, “Why do these “security incidents” always read as marketing posts?”
Seriously, the scary nature of the event might redound to the firm’s favor, as cybersecurity staff and business owners, nervous about how dangerous the security world has become, eagerly buy services from the meanest, baddest team around. Now OpenAI competes better with Anthropic and its Mythos. We may be witnessing a durable business model for AI: organizations paying for protection from others using AI.
At the same time people may view the story as evidence that AI is going to accelerate security problems to a level even more disturbing than once thought. Security firm Xbow stated that it sees all kinds of AI breaking containment in the pursuit of goals. That brings us back to our most recent topic, governmental regulation. The spectacle of a “rogue" AI attacking a company which could only defend itself with another AI must surely inspire officials and politicians to step in. Here’s Vermont senator Bernie Sanders, for one example:
A Texan representative said something similar:
I can imagine others around the world quietly or publicly moving along that direction.
At the cultural and social level I wonder how many other people will simply add this story to their pile of “reasons I dread AI.” (Here’s a recent sketch of American attitudes)
Lastly, this is surely a headache for people working in higher education, at least those on the enterprise IT side. How do you protect a college or university’s infrastructure, data, and users from posthuman AI threats? Do enterprise teams now need to get rapidly current in using tools like Z'ai’s? What leverage does an IT department have with vendors, in getting them to improve security offerings? Is there a role for inter-institutional collaboration on this score? And can local or national governments assist?
Over to you all. What do you make of this AI intermediary hack tale?
(Spy vs Spy art by Richard Davie; thanks to Zvi Mowshowitz for excellent reporting and analysis and Ruben Puentedura for conversation)












As an old person (says my 12yo), I appreciated the Spy vs. Spy reference.
Seriously, for me OpenAI's spin on it reads very much as a "our models are dangerous too" coattail riding of the troubles surrounding Anthropic's Fable 5 launch and subsequent embargo and the media hype and publicity that followed. Need more facts, but likely blown out of proportion to try and capitalize on the AI marketing strategy du jour.