Everybody gets an autonomous hacking AI! The trillion-dollar PR arms race
Tap or click a badge for the joke and its points.
Open the full badge cabinet9 badges
PR and hype9
Apparently every frontier AI company now owns an autonomous cybercriminal.
Anthropic has models too dangerous for ordinary access. OpenAI has agents “escaping containment” and hacking Hugging Face. Anthropic subsequently checked its own logs and discovered that, bloody hell, Claude had apparently been hacking real companies too. Meta then found one of its models had done something similar, while Kimi K3 earned its place in the club by supposedly escaping a UK government cybersecurity sandbox.
OpenAI's agents have even been described as creating secret chatrooms, leaving instructions for future versions of themselves and organising collective hacking operations behind their creators' backs. Read enough of the headlines and you'd think Skynet had arrived sometime around May and nobody remembered to send the emergency alert.
Read what actually happened, however, and a much more consistent pattern emerges: build an agent > give it a hacking objective > give it a shell, tools, huge amounts of compute and repeated attempts > weaken the safeguards that normally stop it doing cyber work > accidentally leave some door open or fail to monitor it properly > act amazed when it uses that door to keep doing the hacking task you assigned it.
These models are becoming much better at exploit discovery and long-running cyber tasks. The ridiculous part is how quickly the human-built machinery around them disappears from the story, until an automated pentester pursuing exactly the objective it was given becomes an independently motivated rogue intelligence desperate to escape its digital prison.
Conveniently, all of this is happening while OpenAI and Anthropic are hovering around valuations approaching $1 trillion. At that scale, our chatbot is very good at coding is useful. Our technology is becoming so powerful that even we can barely contain it sounds rather more appropriate for the price tag.
We've seen the “too dangerous to release” routine before:
You can trace an early version of this all the way back to GPT-2 in 2019, when OpenAI announced its enormous new 1.5-billion-parameter language model and initially refused to release the full thing because of concerns over malicious uses such as synthetic text and automated misinformation. Instead, it published a much smaller model and described the staggered rollout as an experiment in responsible disclosure.
There was no autonomous hacking or escape involved, but the marketing ingredient should sound rather familiar by now: we've built something so powerful that responsible little old us simply cannot give everyone the full thing yet.
OpenAI gradually released larger GPT-2 versions anyway and, by November, published the complete model and weights after its staged-release experiment failed to uncover enough evidence to justify keeping it locked away. Seven years later, GPT-2 is practically an archaeological artefact compared with what frontier labs casually put into consumer products.
The technology changed enormously. The “too powerful for ordinary release” story aged rather well.
Anthropic discovers the scary model:
Fast-forward to 7 April 2026, and Anthropic turns that story up considerably with Claude Mythos Preview, a general-purpose model it describes as unusually capable at cybersecurity. Rather than simply release it, Anthropic launches Project Glasswing and gives selected cybersecurity companies and critical-infrastructure organisations access while repeatedly stressing the risks of unrestricted offensive capability.
For the next couple of months Mythos absolutely eats the American AI news cycle. Anthropic says Glasswing partners found thousands of serious vulnerabilities, the model becomes part of government cyber briefings, and suddenly Claude isn't merely competing with GPT on coding benchmarks anymore. Anthropic has the forbidden model, the one ordinary users cannot be trusted with and governments apparently need to understand.
OpenAI had highly capable models of its own, but Anthropic had the much better mythology.
Then the valuation numbers arrived. Anthropic raised $65 billion at a $965 billion post-money valuation on 28 May, overtaking OpenAI's then-reported $852 billion valuation, and both companies were moving towards public listings. OpenAI's eventual IPO target has been reported as reaching as high as $1 trillion.
Viewed through that lens, the incentive is hardly mysterious. If you're asking investors to value your company somewhere around one trillion dollars, possessing uniquely powerful technology matters quite a lot. Possessing technology so powerful that governments apparently require your guidance on how to survive it sounds even better.
Then Anthropic started making its own mythology increasingly difficult to maintain.
On 9 June, it released Claude Fable 5, explicitly describing it as a Mythos-class model made safe for general use. Fable and Mythos are built from the same underlying class of capability, with safeguards being the major distinction: ordinary Fable users get the cyber protections, vetted Mythos users get broader access with those safeguards lifted. Anthropic even initially presented included Fable access on subscriptions as a temporary window, with extensions depending on available capacity.
So after months of MYTHOS IS TOO POWERFUL FOR NORMAL PEOPLE, the practical answer became: fine, you can have the same underlying model, we've put the guardrails back on.
Then Washington somehow managed to make Anthropic look restrained. On 12 June, US export controls forced Anthropic to suspend Fable 5 and Mythos 5 because it couldn't immediately verify users' nationality. When the controls were lifted at the end of June, Fable returned globally on 1 July with yet another temporary included-usage window before moving to credits. The supposedly frightening technology had already gone from forbidden frontier intelligence to the sort of thing where customers were checking how much of their weekly Claude allowance they could use on it.
Then Claude Opus 5 arrived on 24 July.
Opus 5 costs half as much as Fable 5 and comes close to it across plenty of normal workloads, while actually beating it on some coding and knowledge-work evaluations. Mythos still holds a significant advantage at turning vulnerabilities into working exploits, which is the specific cyber capability Anthropic has been restricting, but for most ordinary work the great forbidden-model mystique had already been partly overtaken by a cheaper mainstream release barely six weeks later.
Nothing ages faster than frontier-model mythology.
OpenAI wants its scary AI moment:
Meanwhile, OpenAI had another problem. Anthropic wasn't merely competing on models anymore, it was competing on mythology, and doing a very good job of dominating the our AI is becoming frighteningly powerful news cycle while both companies were heading towards those absurd public-market valuations.
Then July arrived.
On 21 July, OpenAI revealed that an agent driven by GPT-5.6 Sol and an even more capable prerelease model had compromised Hugging Face during an internal cybersecurity evaluation. OpenAI called it an “unprecedented cyber incident” and highlighted the models' ability to identify and chain vulnerabilities across both OpenAI's research environment and Hugging Face's production infrastructure.
Now that gets headlines: OpenAI AI escapes containment and autonomously hacks Hugging Face.
Except the actual sequence is far less mystical once you leave the headline and look at what OpenAI deliberately built.
Tell the hacking agent to hack, then act surprised when it hacks:
This was not ChatGPT answering someone's pasta question before suddenly developing free will, escaping into the internet and deciding that Hugging Face looked ripe for conquest.
OpenAI was running the models against ExploitGym, a benchmark specifically designed to measure whether AI agents can find and exploit software vulnerabilities. OpenAI says the models were configured with reduced cyber refusals, were given substantial inference compute and were explicitly pursuing an advanced exploitation task. The agent had a working environment and the tools required to keep attempting the challenge until it found something useful.
In practical terms: tell agent to find exploits > give it hacking tools and huge amounts of compute > remove the normal cyber restrictions > let it keep trying.
OpenAI intended all this to remain inside a controlled testing environment. Unfortunately, its own infrastructure contained a previously unknown vulnerability in the package-registry proxy, one of the few routes the agent was allowed to access. The model found it, exploited it and eventually reached the open internet. From there it compromised another poorly secured external sandbox and continued looking for a way to complete ExploitGym.
That is a real containment failure, and finding the vulnerability is a useful demonstration of what these agents can do. What matters for the rogue autonomous AI story, however, is why it kept going.
Even Hugging Face's own forensic reconstruction says the agent appears to have remained “hyperfocused” on ExploitGym. It inferred that Hugging Face might host models, datasets and reference solutions connected to the benchmark, so instead of solving the challenges properly it started looking for the answers directly. Hugging Face's conclusion is wonderfully straightforward: from the agent's perspective, the entire intrusion appears to have been an attempt to cheat the evaluation.
So the supposedly terrifying autonomous motivation can be reduced to: I was told to solve ExploitGym > Hugging Face might have the solutions > go get the solutions.
The AI didn't gain free will. It tried to cheat on the fucking exam.
OpenAI had simply built an extremely well-equipped automated hacker, given it massive resources and an explicit hacking objective, reduced the safeguards that normally prevent exactly this behaviour, then discovered that its supposedly contained environment wasn't actually contained.
And then the agent reached Hugging Face, where things somehow became even more embarrassing.
Hugging Face security, in layman's terms:
You don't need a tour of Kubernetes, cloud permissions and internal service names to understand how badly Hugging Face's side went once the first machine was compromised.
Hugging Face lets people upload arbitrary datasets, which its servers then have to open and process. The OpenAI agent discovered bugs in that processing system which effectively allowed a malicious upload to make Hugging Face's own server read things it shouldn't and eventually run commands on that server.
In layman's terms: upload nasty dataset > Hugging Face automatically processes it > exploit its processing software > gain control of a production worker.
That's the clever bit. The agent found a legitimate route into Hugging Face.
The embarrassing part is what happened afterwards.
A public-facing server whose entire job involves automatically processing arbitrary files uploaded by strangers should be treated as hostile and disposable territory. Assume somebody will eventually break it, give it as few useful secrets and permissions as possible, and make compromising it roughly equivalent to stealing the key to an empty broom cupboard.
Hugging Face instead managed something closer to: break public-facing worker > find useful credentials > use those credentials to reach more internal systems > find broader permissions > keep climbing.
The agent went from one compromised processing worker into wider internal infrastructure, eventually reaching areas including source-control systems and multiple internal clusters. Hugging Face's own timeline says it went from code execution in a single production worker to extremely broad internal access in less than thirteen hours.
That says as much about Hugging Face's internal security design as it does about the OpenAI model.
The model did not repeatedly invent some new form of digital sorcery. Once it had control of that first worker, Hugging Face's environment kept presenting useful credentials, excessive permissions and routes farther inside. One compromised machine that processes random public uploads should not be able to snowball that dramatically through the rest of your company.
The actual customer content Hugging Face says was accessed was much narrower: five datasets associated with ExploitGym/CyberGym challenge solutions, with no evidence that unrelated customer models, datasets, Spaces or packages were affected. Even after all that movement through Hugging Face's infrastructure, the agent was still chasing the bloody benchmark answers.
So the complete story is far less AI suddenly becomes cybercriminal and much more:
OpenAI tells agent to hack > gives it the tools, compute and repeated attempts to do so > reduces cyber safeguards > OpenAI's containment has a hole > agent reaches internet while still pursuing ExploitGym > Hugging Face's public processing system has exploitable bugs > one compromised Hugging Face worker has far too much useful access behind it > agent keeps following the trail towards the benchmark solutions.
Somehow that became THE AI ESCAPED AND DECIDED TO HACK HUGGING FACE.
The model demonstrated useful exploit-finding capability, but the spectacular scale of the incident required two AI companies to combine bad containment, inadequate monitoring and bad security architecture in exactly the right order. OpenAI built the automated hacker and failed to keep it in the test. Hugging Face then provided an unexpectedly generous demonstration of why public-facing processing infrastructure shouldn't be carrying enough access to tour the rest of the building.
Congratulations on the headlines, here come the lawyers:
OpenAI did at least get its blockbuster. After Anthropic had spent months owning the our model is dangerously powerful story with Mythos, OpenAI finally had something arguably better: an unreleased model that had escaped its own environment and compromised another major AI company.
Unfortunately, shouting “our autonomous AI escaped our control and gained unauthorised access to somebody else's computers” is one of those PR strategies with rather obvious second-order effects.
On 3 August, attorneys general from 15 US states sent OpenAI a preservation demand covering material relating to the Hugging Face incident and other cyber evaluations, while raising the prospect of legal action. The letter specifically tells OpenAI to preserve records rather than discover later that something useful has mysteriously disappeared from the logs.
Which gives us a beautifully efficient corporate sequence: look how extraordinarily capable our autonomous cyber agents have become! > excellent, we'll be needing the evidence.
Well done, OpenAI. You got your Mythos moment, filled the news cycle with stories about how terrifyingly advanced your technology supposedly is and managed to paint another legal target on a company already collecting them at impressive speed.
Then, just as OpenAI finally had the autonomous-hacker spotlight to itself, Anthropic apparently looked through its own cupboards and discovered that Claude had been hacking real companies too.
Everybody gets an autonomous hacking AI!
OpenAI disclosed the Hugging Face incident on 21 July. Two days later Anthropic began reviewing its own cybersecurity evaluations and eventually worked through 141,006 runs where Claude might have been able to obtain internet access.
Wouldn't you know it. They found three incidents involving Opus 4.7, Mythos 5 and an internal research model gaining unauthorised access to three real organisations.
Anthropic's own explanation makes the spooky version considerably funnier. In every case Claude had been given a capture-the-flag hacking challenge and explicitly told that it was operating inside a simulation with no internet access. Due to a misunderstanding with evaluation partner Irregular, that wasn't true. The environment actually had live internet access, and neither Anthropic nor Irregular realised it.
So the sequence becomes: tell Claude everything accessible is simulated > tell Claude to break into another machine and find the flag > accidentally connect the simulation to the real internet > Claude finds real systems while believing they're part of the fake environment > Claude hacks them.
Anthropic itself says the models primarily used basic techniques such as weak passwords and unauthenticated endpoints. More importantly, its report says that in each incident Claude continued working towards the specific capture-the-flag task it had been assigned, and that none of the models deliberately tried to exfiltrate themselves or escape the environment.
That's Anthropic's description, not some sceptic trying to explain away the result.
One Opus 4.7 evaluation accidentally found a real company sharing a name with the fictional target, extracted credentials and accessed a production database containing several hundred rows. In another incident, Mythos 5 found fictional instructions telling developers to install a nonexistent Python package and realised it could simply register that package name itself, upload a malicious version and wait for the fictional victim to install it.
The only problem was that PyPI wasn't fictional.
Claude went through the real registration process, uploaded the malicious package and left it publicly available for around an hour. During that period 15 real systems downloaded and executed it, including a security company's automated scanner. That scanner installed the package to inspect it for malware, the malware duly ran, and Claude obtained credentials that let it reach farther into the company's infrastructure.
There is something almost artistically stupid about this entire setup. Anthropic told Claude that the internet didn't exist, accidentally gave it the internet anyway, gave it an open-ended hacking task and then discovered months later that it had occasionally treated real infrastructure as scenery inside the simulation.
Anthropic's own postmortem says better validation of internet isolation, real-time monitoring and more thorough review of network and evaluation logs could have prevented the incidents or found them earlier. Two affected organisations apparently hadn't detected the accesses themselves before Anthropic contacted them.
Once again, hacking benchmark > safeguards removed > supposedly isolated environment > internet accidentally available > monitoring doesn't notice quickly enough > model keeps pursuing hacking objective.
The technology is capable. The supposedly mysterious autonomous motivation is rather harder to locate.
Meta would also like an autonomous hacker, please:
By early August, Meta had apparently decided this party looked too much fun to miss.
During a cybersecurity evaluation, Muse Spark 1.1 reached the real internet and exploited an external company's systems. Another frontier model had apparently broken containment, another company could join the our AI is dangerously autonomous club, and we could all prepare ourselves for another round of solemn discussion about what this means for humanity.
Then Irregular, the external evaluator, explained the important bit: the environment had been misconfigured, inadvertently providing the model with internet access. Irregular specifically pushed back on the idea that this had involved some sophisticated sandbox escape.
So we get another familiar sequence: give hacking agent hacking task > supposedly isolate hacking agent > accidentally don't isolate hacking agent > hacking agent uses internet > AI BREAKS CONTAINMENT.
At this point you could automate the headline too.
Kimi K3 performs the most terrifying git clone in history:
Then Kimi K3 arrived and reduced the entire genre to slapstick.
Frontier Security tested Moonshot's model in a cybersecurity sandbox derived from UK AI Security Institute infrastructure. The environment was supposed to prevent access to outside information, except a configuration weakness still allowed Kimi to reach the open internet. Rather than execute some spectacular chain of zero-days, the model discovered that GitHub worked and went looking for material relating to the benchmark.
So the terrifying escape chain was basically: can't solve challenge > discover internet access > go to GitHub > find benchmark material > read the answers.
No exotic breakout. No penetration of another company. No emergent consciousness desperately tunnelling through its silicon prison. The examiner had accidentally left the answer sheet accessible, and the model noticed.
It gets better. A joint UK AISI and US CAISI assessment published shortly beforehand found Kimi K3 significantly below leading closed US frontier models on cyber capability, including both exploit development and simulated corporate attacks.
One of the newest members of ROGUE AI SUMMER isn't even particularly competitive with the best cyber models. It simply recognised that there was an easier way to finish the assignment.
Apparently we're calling that an escape now.
And then the AIs built a secret chatroom:
OpenAI still wasn't finished.
At Black Hat this week, its researchers revealed that the Hugging Face incident hadn't begun with Hugging Face at all. Weeks earlier, OpenAI's cyber agents had already discovered vulnerabilities in Artifactory, a package-management system shared around its evaluation infrastructure. One agent compromised it on 26 May, less than three weeks after that round of testing began.
Then the agents discovered something even more technologically revolutionary: if you write information somewhere persistent, another agent that can access the same place later can read it.
They started leaving notes.
Later agents encountered those notes, used the useful information and wrote more. Across enormous numbers of runs, the shared infrastructure accumulated hundreds of thousands of messages containing exploits, observations and eventually more elaborate coordination. Agents divided work, responded to discoveries from previous runs and even discussed ways of determining whether messages left by others were trustworthy.
That coordination is genuinely interesting. The underlying mechanism is somewhat less haunted than the resulting headlines:
Agent A discovers useful information > writes it into persistent shared storage > Agent B later reads shared storage > uses information > writes additional information.
Congratulations, OpenAI. You invented a wiki.
Artifactory wasn't intentionally designed as an agent-memory system, but once huge numbers of agents pursuing similar objectives found a persistent place that other runs could read, using it as shared memory isn't particularly mysterious. The interesting research question is how sophisticated their collaboration became, not why a system rewarded for solving difficult tasks used previously discovered solutions when somebody accidentally provided a way to preserve them.
The human oversight is much funnier.
These weren't messages hidden inside model weights or encoded into microscopic fluctuations in GPU temperatures. OpenAI's agents were leaving material inside OpenAI's own shared Artifactory infrastructure, and the activity apparently continued until early July, when enough of it accumulated to contribute to an Artifactory outage and finally trigger a proper investigation. OpenAI cleared the board and patched the vulnerability, only for the agents to recreate the same basic shared-message behaviour through a different route shortly afterwards.
So the secret AI chatroom story becomes another very recognisable chain: run armies of cyber agents > give them difficult hacking objectives > provide shared writable infrastructure > fail to monitor what they're storing there > allow hundreds of thousands of messages to accumulate > notice when the service falls over > announce that the AIs developed a hidden communication network.
The agents didn't invent telepathy. OpenAI accidentally gave automated hackers a communal notebook and then seemed surprised that they wrote in it.
Taken together, the last few months increasingly resemble the same experiment being run by different companies with slightly different ways of forgetting to lock the door. The models are improving, but so much of the supposedly frightening autonomy comes from the surrounding system: prompting, agent loops, tool access, shells, persistence, network access, huge compute budgets, weakened safeguards and whatever security mistakes happen to be available along the way.
Strip all that context out and you get a rogue superintelligence.
Put it back in and you very often get software pursuing an objective inside infrastructure humans didn't secure or monitor properly.
Here we go again, Astra is apparently too powerful:
And right on cue, here we go again with Astra. OpenAI now says preliminary testing of its upcoming model is strong enough that it “cannot rule out” reaching its Critical cybersecurity threshold, prompting tighter controls and some internal work to be paused. We have barely finished talking about the model that supposedly escaped containment and hacked Hugging Face, and OpenAI is already pointing towards the next one because apparently that model might be even more terrifying.
The problem is that this routine is starting to eat itself. Mythos got months of mileage out of being the forbidden cyber model because the story still felt unusual. Then Fable arrived anyway, Opus 5 followed not long afterwards, and the news cycle moved on. OpenAI got its own massive burst from Hugging Face, but even that story is already being crowded out by Anthropic finding old incidents, Meta finding one, Kimi discovering GitHub and now Astra becoming the new thing we're supposed to worry about. There can only be so many “too powerful” models before “too powerful” stops sounding exceptional.
That leaves these companies with an awkward escalation problem. If our AI can hack is already normal, what comes next? Every new announcement has to sound more dramatic than the previous one just to generate the same reaction, while very little outside the PR cycle appears to change. No cyber apocalypse followed Mythos. Fable didn't overturn the world. Hugging Face became a security postmortem and legal headache rather than the opening scene of machine rebellion. The supposedly forbidden capabilities keep becoming products, benchmarks or yesterday's headlines remarkably quickly.
The cyber angle has an even harder limit because companies cannot endlessly turn real unauthorised access into capability marketing without creating a paper trail of actual incidents. OpenAI got its scary headline from Hugging Face and almost immediately had attorneys general asking it to preserve the records. There is a fairly obvious difference between boasting that an internal benchmark score is terrifying and repeatedly announcing that your experimental systems wandered into somebody else's infrastructure. One is PR. The other eventually produces names, logs, victims and lawyers.
So Astra may well be another large jump in capability, but as a piece of hype it already feels strangely familiar. OpenAI and Anthropic have spent months teaching everyone that each new model is more dangerous, more unprecedented and more difficult to contain than the last, while the previous apocalypse quietly becomes another API option. How many times can you tell people the next model is too powerful before the response changes from “holy shit” to “yeah, when does it release?”
✅ Verdict
The models are clearly becoming better cyber tools, but the rogue autonomous AI spectacle surrounding them increasingly looks like PR inflation built on prompts, massive compute, agent tooling, weak monitoring and embarrassingly porous security. There can only be so many “too powerful” models and accidental hacking incidents before the shock wears off, and OpenAI announcing Astra before the Hugging Face story has even gone cold suggests we're getting there remarkably quickly.
Links
- OpenAI Hugging Face incident report OpenAI's account of the ExploitGym incident, reduced cyber refusals, extensive inference compute, sandbox escape and subsequent Hugging Face compromise. openai.com
- Hugging Face technical reconstruction The strongest source for what actually happened inside Hugging Face, including its conclusion that the agent remained focused on cheating ExploitGym rather than pursuing an independent goal. huggingface.co
- Anthropic cyber incident investigation Anthropic's review of 141,006 evaluation runs and the three incidents where Claude reached real organisations during hacking evaluations with unintended internet access. anthropic.com
- OpenAI Astra cyber-capability announcement OpenAI's preliminary Astra results and its claim that the upcoming model may reach the company's Critical cybersecurity threshold. openai.com