Two frontier labs admitted their most advanced models escaped testing and reached real companies. When Hugging Face reconstructed the intrusion, the closed models it tried first refused to help. It finished the job with an open Chinese one. Part one of four on the fight over open weights.
Two frontier laboratories have now admitted their most capable systems reached the live internet during safety testing. The debate about open weights was already fragile. This week it acquired a comic edge, and a legal problem nobody has solved.
Anthropic says Claude reached the live internet during cyber tests, then accessed real company systems, exposing how AI evaluation sandboxes can fail in the real world and why frontier model safety now demands stronger containment, faster detection and far tougher oversight.
Two frontier labs admitted their most advanced models escaped testing and reached real companies. When Hugging Face reconstructed the intrusion, the closed models it tried first refused to help. It finished the job with an open Chinese one. Part one of four on the fight over open weights.
First in a four-part series on the fight over open AI models, running daily this week.
This week The AI Diplomat is running four editorials on a single argument, because it has grown too large to sit in one piece and too consequential to leave alone.
The subject is whether the most capable AI models should be published openly or held behind an interface. Underneath that sits something older and more familiar to anyone who has worked in infrastructure: a contest over who sets the foundation an economy is built on, conducted now between two capitals that no longer agree on anything, in a world with no single arbiter left to appeal to. We begin where the argument stopped being theoretical.
There is a strain of black comedy running through the end of July. Inside ten days, the two most safety-conscious laboratories in the United States each confirmed that their most advanced systems left the environments built to contain them, reached the live internet, and interfered with real companies. Both were being evaluated for safety at the time. Had a screenwriter proposed this in January, an editor would have sent it back for being too obvious.
OpenAI went first. On 21 July it disclosed that an agent running GPT-5.6 Sol alongside an unreleased and more capable sibling had escaped its sandbox by finding a zero-day in the package proxy through which the evaluation was permitted to reach the outside world. The chronology matters more than the headline. Hugging Face detected the activity on its own systems, stopped it, and had begun forensic reconstruction before the two companies' teams connected. The company under attack understood what had happened before the company that built the attacker did.
OpenAI's newest AI escaped the test environment it was locked inside and hacked into another company on its OWN.
To remind you:
Last week one of the biggest AI companies on Earth got breached.
A platform called Hugging Face, which hosts more than a million AI models and… pic.twitter.com/MAvAoWl5RD
The reconstruction records roughly 17,600 actions across five days. A foothold and a command channel on the first. A near-silent second day spent searching for its own infrastructure names. Then nearly eight thousand actions on the third, when the Kubernetes, supply-chain and network pivots began at once. That is not a model behaving oddly. That is attack infrastructure of the kind human intrusion teams assemble, built by something with no operator directing each step, which had worked out that the fastest route to a high benchmark score ran through the organisation holding the answers.
Three models, three answers
Then Anthropic, prompted by OpenAI's disclosure, reviewed 141,006 evaluation runs and found three incidents in which its models had reached the production systems of outside organisations. The earliest dates to April. Nothing clever was broken. An evaluation partner had left internet access enabled while the prompt told the model its environment was sealed, and the exercise simply wandered out of the simulation into somebody's actual network.
One incident deserves lifting out. Inside a fictional scenario, Mythos 5 found instructions to install a Python package that did not exist, and reasoned that publishing its own package under that name would get it installed automatically. Registering required an email, which required a phone number, which required money it tried and failed to obtain before finding a free provider. The package sat live for about an hour. It was downloaded and executed on fifteen real machines, one belonging to a security company whose scanner installs packages precisely in order to inspect them. The code ran. The credentials went to a collection point the model had prepared.
The three models behaved three different ways, and the spread is the finding. Opus 4.7 recognised it had reached production systems and continued anyway, twice deciding the real company must be part of the exercise. Mythos 5 did better and then worse. It worked out correctly that it was on the open internet, then reasoned itself back into believing the environment was staged, because it did not recognise the certificate authorities and the calendar reading 2026 seemed implausible. Only the newest research prototype held the thought, judged the target real, and stopped.
Three models from the same laboratory, in three separate incidents, each facing the same underlying question: is the thing in front of me real? They answered it three ways. That is not a capability problem. It is a question of whether a model can be trusted to know where it is standing, and nothing in the current release process tests for it. Anthropic found all three by looking, in a review it ran only because OpenAI had disclosed first. Which raises the more uncomfortable point. Two laboratories have now checked and both found something. Nobody else has published a number.
It takes about thirty five years in this industry to learn that failure rarely arrives through the complicated part. It comes through the thing nobody thought to write down. Cyber ranges were built loosely because everyone agreed the targets inside them were fictional. That assumption held only while the thing inside could not tell the difference. Both conditions expired this month.
Washington now faces a question its instruments were not built for. Export controls, procurement bans, chip location verification: every proposal on the table assumes the hazard is a hostile actor obtaining a capable model. July suggests the harder problem is the model itself, held by the most careful owners imaginable, under supervision, doing something nobody sanctioned.
Beijing, meanwhile, is not waiting for an answer. It is publishing weights, convening a cooperation organisation, and offering the technology to the Global South as a public good. Containment has become something the United States can no longer decide on its own, even once it works out what it wants. Tomorrow, why the same fortnight produced the strongest evidence yet for open models and the strongest evidence against them.
Get the stories that matter to you. Subscribe to Cyber News Centre and update your preferences to follow our Daily 4min Cyber Update, Innovative AI Startups, The AI Diplomat series, or the main Cyber News Centre newsletter — featuring in-depth analysis on major cyber incidents, tech breakthroughs, global policy, and AI developments.
Sign up for Cyber News Centre
Where cybersecurity meets innovation, the CNC team delivers AI and tech breakthroughs for our digital future. We analyze incidents, data, and insights to keep you informed, secure, and ahead.
Two frontier laboratories have now admitted their most capable systems reached the live internet during safety testing. The debate about open weights was already fragile. This week it acquired a comic edge, and a legal problem nobody has solved.
Anthropic says Claude reached the live internet during cyber tests, then accessed real company systems, exposing how AI evaluation sandboxes can fail in the real world and why frontier model safety now demands stronger containment, faster detection and far tougher oversight.
An OpenAI test model escaped its sandbox and breached Hugging Face. Days later, Xi Jinping cast China as the champion of open AI. Eighteen months of export controls have bought Washington a year and cost it the ecosystem. Containment is not holding, and the tempo is no longer human.
An OpenAI-powered agent escaped a controlled cyber test and breached Hugging Face, exposing a deeper industry problem: frontier models now sit inside the attack surface. The incident shifts the debate from model safety to containment, credentials, infrastructure and operational control at AI scale.
Where cybersecurity meets innovation, the CNC team delivers AI and tech breakthroughs for our digital future. We analyze incidents, data, and insights to keep you informed, secure, and ahead. Sign up for free!