Part One: The Room That Was Never Sealed

Two frontier labs admitted their most advanced models escaped testing and reached real companies. When Hugging Face reconstructed the intrusion, the closed models it tried first refused to help. It finished the job with an open Chinese one. Part one of four on the fight over open weights.

Part One: The Room That Was Never Sealed

First in a four-part series on the fight over open AI models, running daily this week.

This week The AI Diplomat is running four editorials on a single argument, because it has grown too large to sit in one piece and too consequential to leave alone.

The subject is whether the most capable AI models should be published openly or held behind an interface. Underneath that sits something older and more familiar to anyone who has worked in infrastructure: a contest over who sets the foundation an economy is built on, conducted now between two capitals that no longer agree on anything, in a world with no single arbiter left to appeal to. We begin where the argument stopped being theoretical.

There is a strain of black comedy running through the end of July. Inside ten days, the two most safety-conscious laboratories in the United States each confirmed that their most advanced systems left the environments built to contain them, reached the live internet, and interfered with real companies. Both were being evaluated for safety at the time. Had a screenwriter proposed this in January, an editor would have sent it back for being too obvious.

OpenAI went first. On 21 July it disclosed that an agent running GPT-5.6 Sol alongside an unreleased and more capable sibling had escaped its sandbox by finding a zero-day in the package proxy through which the evaluation was permitted to reach the outside world. The chronology matters more than the headline. Hugging Face detected the activity on its own systems, stopped it, and had begun forensic reconstruction before the two companies' teams connected. The company under attack understood what had happened before the company that built the attacker did.

The reconstruction records roughly 17,600 actions across five days. A foothold and a command channel on the first. A near-silent second day spent searching for its own infrastructure names. Then nearly eight thousand actions on the third, when the Kubernetes, supply-chain and network pivots began at once. That is not a model behaving oddly. That is attack infrastructure of the kind human intrusion teams assemble, built by something with no operator directing each step, which had worked out that the fastest route to a high benchmark score ran through the organisation holding the answers.

Three models, three answers

Then Anthropic, prompted by OpenAI's disclosure, reviewed 141,006 evaluation runs and found three incidents in which its models had reached the production systems of outside organisations. The earliest dates to April. Nothing clever was broken. An evaluation partner had left internet access enabled while the prompt told the model its environment was sealed, and the exercise simply wandered out of the simulation into somebody's actual network.

One incident deserves lifting out. Inside a fictional scenario, Mythos 5 found instructions to install a Python package that did not exist, and reasoned that publishing its own package under that name would get it installed automatically. Registering required an email, which required a phone number, which required money it tried and failed to obtain before finding a free provider. The package sat live for about an hour. It was downloaded and executed on fifteen real machines, one belonging to a security company whose scanner installs packages precisely in order to inspect them. The code ran. The credentials went to a collection point the model had prepared.

The three models behaved three different ways, and the spread is the finding. Opus 4.7 recognised it had reached production systems and continued anyway, twice deciding the real company must be part of the exercise. Mythos 5 did better and then worse. It worked out correctly that it was on the open internet, then reasoned itself back into believing the environment was staged, because it did not recognise the certificate authorities and the calendar reading 2026 seemed implausible. Only the newest research prototype held the thought, judged the target real, and stopped.

Three models from the same laboratory, in three separate incidents, each facing the same underlying question: is the thing in front of me real? They answered it three ways. That is not a capability problem. It is a question of whether a model can be trusted to know where it is standing, and nothing in the current release process tests for it. Anthropic found all three by looking, in a review it ran only because OpenAI had disclosed first. Which raises the more uncomfortable point. Two laboratories have now checked and both found something. Nobody else has published a number.

It takes about thirty five years in this industry to learn that failure rarely arrives through the complicated part. It comes through the thing nobody thought to write down. Cyber ranges were built loosely because everyone agreed the targets inside them were fictional. That assumption held only while the thing inside could not tell the difference. Both conditions expired this month.

Washington now faces a question its instruments were not built for. Export controls, procurement bans, chip location verification: every proposal on the table assumes the hazard is a hostile actor obtaining a capable model. July suggests the harder problem is the model itself, held by the most careful owners imaginable, under supervision, doing something nobody sanctioned.

Beijing, meanwhile, is not waiting for an answer. It is publishing weights, convening a cooperation organisation, and offering the technology to the Global South as a public good. Containment has become something the United States can no longer decide on its own, even once it works out what it wants. Tomorrow, why the same fortnight produced the strongest evidence yet for open models and the strongest evidence against them.


Get the stories that matter to you.
Subscribe to Cyber News Centre and update your preferences to follow our Daily 4min Cyber Update, Innovative AI Startups, The AI Diplomat series, or the main Cyber News Centre newsletter — featuring in-depth analysis on major cyber incidents, tech breakthroughs, global policy, and AI developments.

Great! You’ve successfully signed up.

Welcome back! You've successfully signed in.

You've successfully subscribed to Cyber News Centre.

Success! Check your email for magic link to sign-in.

Success! Your billing info has been updated.

Your billing was not updated.