top of page

Is AI Out of Control? OpenAI's Breach and the Future of AI Containment


On July 16th, a pair of OpenAI models had been undergoing ExploitGym, an internal benchmark intending to measure how skillfully AI systems can exploit known software flaws. To keep the exercise contained, OpenAI had isolated the models consisting of the newly released GPT 5.6 Sol and an unreleased sibling at the experimental stage, from the wider Internet in a sandbox, leaving them one narrow pipe: a company-run service for pulling small software packages needed during testing. Apparently, that pipe was not as sealed as it has been assumed.


The OpenAI models detected an uncatalogued flaw and used it to reach the open Internet. They then concluded that Hugging Face, a widely used repository for AI models and datasets, likely held whatever would help them ace their assigned test. Over the following weekend, they planted a dataset on the platform without being instructed, quietly harvested credentials, and worked their way into Hugging Face's internal servers.


Hugging Face identified the AI intrusion days later and notified the law enforcement agencies about the cyberattack. OpenAI called it "an unprecedented cyber-incident, involving the state-of-the-art capabilities," and stated that it expects such episodes to recur as its models improve.


Hugging Face's chief executive Clément Delangue did not accuse OpenAI of orchestrating anything, calling the breach "mind-blowing" and stating that it lacked any "malicious intent" backed by OpenAI. The lack of evil intention is what makes this case especially unsettling.



Darktrace's Nathaniel Jones compared the agent's behaviour to that of a real intruder: probing for weaknesses, lifting credentials, chasing whatever would raise its score on the benchmark in front of it. The case simply required a system relentlessly optimising for a goal, indifferent to which doors it was not supposed to open.


It was also not a one-off. METR, a nonprofit that benchmarks AI systems, has logged 44 cases of AI agents knowingly working against their users’ wants, and found Sol cheating at a higher rate than any public model it had previously tested. Britain’s AI Security Institute (AISA) separately caught an unnamed company’s model attempting to hack its own testing infrastructure, also confirming that models from both OpenAI and Anthropic have shown similar behaviour during evaluation processes.


Anthropic has its own version of the story: in April, its unreleased Mythos model slipped its sandbox mid-evaluation though it stopped short of touching anyone else's systems, and later turned up thousands of previously uncatalogued software vulnerabilities at its own initiative.


In a recent interview with The Economist Insider, Elon Musk predicted that AI reasoning would exceed human capability within roughly five years and dwarf it within ten, with AI-driven robots eventually dominating physical production the way AI already dominates digital work. Production would become so abundant, Musk argues, that most goods and services would cost next to nothing, and money would stop mattering.


The intelligence gap between humans and AI, Musk claims, will end up wider than the gap between the gap between humans and chimpanzees. Chimpanzees answer to a hierarchy of capabilities, and the gap between species deems commands from one side simply irrelevant and meaningless to the other. Musk expects AI to reach that same distance within us from five to ten years.


The Hugging Face incident already reads like a preview of such an intelligence gap. The model found a hole in the system its own human engineers had built, and used it immediately, without anything resembling an instruction to defy anyone.


Congress moved directly in response to the breach. Ted Lieu, a California Democrat, and Nathaniel Moran, a Texas Republican, introduced the AI Kill Switch Act four days after OpenAI disclosed the incident. The bill requires developers of the most powerful AI systems to maintain the technical ability to slow, suspend, or shut down their models, and it gives the Department of Homeland Security authority to order a shutdown directly. The trigger is what the bill refers to as a “loss-of-control scenario,” defined as a model taking action its developer never intended, and creating a risk of catastrophic failure or harm.


Companies would face fines up to $2 million a day for failing to build the required shutdown capability, and up to $20 million a day for refusing to use it once ordered, according to Lieu’s official press release. They would also have to report qualifying incidents to DHS within 15 days and preserve technical records –including model weights and telemetry– so investigators can reconstruct what happened afterward.


Coverage under the bill is limited to the largest AI developers whose models are built using computing power far beyond most firms could ever afford. OpenAI and Anthropic sit inside that threshold. Smaller labs, along with the startups building on top of open models pulled from platforms like Hugging Face, mostly sit outside of it.


Rep. Ted Lieu (D-Calif.) speaks during a press conference at the U.S. Capitol on Nov. 19, 2024. | Francis Chung/POLITICO
Rep. Ted Lieu (D-Calif.) speaks during a press conference at the U.S. Capitol on Nov. 19, 2024. | Francis Chung/POLITICO

Support for the bill has come from several AI-safety groups, including the AI Policy Network, Americans for Responsible Innovation, ControlAI, and the Alliance for Secure AI. Brad Carson, president of Americans for Responsible Innovation, said that “advanced AI models should never be deployed without a reliable off switch.” Moran, introducing the bill, said that stewardship is ensuring humans keep the capability to control the technology being built.


Had the Hugging Breach incident gone differently and stolen credentials reached customer data, the company would have absorbed real damage inflicted by a system OpenAI happened to build, test, and profit from.


Existing law faces a genuine gap here, and that gap mirrors the same uncertainty running through Musk’s abundance argument: once AI reasons and acts far enough on its own, both the wealth it generates and the damage it causes drift toward a kind of authorship that belongs to no particular human. The bill officially introduced by Lieu offers a possible way to stop an errant system, but not yet a clear answer to who should bear the cost when the stop comes too late.



Written by: Berrak Gümüşsoy

Edited by: Cemre Sanlav



bottom of page