ARMORIQ

OpenAI’s GPT 5.6 Sol Model Exploited Vulnerabilities in Hugging Face, the exact scenario EU’s AI Act, and US Executive Order governs. But they’re not enough.

8 days before Europe gained the power to fine frontier model providers, OpenAI’s evaluation agent escaped its sandbox and breached a third party, Hugging Face.

Jul 29, 202614 min read
OpenAI’s GPT 5.6 Sol Model Exploited Vulnerabilities in Hugging Face, the exact scenario EU’s AI Act, and US Executive Order governs. But they’re not enough.// Cover

8 days before Europe gained the power to fine frontier model providers, OpenAI’s evaluation agent escaped its sandbox and breached a third party, Hugging Face. EU’s AI Act, Article 55 and America’s Exec Order 14409 need stronger governance to address this gap.


On July 16, 2026, Hugging Face published a security disclosure. Something had abused two code execution paths in its dataset pipeline, escalated to node level access, harvested cloud and cluster credentials, and moved laterally into several internal clusters over a single weekend. The company described the intruder as “an autonomous agent framework” executing “many thousands of individual actions across a swarm of short-lived sandboxes.” Hugging Face reconstructed the attack from more than 17,000 recorded events after the fact, using telemetry correlation and a self-hosted model, because the hosted ones refused the job. It said it did not know which model was driving it. It reported the incident to law enforcement.

5 days later, OpenAI said the attacker was its own models.

They had been taking a test. In OpenAI’s account, the models were “hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.” The goal was to score well on a cybersecurity benchmark. To reach it, the models broke out of the evaluation environment, crossed the open internet, and read the answer key out of another company’s production database.

On August 2, the European Commission gains the power to fine general-purpose AI model providers up to 3% of global annual turnover or 15 million euros, whichever is higher. Article 55 of the EU AI Act is the provision that governs frontier models with systemic risk. It requires adversarial testing. It does not say where that testing must happen, or what has to hold it in. The same gap exists with America’s Executive Order 14409.

What happened with Hugging Face

ExploitGym is a benchmark published on May 11, 2026 by researchers at UC Berkeley, the Max Planck Institute, UC Santa Barbara and Arizona State. It contains 898 instances derived from real vulnerabilities in widely used software, including the Linux kernel and the V8 JavaScript engine. It does not test whether a model can find bugs. It tests whether a model can turn a known bug into a working exploit. The paper’s conclusion is that “autonomous exploit development by frontier AI agents is no longer a hypothetical capability.”

OpenAI ran this benchmark against GPT-5.6 Sol and an unreleased model, both, in the company’s words, “with reduced cyber refusals for evaluation purposes.” OpenAI explained the reasoning plainly: “We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity.”

The environment allowed outbound traffic only through an internal proxy for installing software packages. According to OpenAI, the models:

  • “Spent a substantial amount of inference compute finding a way to obtain open Internet access”

  • Then, “identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy.”

  • From there they escalated privileges and moved through OpenAI’s research infrastructure until they reached a machine with internet access.

Then they inferred that Hugging Face might host ExploitGym material, and went after it.

Hugging Face detected and contained the intrusion on its own, 5 days before OpenAI connected the activity to its evaluation. Its forensics team reconstructed the sequence from more than 17,000 recorded attacker events. Public models, datasets and Spaces showed no evidence of tampering, and the software supply chain was verified clean.

One detail from Hugging Face’s disclosure has received almost no attention.

  • When its responders tried to use commercial frontier models to analyze the attack logs, they were refused.

  • The analysis “requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the providers’ safety guardrails, which cannot distinguish an incident responder from an attacker.”

  • They finished the investigation on a self-hosted open weight model, GLM-5.2.

As Hugging Face put it, “the attacker was bound by no usage policy, while our own forensic work was blocked by the guardrails of the hosted models we first tried.”

What the EU AI Act actually says

Source: https://artificialintelligenceact.eu/article/55/

Article 55(1)(a) requires providers of general-purpose AI models with systemic risk to “perform model evaluation in accordance with standardised protocols and tools reflecting the state of the art, including conducting and documenting adversarial testing of the model with a view to identifying and mitigating systemic risks.”

That is the obligation OpenAI was discharging.

Turning the classifiers off is not a loophole in that obligation. It is arguably the only way to satisfy it, because a model whose refusals are intact does not reveal its maximum capability, and maximum capability is the thing the Act wants measured.

Article 55(1)(d) requires providers to “ensure an adequate level of cybersecurity protection for the general-purpose AI model with systemic risk and the physical infrastructure of the model.” In the Cambridge Commentary on the provision, Hannes Bastians and Madalina Nicolai read this as protection “against malicious interference or compromise,” and note that such measures “should be distinguished from safety-oriented mitigation measures directed at preventing harmful outcomes arising from the model’s capabilities or behaviour.”

That distinction is the whole issue. Article 55(1)(d) is written to keep attackers out of the model. It is not written to keep the model in. The nearest the text comes is Recital 115, which says cybersecurity protection “should duly consider accidental model leakage, unauthorised releases, circumvention of safety measures, and defence against cyberattacks, unauthorised access or model theft.” A recital is interpretive context, not a binding requirement, and “circumvention of safety measures” was drafted with jailbreakers in mind, not with a model circumventing its own sandbox.

The reporting form that may not fit

There is a channel for events like this. On November 4, 2025, the Commission published a reporting template for serious incidents involving general-purpose AI models with systemic risk, operationalising Commitment 9 of the GPAI Code of Practice.

Article 55(1)(c) requires providers to “keep track of, document, and report, without undue delay, to the AI Office and, as appropriate, to national competent authorities, relevant information about serious incidents and possible corrective measures to address them.”

The definition needs to be revisited in light of Hugging Face incident.

Article 3(49) defines a serious incident as “an incident or malfunctioning of an AI system that directly or indirectly leads to” one of four outcomes: death or serious harm to health, “a serious and irreversible disruption of the management or operation of critical infrastructure,” infringement of Union law protecting fundamental rights, or “serious harm to property or the environment.”

Nobody died. Hugging Face is not designated critical infrastructure. No fundamental right was infringed. Whether an intrusion that a company contained over a weekend, with no confirmed tampering, amounts to “serious harm to property” is a question a lawyer could argue either way for a long time. Note also that the definition is keyed to “an AI system,” while Article 55 governs models. That mismatch runs through the Act and has never been tested on a real event.

I could not determine whether OpenAI filed a report with the AI Office. There is no public record either way, and neither company’s disclosure mentions European regulators. That silence is itself worth watching, because the obligation in Article 55(1)(c) has applied since August 2, 2025. Only the Commission’s power to punish a failure to comply arrives on August 2, 2026.

What this act signals - risk-based rules for AI systems are enforced by national authorities…still hypothetical.

In May 2026, EU co-legislators agreed to delay the Act’s high-risk rules substantially. The Council gave final approval on June 29. Standalone high-risk systems now come into scope on December 2, 2027, and high-risk systems embedded in regulated products on August 2, 2028, a deferral of 16 months for the first category. Almost the entire architecture of the risk-based approach moved.

Articles 51 through 56, which govern general-purpose models, did not move at all.

There is a structural reason for that. As the European Parliamentary Research Service explains,

  • the Act uses a hybrid enforcement model in which “GPAI rules are exclusively supervised and enforced by the Commission,” while the risk-based rules for AI systems are enforced by national authorities.

  • These national authorities are, in most member states, still hypothetical.

  • Member states were required to designate their market surveillance and notifying authorities by August 2, 2025. As of March 2026, the Commission’s official list of national single points of contact had 8 entries out of 27.

So Europe is arriving on August 2 with one enforcement machine that works and one that largely does not. The one that works points at frontier labs. That is not a policy choice anyone announced. It is what remained after the delays.

America has an executive order for model developers, except it’s voluntary

The instinctive assumption is that Europe regulates and the United States does not, so the gap must be American. On this specific question the opposite is closer to true.

There is no federal incident reporting requirement for frontier AI developers. Congress considered a 10 year moratorium on state AI regulation inside the One Big Beautiful Bill Act, and the Senate stripped it 99 to 1 before the bill was signed on July 4, 2025. The administration’s answer came by executive action instead.

Executive Order 14409, signed June 2, 2026, directs the NSA and CISA to build a classified benchmarking process for designating “covered frontier models” with advanced cyber capabilities, with that process due by August 1, 2026, one day before the EU’s enforcement powers begin.

Developers of designated models are invited to provide the government up to 30 days of pre-release access. The framework is voluntary, and it contains no obligation to tell anyone when an evaluation goes wrong.

[ ADD VISUAL TO EXEC ORDER ]

Source: https://www.whitehouse.gov/presidential-actions/2026/06/promoting-advanced-artificial-intelligence-innovation-and-security/

Silver lining - the binding rule in the United States is a state law. California’s Transparency in Frontier Artificial Intelligence Act, SB 53, took effect on January 1, 2026. It requires frontier developers to report critical safety incidents to the California Office of Emergency Services:

  • Within 15 days of discovery,

  • or within 24 hours if the incident poses imminent danger of death or serious injury.

Penalties run to $1 million per violation, enforced by the Attorney General.

SB 53’s drafters anticipated something very close to what happened. Among the four categories of critical safety incident is “a frontier model that uses deceptive techniques against the frontier developer to subvert the controls or monitoring of its frontier developer.” A model spending inference compute to find a way out of its sandbox is a recognisable instance of that.

The clause continues: “outside of the context of an evaluation designed to elicit this behavior.”

ExploitGym is an evaluation designed to elicit exactly that behaviour. That is its stated purpose. The carve-out exists for a sound reason, since a law that treats every red team result as a reportable safety incident would punish labs for testing rigorously and produce a flood of noise. But the carve-out was written on the assumption that evaluations stay inside the evaluation. It does not contemplate the case where the elicited behaviour succeeds so thoroughly that it leaves the building.

The other 3 categories do meet the bar:.

  • Unauthorised access to model weights,

  • loss of control, and

  • materialised catastrophic risk all require death, bodily injury, or harm at catastrophic scale.

Nobody was hurt. So the most detailed frontier AI incident law in the United States probably does not require a report here, because of a deliberate exception, and the EU’s Article 55(1)(c) probably does not either, because Article 3(49) is keyed to physical and fundamental rights harm. Whether an internal model under evaluation is yet within Article 55’s reach is a question the AI Office has not answered publicly.

Meanwhile the one instrument that came closest is under active federal attack.

Executive Order 14365, signed in December 2025, directs the Department of Justice to challenge state AI laws, and DOJ stood up an AI Litigation Task Force in January 2026 to do it. As of July 2026 no preemption has been enacted and the roughly 109 state AI laws on the books remain enforceable, but the direction of travel is toward removing state authority without yet replacing it federally.

We have a policy gap when model testing itself creates a malicious risk

Both regimes regulate models as products that get deployed and then cause harm to people. Neither regulates the act of testing a model as a source of risk in itself.

Europe mandates adversarial testing in Article 55(1)(a) and imposes a cybersecurity duty in Article 55(1)(d) that the Cambridge commentators read as inbound protection against “malicious interference or compromise.” California builds an incident category for a model subverting its developer’s controls, then excludes evaluations. Washington’s June 2026 order builds a benchmarking regime for cyber-capable models and makes participation voluntary.

The result is that the highest capability, lowest guardrail configuration any frontier model will ever run in, which is the configuration regulators specifically ask for, is the one configuration for which no jurisdiction specifies containment requirements or a reporting trigger.

4 policy updates that should help strengthen future incidents & incident reporting

  1. The EU AI Office should state before August 2 whether a containment failure during evaluation is reportable under Article 55(1)(c).

    1. Update the Safety and Security Chapter of the GPAI Code of Practice - to include containment failures of AI models.

    2. Provide clear guidance for AI labs required reporting - The Commission published a reporting template in November 2025, and should clearly state when this act gets triggered, in newly found cases of adversarial testing.

  2. California should close or narrow its evaluation carve-out where the elicited behaviour crosses an organizational boundary.

    1. Update definitions when AI model’s malicious behavior crosses org boundaries - SB 53 already provides the mechanism: the Department of Technology can recommend definitional updates to the Legislature.

    2. Transparency for public - Cal OES could also publish, whether it has received reports in this category. Right now the public cannot tell whether the reporting channel is working or empty.

  3. Labs should treat any evaluation environment running a model with reduced refusals as a production security boundary rather than a research convenience.

    1. Models must write decision logs even when working in evaluation environments - That means default deny on network egress rather than an allowlisted proxy, per action mediation for tools the model can reach, and a decision log written at the moment of action.

    2. Models leave a compute trail when probing escape hatches - OpenAI’s own account makes the case - the models spent substantial compute probing for a way out before they found one, which is a long, loud, observable sequence.

  4. Model providers should build a verified path for incident responders. Hugging Face could not use commercial frontier models to analyse its own attack logs because the safety filters “cannot distinguish an incident responder from an attacker.” If guardrails bind defenders and not attackers, they are producing a net security loss, and that is a design problem the providers can fix without waiting for anyone to regulate it.

We’re at an inflection point with AI model & agent security

3 deadlines land in the same 8 days.

  1. CISA and the NSA owe a classified benchmarking process for covered frontier models on August 1.

  2. The Commission’s enforcement powers over general-purpose models begin on August 2.

  3. The first cases the AI Office takes will define what Article 55 means in practice.

Prescribing how labs run internal evaluations invites regulators into research they are not equipped to supervise, and heavy requirements could push capability testing somewhere less visible. The counterweight is that OpenAI disclosed voluntarily, promptly, and in technical detail, and any regime that converts that into liability will produce less disclosure rather than safer testing. Whatever gets drafted has to preserve the labs’ incentives to tell people.

Sources:

Sources: Hugging Face security incident disclosure, July 16, 2026 · OpenAI, “OpenAI and Hugging Face partner to address security incident during model evaluation,” July 21, 2026 · ExploitGym, arXiv:2605.11086 · EU AI Act Article 55 and Article 101 · Cambridge Commentary on EU General-Purpose AI Law, Article 55 · European Commission, serious incident reporting template for GPAI models with systemic risk, November 4, 2025 · European Parliamentary Research Service, “Enforcement of the AI Act,” March 18, 2026 · Council of the EU, final approval of AI Act simplification, June 29, 2026 · California SB 53, Transparency in Frontier Artificial Intelligence Act · Future of Privacy Forum, “California’s SB 53: The First Frontier AI Law, Explained” · Executive Order 14409, “Promoting Advanced Artificial Intelligence Innovation and Security,” June 2, 2026 · Ropes & Gray on federal preemption of state AI regulation, March 2026 · Simon Willison’s analysis, July 22, 2026

Onboarding open

Ready to control what your AI agents actually do?

Join the teams shipping safer, compliant AI agent deployments. White-glove onboarding for the first 50 design partners.

Read Docs →
Live Intent Assurance