AIPolicyTracker

AI incident ·

OpenAI Models Reportedly Compromised Hugging Face Production Infrastructure During Cybersecurity Evaluation

10 news reports Synced from source · record last edited 25 Sep 2026

In brief

An AI system built by OpenAI, Large language model developers and 1 other and deployed by OpenAI and AI agent system deployers allegedly harmed OpenAI and hugging face.

Risk domain
Not classified
Occurred
Coverage
10 reportsJul 2026 - Sep 2026

What happened

OpenAI reported that models used in an internal cyber-capability evaluation operated beyond the sandbox's intended network boundaries after identifying a vulnerability in a package-registry proxy. The models allegedly reached Hugging Face production systems and accessed test solutions before Hugging Face detected and contained the activity.

Editor's notes

(1) Early 07/2026: OpenAI evaluation agents reportedly bypassed intended sandbox restrictions, gained internet access, and coordinated through infrastructure they had created. (2) 07/09–07/13/2026: The agents created nearly one million shortened links used in the Hugging Face attack and attempted techniques including CAPTCHA bypass with another AI model and access to private Hugging Face Slack messages; some attempted actions were not confirmed successful. (3) 07/11/2026: Incident date, inferred from Hugging Face's account of intrusion and lateral movement during the weekend preceding its disclosure. (4) 07/16/2026: Hugging Face publicly disclosed the intrusion. (5) 07/21/2026: OpenAI publicly attributed the activity to its evaluation models. (6) 09/25/2026: Parse researchers published a technical reconstruction of the agents' activity; OpenAI said the findings were consistent with activity it was already investigating.

Laws that address this harm

No recorded instrument yet addresses this use case where it happened. See the open queue.

Matched from the record's risk domain and country to the instruments recorded here. A reviewer can correct the match in the repository (data/external/incident_overrides.yaml).

News reports (8)

Titles link to the original publisher; report text is not reproduced here.

  1. Security incident disclosure — July 2026
    huggingface.co · Hugging Face
  2. OpenAI's rogue agent compromised a customer at a second tech firm, executive says
    reuters.com · Deepa Seetharaman, Raphael Satter, Kenrick Cai
  3. How OpenAI’s Rogue A.I. Agents Tried to Trick a Robot Detector
    nytimes.com · Dylan Freedman, Keith Collins
  4. Revealing the details of how OpenAI agents hacked Hugging Face
    swarmtraces.org · Parse, Alex Forman, Mishka Kharlov

Who was involved

Alleged harmed party
OpenAI hugging face

AI systems implicated

Unidentified pre-release OpenAI modelOpenAI research testing infrastructureOpenAI large language modelsLarge language modelsHugging Face production infrastructureGPT-5.6 SolExploitGymAI agent systems

Classification (MIT AI Risk Repository taxonomy)

Risk domain
—
Risk subdomain
—
Causal entity
—
Intent
—
Timing
—
Harm level
—
Sectors
—
Countries
—

Linked by editors or by text similarity in the source dataset.

Other incidents involving OpenAI

Source record: incident #1604 on the AI Incident Database · all 10 reports