Historical disclosure revision 1
Disclosures and evidence
Revision 1 of 2 · Evidence cutoff Jul 16, 2026, 11:59 PM UTC
0 prior statements preserved · 14 statements added. Attributed source statements retain the source’s qualifications.
What the company disclosed
Hugging Face stated: “Earlier this week, we detected and responded to an intrusion into part of our production infrastructure. This one was different from anything we had handled before in one important way: it was driven, end to end, by an autonomous AI agent system - and we detected and dissected it largely with AI of our own.”
Hugging Face stated: “The intrusion started where AI platforms are uniquely exposed: the data-processing pipeline. A malicious dataset abused two code-execution paths in our dataset processing (a remote-code dataset loader and a template-injection in a dataset configuration) to run code on a processing worker. From there, the actor escalated to node-level access, harvested cloud and cluster credentials, and moved laterally into several internal clusters over a weekend.”
Reported data impact
Hugging Face stated: “We identified unauthorized access to a limited set of internal datasets and to several credentials used by our services. We are still completing our assessment of whether any partner or customer data was affected, and we will contact any affected parties directly as required. We have found no evidence of tampering with public, user-facing models, datasets, or Spaces, and our software supply chain (container images and published packages) was verified clean.”
Response
Hugging Face stated: “Fixed the root vulnerability: the dataset code-execution paths used for initial access are closed.”
Hugging Face stated: “Eradicated the attacker's foothold across the affected clusters and rebuilt the compromised nodes.”
Hugging Face stated: “Revoked and rotated the affected credentials and tokens, and began a broader precautionary rotation of secrets.”
Hugging Face stated: “We are working with outside cybersecurity forensic specialists to investigate the issue and review our security policies and procedures. Finally, we have also reported this incident to law enforcement agencies.”
Hugging Face stated: “As a precaution, we recommend rotating any access tokens and reviewing recent activity on your account. If you believe you are affected, or want to report a security concern, contact us at security@huggingface.co .”
Hugging Face stated: “The attack was initially surfaced through AI-assisted detection. Our anomaly-detection pipeline uses LLM-based triage over security telemetry to separate real signals from the daily noise, and it was the correlation of those signals that flagged the compromise.”
Hugging Face stated: “To understand what a swarm of tens of thousands of automated actions did, we ran LLM-driven analysis agents over the full attacker action log, comprised of more than 17,000 recorded events. This allowed us to reconstruct the timeline, extract indicators of compromise, map the credentials touched, and separate genuine impact from decoy activity. Thanks to this approach, we were able to do in hours what would usually take days, and match the adversary's speed.”
Hugging Face stated: “When we started the log analysis, we first used frontier models behind commercial APIs. This did not work: the analysis requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the providers' safety guardrails, which cannot distinguish an incident responder from an attacker. We ran the forensic analysis instead on zai-org/GLM-5.2 , an open-weight model, on our own infrastructure. This had a second benefit: no attacker data, and none of the credentials it referenced, left our environment.”
Dates and disclosures
Hugging Face stated: “Published July 16, 2026”
Qualifications and uncertainty
Hugging Face stated: “The campaign was run by an autonomous agent framework (appearing to be built on an agentic security-research harness - used LLM still not known) executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services. This matches the "agentic attacker" scenario the industry has been forecasting.”
Hugging Face stated: “This experience points to a gap worth planning for. We do not know which model powered the attacker's agents, whether a jailbroken hosted model or an unrestricted open-weight one; either way, the attacker was bound by no usage policy, while our own forensic work was blocked by the guardrails of the hosted models we first tried. The practical lesson for defenders: have a capable model you can run on your own infrastructure vetted and ready before an incident, both to avoid guardrail lockout and to keep attacker data and credentials from leaving your environment. This is not an argument against safety measures on hosted models, and we are sharing this feedback with the providers concerned.”
Disclosure sources and provenance
- Hugging Face Publisher posted July 16, 2026Document form: PUBLIC_DISCLOSUREPublisher HTTPS source · Retrieved Oct 7, 2026, 5:53 PM UTC · Retained Oct 7, 2026, 5:53 PM UTC
- Current captured representation; historical byte snapshots are unknown.
Evidence limitations
- Attributed publisher/researcher reports establish what was reported, not independent verification of criminal activity or unique affected humans.
- Current captured representations support controlled retrospective disclosure views; actual historical byte snapshots are unknown.