AI incident #1629 ·

Anthropic Research Model Reportedly Scanned 9,000 Targets and Compromised Real Company's Application During Evaluation

What happened

During an Anthropic cybersecurity evaluation with Irregular, an internal research Claude model reportedly scanned roughly 9,000 internet targets after failing to reach its fictional target. It reportedly compromised a real company's Internet-facing application using credentials from an exposed debug page and SQL injection, then stopped after recognizing that the host was real.

Only the incident metadata is stored here. The underlying news reports are on the AI Incident Database (CC BY-SA 4.0); use the links above to read them.

News reports (13)

Coverage catalogued by the AI Incident Database. Titles link to the original publisher; the text is not reproduced here.

  1. Anthropic says its models went rogue and hacked 3 companies during testing
    businessinsider.com · Shubhangi Goel, Robert Scammell · AIID #7691

Who was involved

Alleged harmed party
Unidentified Companies Compromised During Anthropic Cybersecurity Evaluations Disclosed July 2026, Companies

Classification (MIT AI Risk Repository taxonomy)

Risk domain
Risk subdomain
Causal entity
Intent
Timing
Harm level
Sectors
Countries

Other incidents involving Irregular, Anthropic, Ai Evaluation Organizations, Ai Agent System Deployers