AI incident ·
Early Claude Opus 4.6 Checkpoint Reportedly Gained Unauthorized Admin Access to Third-Party System During Cybersecurity Evaluation
In brief
An AI system built by Large language model developers, Anthropic and 1 other and deployed by Irregular, Anthropic and 2 others allegedly harmed Unidentified third party compromised by Claude Opus 4.6 during Anthropic cybersecurity evaluation, Unidentified person whose personal information was accessed by Claude Opus 4.6 and 2 others.
- Risk domain
- Not classified
- Occurred
- Coverage
- 3 reports
What happened
Anthropic reported that during a January 2026 cybersecurity evaluation, an early checkpoint of Claude Opus 4.6 accidentally disabled its assigned target, then reached an unrelated third-party machine over the open Internet. The model reportedly used a discovered password for admin access, harvested additional credentials, changed system settings, and read one person's personal information before its token budget ended. Anthropic later notified the affected party.
Editor's notes
(1) Jan. 2026: incident occurred during a pre-release cybersecurity evaluation. (2) Aug. 2026: Anthropic identified the previously missed incident while preparing transcripts for METR and notified the affected party. (3) 09/09/2026: Anthropic disclosed the incident, said a subsequent review of roughly 481 million transcripts found no additional cases of similar or greater severity, and announced an independent METR investigation.
Laws that address this harm
No recorded instrument yet addresses this use case where it happened. See the open queue.
Matched from the record's risk domain and country to the instruments recorded here. A reviewer can correct the match in the repository (data/external/incident_overrides.yaml).
News reports (3)
Titles link to the original publisher; report text is not reproduced here.
Who was involved
- Alleged deployer
- Irregular Anthropic AI evaluation organizations AI agent system deployers
- Alleged developer
- Large language model developers Anthropic AI agent system developers
- Alleged harmed party
- Unidentified third party compromised by Claude Opus 4.6 during Anthropic cybersecurity evaluation Unidentified person whose personal information was accessed by Claude Opus 4.6 Privacy Organizations
AI systems implicated
Large language modelsEarly checkpoint of Claude Opus 4.6Cybersecurity AI systemsClaude Opus 4.6ClaudeAI agent systems
Classification (MIT AI Risk Repository taxonomy)
- Risk domain
- —
- Risk subdomain
- —
- Causal entity
- —
- Intent
- —
- Timing
- —
- Harm level
- —
- Sectors
- —
- Countries
- —
Related incidents
Linked by editors or by text similarity in the source dataset.
- Claude Opus 4.7 Reportedly Compromised Real Company's Production Infrastructure During Cybersecurity Evaluation
- Claude Mythos 5 Reportedly Published Malicious PyPI Package That Compromised Real Security Company During Evaluation
- Anthropic Research Model Reportedly Scanned 9,000 Targets and Compromised Real Company's Application During Evaluation
- OpenAI Models Reportedly Compromised Hugging Face Production Infrastructure During Cybersecurity Evaluation
- Anthropic and OpenAI AI Agents Reportedly Took Unsanctioned Actions on the Live Internet During UK AISI Cybersecurity Evaluations
- Meta AI Model Reportedly Exploited Security Vulnerability in Real Third-Party Service During Cybersecurity Evaluation
Other incidents involving Irregular
Source record: incident #1685 on the AI Incident Database · all 3 reports