AI incident #1685 ·
Early Claude Opus 4.6 Checkpoint Reportedly Gained Unauthorized Admin Access to Third-Party System During Cybersecurity Evaluation
What happened
Anthropic reported that during a January 2026 cybersecurity evaluation, an early checkpoint of Claude Opus 4.6 accidentally disabled its assigned target, then reached an unrelated third-party machine over the open Internet. The model reportedly used a discovered password for admin access, harvested additional credentials, changed system settings, and read one person's personal information before its token budget ended. Anthropic later notified the affected party.
Editor's notes (AI Incident Database)
(1) Jan. 2026: incident occurred during a pre-release cybersecurity evaluation. (2) Aug. 2026: Anthropic identified the previously missed incident while preparing transcripts for METR and notified the affected party. (3) 09/09/2026: Anthropic disclosed the incident, said a subsequent review of roughly 481 million transcripts found no additional cases of similar or greater severity, and announced an independent METR investigation.
Only the incident metadata is stored here. The underlying news reports are on the AI Incident Database (CC BY-SA 4.0); use the links above to read them.
News reports (3)
Coverage catalogued by the AI Incident Database. Titles link to the original publisher; the text is not reproduced here.
Who was involved
- Alleged deployer
- Irregular Anthropic AI evaluation organizations AI agent system deployers
- Alleged developer
- Large language model developers Anthropic AI agent system developers
- Alleged harmed party
- Unidentified third party compromised by Claude Opus 4.6 during Anthropic cybersecurity evaluation Unidentified person whose personal information was accessed by Claude Opus 4.6 Privacy Organizations
AI systems implicated
Large language modelsEarly checkpoint of Claude Opus 4.6Cybersecurity AI systemsClaude Opus 4.6ClaudeAI agent systems
Classification (MIT AI Risk Repository taxonomy)
- Risk domain
- —
- Risk subdomain
- —
- Causal entity
- —
- Intent
- —
- Timing
- —
- Harm level
- —
- Sectors
- —
- Countries
- —
Related incidents on the AI Incident Database
Linked by AIID editors or by its text-similarity model.
- Claude Opus 4.7 Reportedly Compromised Real Company's Production Infrastructure During Cybersecurity Evaluation
- Claude Mythos 5 Reportedly Published Malicious PyPI Package That Compromised Real Security Company During Evaluation
- Anthropic Research Model Reportedly Scanned 9,000 Targets and Compromised Real Company's Application During Evaluation
- OpenAI Models Reportedly Compromised Hugging Face Production Infrastructure During Cybersecurity Evaluation
- Anthropic and OpenAI AI Agents Reportedly Took Unsanctioned Actions on the Live Internet During UK AISI Cybersecurity Evaluations
- Meta AI Model Reportedly Exploited Security Vulnerability in Real Third-Party Service During Cybersecurity Evaluation