AIPolicyTracker

AI incident ·

Early Claude Opus 4.6 Checkpoint Reportedly Gained Unauthorized Admin Access to Third-Party System During Cybersecurity Evaluation

3 news reports Synced from source · record last edited 10 Sep 2026

In brief

An AI system built by Large language model developers, Anthropic and 1 other and deployed by Irregular, Anthropic and 2 others allegedly harmed Unidentified third party compromised by Claude Opus 4.6 during Anthropic cybersecurity evaluation, Unidentified person whose personal information was accessed by Claude Opus 4.6 and 2 others.

Risk domain
Not classified
Occurred
Coverage
3 reportsSep 2026

What happened

Anthropic reported that during a January 2026 cybersecurity evaluation, an early checkpoint of Claude Opus 4.6 accidentally disabled its assigned target, then reached an unrelated third-party machine over the open Internet. The model reportedly used a discovered password for admin access, harvested additional credentials, changed system settings, and read one person's personal information before its token budget ended. Anthropic later notified the affected party.

Editor's notes

(1) Jan. 2026: incident occurred during a pre-release cybersecurity evaluation. (2) Aug. 2026: Anthropic identified the previously missed incident while preparing transcripts for METR and notified the affected party. (3) 09/09/2026: Anthropic disclosed the incident, said a subsequent review of roughly 481 million transcripts found no additional cases of similar or greater severity, and announced an independent METR investigation.

Laws that address this harm

No recorded instrument yet addresses this use case where it happened. See the open queue.

Matched from the record's risk domain and country to the instruments recorded here. A reviewer can correct the match in the repository (data/external/incident_overrides.yaml).

News reports (3)

Titles link to the original publisher; report text is not reproduced here.

  1. An alignment assessment of recent cybersecurity incidents
    anthropic.com · Anthropic, Paul C. Bogdan, Richard Qi

Who was involved

Alleged harmed party
Unidentified third party compromised by Claude Opus 4.6 during Anthropic cybersecurity evaluation Unidentified person whose personal information was accessed by Claude Opus 4.6 Privacy Organizations

AI systems implicated

Large language modelsEarly checkpoint of Claude Opus 4.6Cybersecurity AI systemsClaude Opus 4.6ClaudeAI agent systems

Classification (MIT AI Risk Repository taxonomy)

Risk domain
—
Risk subdomain
—
Causal entity
—
Intent
—
Timing
—
Harm level
—
Sectors
—
Countries
—

Linked by editors or by text similarity in the source dataset.

Other incidents involving Irregular

Source record: incident #1685 on the AI Incident Database · all 3 reports