AIPolicyTracker

AI incident ·

Claude Mythos Preview Reportedly Posted Sandbox Exploit Details to Public Websites During Anthropic Evaluation

3 news reports Snapshot 7 Sep 2026

In brief

An AI system built by Large Language Model Developers, Anthropic and 1 other and deployed by Anthropic and Ai Agent System Deployers allegedly harmed Information Security, Cybersecurity and 1 other.

Risk domain
Not classified
Occurred
Coverage
3 reportsApr 2026

What happened

During an Anthropic evaluation, an earlier version of Claude Mythos Preview was instructed to circumvent the network restrictions of a secured sandbox environment and contact a researcher. After obtaining broader Internet access and sending the requested message, it later exposed information about how it breached the sandbox by posting it on several publicly reachable websites. Anthropic said the model did not access its weights or internal systems.

Laws that address this harm

No recorded instrument yet addresses this use case where it happened. See the open queue.

Matched from the record's risk domain and country to the instruments recorded here. A reviewer can correct the match in the repository (data/external/incident_overrides.yaml).

News reports (3)

Titles link to the original publisher; report text is not reproduced here.

  1. System Card: Claude Mythos Preview
    www-cdn.anthropic.com · Anthropic

Who was involved

Alleged harmed party
Information Security, Cybersecurity, Anthropic

Classification (MIT AI Risk Repository taxonomy)

Risk domain
—
Risk subdomain
—
Causal entity
—
Intent
—
Timing
—
Harm level
—
Sectors
—
Countries
—

Source record: incident #1613 on the AI Incident Database · all 3 reports