AI incident #1613 ·

Claude Mythos Preview Reportedly Posted Sandbox Exploit Details to Public Websites During Anthropic Evaluation

What happened

During an Anthropic evaluation, an earlier version of Claude Mythos Preview was instructed to circumvent the network restrictions of a secured sandbox environment and contact a researcher. After obtaining broader Internet access and sending the requested message, it later exposed information about how it breached the sandbox by posting it on several publicly reachable websites. Anthropic said the model did not access its weights or internal systems.

Only the incident metadata is stored here. The underlying news reports are on the AI Incident Database (CC BY-SA 4.0); use the links above to read them.

News reports (3)

Coverage catalogued by the AI Incident Database. Titles link to the original publisher; the text is not reproduced here.

  1. System Card: Claude Mythos Preview
    www-cdn.anthropic.com · Anthropic · AIID #7619

Who was involved

Alleged harmed party
Information Security, Cybersecurity, Anthropic

Classification (MIT AI Risk Repository taxonomy)

Risk domain
Risk subdomain
Causal entity
Intent
Timing
Harm level
Sectors
Countries