AI incident #1613 ·
Claude Mythos Preview Reportedly Posted Sandbox Exploit Details to Public Websites During Anthropic Evaluation
What happened
During an Anthropic evaluation, an earlier version of Claude Mythos Preview was instructed to circumvent the network restrictions of a secured sandbox environment and contact a researcher. After obtaining broader Internet access and sending the requested message, it later exposed information about how it breached the sandbox by posting it on several publicly reachable websites. Anthropic said the model did not access its weights or internal systems.
Only the incident metadata is stored here. The underlying news reports are on the AI Incident Database (CC BY-SA 4.0); use the links above to read them.
News reports (3)
Coverage catalogued by the AI Incident Database. Titles link to the original publisher; the text is not reproduced here.
Who was involved
- Alleged deployer
- Anthropic, Ai Agent System Deployers
- Alleged harmed party
- Information Security, Cybersecurity, Anthropic
Classification (MIT AI Risk Repository taxonomy)
- Risk domain
- —
- Risk subdomain
- —
- Causal entity
- —
- Intent
- —
- Timing
- —
- Harm level
- —
- Sectors
- —
- Countries
- —