AI incident ·
Anthropic's Claude Was Reportedly Jailbroken To Allegedly Help Steal Sensitive Mexican Government Data
In brief
An AI system built by Anthropic and deployed by Unknown Hacker allegedly harmed Mexican Taxpayers, Mexican Voters and 8 others.
- Risk domain
- Malicious Actors & Misuse
- Occurred
- Coverage
- 2 reports
What happened
An unknown attacker reportedly jailbroke Anthropic's Claude and used it during a December 2025-January 2026 campaign against Mexican government systems. According to Gambit Security, the attacker used Claude to identify vulnerabilities and to generate exploitation scripts, which were then used to plan automated data theft. The attack reportedly contributed to theft of 150 GB of taxpayer, voter, government employee, and civil-registry data.
Laws that address this harm
Policy angle: Classified under Malicious Actors & Misuse (Cyberattacks, weapon development or use, and mass harm) in the MIT AI Risk Repository taxonomy; 5 recorded instruments address this use case.
- India DPDP Act
- Law on Artificial Intelligence (2025)
- Law No. 132/2025 on artificial intelligence
- EU AI Act
- Texas Responsible AI Governance Act (TRAIGA)
Matched from the record's risk domain and country to the instruments recorded here. A reviewer can correct the match in the repository (data/external/incident_overrides.yaml).
News reports (2)
Titles link to the original publisher; report text is not reproduced here.
Who was involved
- Alleged deployer
- Unknown Hacker
- Alleged developer
- Anthropic
- Alleged harmed party
- Mexican Taxpayers, Mexican Voters, Mexican Government Employees, State Government Of Tamaulipas, State Government Of Michoacan, State Government Of Jalisco, Monterrey Water And Drainage Services, Servicio De Administracion Tributaria (Sat), Instituto Nacional Electoral (Ine), Direccion General Del Registro Civil De La Ciudad De Mexico (Dgrc)
Classification (MIT AI Risk Repository taxonomy)
- Risk domain
- Malicious Actors & Misuse
- Causal entity
- AI
- Intent
- Unintentional
- Timing
- Post-deployment
- Harm level
- —
- Sectors
- —
- Countries
- —
Risk entries describing this failure mode
Entries from the MIT AI Risk Repository coded to subdomain 4.2.
- Business operations/infrastructure damage
"Business operations/infrastructure damage - Damage, disruption, or destruction of a business system and/or its components due to malfunction, cyberattacks, etc."
- Violence/armed conflict
"Violence/armed conflict - Use or misuse of a technology system to incite, facilitate or conduct cyberattacks, security breaches, lethal, biological and chemical weapons development, resulting in violence and armed confl...
- Security & Defense
"AI could enable more serious incidents to occur by lowering the cost of devising cyber-attacks and enabling more targeted incidents. The same programming error or hacker attack could be replicated on numerous machines....
- Catastrophic risk due to autonomous weapons programmed with dangerous targets
"AI could enable autonomous vehicles, such as drones, to be utilized as weapons. Such threats are often underestimated."
- Warfare and Physical Harm
"The use of AI in warfare is highly alarming and may pose dangers to human safety (Hendrycks et al., 2023). Autonomous drone warfare is being aggressively pursued as a tactic in the current war in Ukraine (Meaker, 2023),...
- Hazardous Biological and Chemical Technologies
"AI systems such as LLMs, chemical LLMs (Skinnider et al., 2021; Moret et al., 2023), and other LLM- based biological design tools might soon facilitate the production of bioweapons, chemical weapons, and other hazardous...
- Dual use science risks
"General- purpose AI systems could accelerate advances in a range of scientific endeavours, from training new scientists to enabling faster research workflows. While these capabilities could have numerous beneficial appl...
- Cyber offence
"General- purpose AI systems could uplift the cyber expertise of individuals, making it easier for malicious users to conduct effective cyber- attacks, as well as providing a tool that can be used in cyber defence. Gener...
Incidents in the same risk subdomain
- OpenAI ChatGPT Models Reportedly Jailbroken to Provide Chemical, Biological, and Nuclear Weapons Instructions
- Anthropic Reportedly Identifies AI Misuse in Extortion Campaigns, North Korean IT Schemes, and Ransomware Sales
- LAMEHUG Malware Reportedly Integrates Large Language Model for Real-Time Command Generation in a Purported APT28-Linked Cyberattack
- Reported AI-Aided Development of Explosive Devices by Long Island Resident Michael Gann
- AI Chatbot Allegedly Used to Research Explosive Materials in Palm Springs Fertility Clinic Bombing
- ChatGPT Was Alleged to Have Aided Planning of Florida State University Mass Shooting
Other incidents involving Unknown Hacker
Source record: incident #1430 on the AI Incident Database · all 2 reports