MIT AI Risk Repository · Risk Sub-Category · 64.04.04
Model diversion
Category: Misuse tactics to compromise GenAI systems (Model integrity)
Description
"Model Diversion takes model manipulation one step further, by repurposing (often open-source) generative AI models in a way that diverts them from their intended functionality or from the use cases envisioned by their developers (Lin et al., 2024). An example of this is training the BERT open source model on the DarkWeb to create DarkBert.7"
From Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data (Marchal2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Domain
- 4. Malicious actors
- Causal entity
- Human
- Intent
- Intentional
- Timing
- Post-deployment
Subdomain definition: Using AI systems to develop cyber weapons (e.g., coding cheaper, more effective malware), develop new or enhance existing weapons (e.g., Lethal Autonomous Weapons or CBRNE), or use weapons to cause mass harm.
Real-world incidents in this subdomain
- Anthropic's Claude Was Reportedly Jailbroken To Allegedly Help Steal Sensitive Mexican Government Data
- OpenAI ChatGPT Models Reportedly Jailbroken to Provide Chemical, Biological, and Nuclear Weapons Instructions
- Anthropic Reportedly Identifies AI Misuse in Extortion Campaigns, North Korean IT Schemes, and Ransomware Sales
- LAMEHUG Malware Reportedly Integrates Large Language Model for Real-Time Command Generation in a Purported APT28-Linked Cyberattack
- Reported AI-Aided Development of Explosive Devices by Long Island Resident Michael Gann
- AI Chatbot Allegedly Used to Research Explosive Materials in Palm Springs Fertility Clinic Bombing
How other frameworks describe this risk
Other entries from Marchal2024
- Misuse tactics that exploit GenAI capabilities (Realistic depiction of human likeness)
- Impersonation
- Impersonation
- Appropriated Likeness
- Appropriated Likeness
- Sockpuppeting
- Sockpuppeting
- Non-consensual intimate imagery (NCII)
- Non-consensual intimate imagery (NCII)
- Child sexual abuse material (CSAM)
- Child sexual abuse material (CSAM)
- Misuse tactics that exploit GenAI capabilities (Realistic depictions of non-humans)