MIT AI Risk Repository · Risk Sub-Category · 64.04.03
Jailbreaking
Category: Misuse tactics to compromise GenAI systems (Model integrity)
Description
"Jailbreaking aims to bypass or remove restrictions and safety filters placed on a GenAI model completely (Chao et al., 2023; Shen et al., 2023). This gives the actor free rein to generate any output, regardless of its content being harmful, biassed, or offensive. All three of these are tactics that manipulate the model into producing harmful outputs against its design. The difference is that prompt injections and adversarial inputs usually seek to steer the model towards producing harmful or incorrect outputs from one query, whereas jailbreaking seeks to dismantle a model’s safety mechanisms
From Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data (Marchal2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Domain
- 2. Privacy & Security
- Causal entity
- Human
- Intent
- Intentional
- Timing
- Post-deployment
Subdomain definition: Vulnerabilities in AI systems, software development toolchains, and hardware that can be exploited, resulting in unauthorized access, data and privacy breaches, or system manipulation causing unsafe outputs or behavior.
Real-world incidents in this subdomain
- COEMPT Quality Assurance Engineers Allegedly Violated Indian CBSE Student Data Privacy Rights by Processing It with Google Gemini
- Hidden Prompt Injection in Brazilian Labor-Court Petition Reportedly Tried to Manipulate Galileu
- Meta Internal AI Agent Reportedly Gave Advice That Allegedly Exposed Sensitive Data to Unauthorized Employees
- CodeWall's Autonomous Agent Reportedly Obtained Unauthorized Access to McKinsey's Lilli AI Platform Database
- Anthropic Said DeepSeek, Moonshot, and MiniMax Used Fraudulent Accounts and Proxies to Illicitly Distill Claude Capabilities at Scale
- DJI Romo Cloud Authorization Bug Reportedly Exposed Camera, Microphone, and Home-Mapping Data From Nearly 7,000 Robot Vacuums
How other frameworks describe this risk
Other entries from Marchal2024
- Misuse tactics that exploit GenAI capabilities (Realistic depiction of human likeness)
- Impersonation
- Impersonation
- Appropriated Likeness
- Appropriated Likeness
- Sockpuppeting
- Sockpuppeting
- Non-consensual intimate imagery (NCII)
- Non-consensual intimate imagery (NCII)
- Child sexual abuse material (CSAM)
- Child sexual abuse material (CSAM)
- Misuse tactics that exploit GenAI capabilities (Realistic depictions of non-humans)