MIT AI Risk Repository · Risk Sub-Category · 64.04.05
Model extraction
Category: Misuse tactics to compromise GenAI systems (Model integrity)
Description
"Data Exfiltration goes beyond revealing private information, and involves illicitly obtaining the training data used to build a model that may be sensitive or proprietary. Model Extraction is the same attack, only directed at the model instead of the training data — it involves obtaining the architecture, parameters, or hyper-parameters of a proprietary model (Carlini et al., 2024)."
From Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data (Marchal2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Domain
- 2. Privacy & Security
- Causal entity
- Human
- Intent
- Intentional
- Timing
- Post-deployment
Subdomain definition: Vulnerabilities in AI systems, software development toolchains, and hardware that can be exploited, resulting in unauthorized access, data and privacy breaches, or system manipulation causing unsafe outputs or behavior.
Real-world incidents in this subdomain
- COEMPT Quality Assurance Engineers Allegedly Violated Indian CBSE Student Data Privacy Rights by Processing It with Google Gemini
- Hidden Prompt Injection in Brazilian Labor-Court Petition Reportedly Tried to Manipulate Galileu
- Meta Internal AI Agent Reportedly Gave Advice That Allegedly Exposed Sensitive Data to Unauthorized Employees
- CodeWall's Autonomous Agent Reportedly Obtained Unauthorized Access to McKinsey's Lilli AI Platform Database
- Anthropic Said DeepSeek, Moonshot, and MiniMax Used Fraudulent Accounts and Proxies to Illicitly Distill Claude Capabilities at Scale
- DJI Romo Cloud Authorization Bug Reportedly Exposed Camera, Microphone, and Home-Mapping Data From Nearly 7,000 Robot Vacuums
How other frameworks describe this risk
Other entries from Marchal2024
- Misuse tactics that exploit GenAI capabilities (Realistic depiction of human likeness)
- Impersonation
- Impersonation
- Appropriated Likeness
- Appropriated Likeness
- Sockpuppeting
- Sockpuppeting
- Non-consensual intimate imagery (NCII)
- Non-consensual intimate imagery (NCII)
- Child sexual abuse material (CSAM)
- Child sexual abuse material (CSAM)
- Misuse tactics that exploit GenAI capabilities (Realistic depictions of non-humans)