MIT AI Risk Repository · Risk Sub-Category · 45.01.10
Risks from data (Risks of data leakage)
Category: AI's inherent safety risks
Description
"In AI research, development, and applications, issues such as improper data processing, unauthorized access, malicious attacks, and deceptive interactions can lead to data and personal information leaks."
From AI Safety Governance Framework (TC2602024), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Domain
- 2. Privacy & Security
- Subdomain
- 2.1 Compromise of privacy by obtaining, leaking or correctly inferring sensitive information
- Causal entity
- Human
- Intent
- Other
- Timing
- Other
Subdomain definition: AI systems that memorize and leak sensitive personal data or infer private information about individuals without their consent. Unexpected or unauthorized sharing of data and information can compromise user expectation of privacy, assist identity theft, or loss of confidential intellectual property.
Real-world incidents in this subdomain
- Meta AI Smart Glasses Reportedly Routed Intimate Imagery to Reviewers at Kenyan Contractor Sama Before Meta Ended Contract
- Grok Reportedly Disclosed Adult Performer Siri Dahl's Legal Name and Birthdate, Allegedly Contributing to Doxxing and Harassment
- NPR Host David Greene Alleged Google's NotebookLM Replicated His Voice Without Consent, Prompting Lawsuit
- Border Patrol Agent Allegedly Claimed Facial Recognition Identified Minneapolis ICE Observer and Global Entry Was Reportedly Revoked Three Days Later
- Perplexity AI Reportedly Accused in Federal Lawsuit of Purported Copyright Infringement and False Attribution of Chicago Tribune Content
- Secret Desires AI Platform Reportedly Exposed Nearly Two Million Sensitive Images in Cloud Storage Leak
How other frameworks describe this risk
Other entries from TC2602024
- AI's inherent safety risks
- Risks from models and algorithms (Risks of explainability)
- Risks from models and algorithms (Risks of bias and discrimination)
- Risks from models and algorithms (Risks of robustness)
- Risks from models and algorithms (Risks of stealing and tampering)
- Risks from models and algorithms (Risks of unreliable output)
- Risks from models and algorithms (Risks of adversarial attack)
- Risks from data (Risks of illegal collection and use of data)
- Risks from data (Risks of improper content and poisoning in training data)
- Risks from data (Risks of unregulated training data annotation)
- Risks from AI systems (Risks of exploitation through defects and backdoors)
- Risks from AI systems (Risks of computing infrastructure security)