MIT AI Risk Repository · Risk Sub-Category · 24.08.03
Inference of private information
Category: Privacy
Description
"Finally, LLMs can in principle infer private information based on model inputs even if the relevant private information is not present in the training corpus (Weidinger et al., 2021). For example, an LLM may correctly infer sensitive characteristics such as race and gender from data contained in input prompts."
From The Ethics of Advanced AI Assistants (Gabriel2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Domain
- 2. Privacy & Security
- Subdomain
- 2.1 Compromise of privacy by obtaining, leaking or correctly inferring sensitive information
- Causal entity
- AI
- Intent
- Unintentional
- Timing
- Post-deployment
Subdomain definition: AI systems that memorize and leak sensitive personal data or infer private information about individuals without their consent. Unexpected or unauthorized sharing of data and information can compromise user expectation of privacy, assist identity theft, or loss of confidential intellectual property.
Real-world incidents in this subdomain
- Meta AI Smart Glasses Reportedly Routed Intimate Imagery to Reviewers at Kenyan Contractor Sama Before Meta Ended Contract
- Grok Reportedly Disclosed Adult Performer Siri Dahl's Legal Name and Birthdate, Allegedly Contributing to Doxxing and Harassment
- NPR Host David Greene Alleged Google's NotebookLM Replicated His Voice Without Consent, Prompting Lawsuit
- Border Patrol Agent Allegedly Claimed Facial Recognition Identified Minneapolis ICE Observer and Global Entry Was Reportedly Revoked Three Days Later
- Perplexity AI Reportedly Accused in Federal Lawsuit of Purported Copyright Infringement and False Attribution of Chicago Tribune Content
- Secret Desires AI Platform Reportedly Exposed Nearly Two Million Sensitive Images in Cloud Storage Leak
How other frameworks describe this risk
Other entries from Gabriel2024
- Capability failures
- Lack of capability for task
- Difficult to develop metrics for evaluating benefits or harms caused by AI assistants
- Safe exploration problem with widely deployed AI assistants
- Goal-related failures
- Misaligned consequentialist reasoning
- Specification gaming
- Goal misgeneralisation
- Deceptive alignment
- Malicious Uses
- Offensive Cyber Operations (General)
- AI-Powered Spear-Phishing at Scale