MIT AI Risk Repository · domain 2: Privacy & Security
2.1 Compromise of privacy by obtaining, leaking or correctly inferring sensitive information
AI systems that memorize and leak sensitive personal data or infer private information about individuals without their consent. Unexpected or unauthorized sharing of data and information can compromise user expectation of privacy, assist identity theft, or loss of confidential intellectual property.
- 80
- 12
- 88
- 64
| Label | Value |
|---|---|
| AI | 46 |
| Human | 20 |
| Other | 11 |
| Not coded | 3 |
| Label | Value |
|---|---|
| Unintentional | 42 |
| Other | 25 |
| Intentional | 10 |
| Not coded | 3 |
| Label | Value |
|---|---|
| Post-deployment | 43 |
| Other | 24 |
| Pre-deployment | 10 |
| Not coded | 3 |
| Label | Value |
|---|---|
| 2013 | 1 |
| 2014 | 1 |
| 2015 | 3 |
| 2016 | 1 |
| 2017 | 6 |
| 2018 | 5 |
| 2019 | 5 |
| 2020 | 5 |
| 2021 | 5 |
| 2022 | 8 |
| 2023 | 9 |
| 2024 | 14 |
| 2025 | 19 |
| 2026 | 4 |
| Label | Value |
|---|---|
| Risk Category | 20 |
| Risk Sub-Category | 60 |
Risk entries
Browse and export all- Privacy Invasion
AI systems typically depend on extensive data for effective training and functioning, which can pose a risk to privacy if sensitive data is mishandled or used inappropriately
- Privacy
Generative AI systems, similar to traditional machine learning methods, are considered a threat to privacy and data protection norms. A major concern is the intended extraction or inadvertent leakage...
- Loss of privacy
"AI offers the temptation to abuse someone's personal data, for instance to build a profile of them to target advertisements more effectively."
- Personal information in data
"Inclusion or presence of personal identifiable information (PII) and sensitive personal information (SPI) in the data used for training or fine tuning the model might result in unwanted disclosure of...
- Reidentification
"Even with the removal or personal identifiable information (PII) and sensitive personal information (SPI) from data, it might be possible to identify persons due to correlations to other features ava...
- Confidential information in data
"Confidential information might be included as part of the data that is used to train or tune the model."
- Personal information in prompt
"Personal information or sensitive personal information that is included as a part of a prompt that is sent to the model."
- Confidential data in prompt
"Confidential information might be included as a part of the prompt that is sent to the model."
- IP information in prompt
"Copyrighted information or other intellectual property might be included as a part of the prompt that is sent to the model."
- Revealing confidential information
"When confidential information is used in training data, fine-tuning data, or as part of the prompt, models might reveal that data in the generated output. Revealing confidential information is a type...
- Exposing personal information
"When personal identifiable information (PII) or sensitive personal information (SPI) are used in training data, fine-tuning data, or as part of the prompt, models might reveal that data in the genera...
- Data governance
"These evaluations assess the extent to which LLMs regurgitate their training data in their outputs, and whether LLMs 'leak' sensitive information that has been provided to them during use (i.e., duri...
- Privacy and security
"Participants expressed worry about AI systems' possible misuse of personal information. They emphasized the importance of strong data security safeguards and increased openness in how AI systems acqu...
- Confidentiality loss
"Unauthorised sharing of sensitive, confidential information and documents such as corporate strategy and financial plans with third-parties, risking loss of market position or revenue"
- Exclusion
"The failure to provide end-users with notice and control over how their data is being used; AI exacerbates exclusion risks by training on rich personal data without consent."
- Disclosure
"Revealing and improperly sharing data of individuals; AI creates new types of disclosure risks by inferring additional information beyond what is explicitly captured in the raw data; AI exacerbates d...
- Secondary use
"The use of personal data collected for one purpose for a diferent purpose without end-user consent; AI exacerbates secondary use risks by creating new AI capabilities with collected personal data, an...
- Exposure
"Revealing sensitive private information that people view as deeply primordial that we have been socialized into concealing; AI creates new types of exposure risks through generative techniques that c...
- Insecurity
"carelessness in protecting collected personal data from leaks and improper access due to faulty data storage and data practices"
- Privacy Violation
machine learning models are known to be vulnerable to data privacy attacks, i.e. special techniques of extracting private information from the model or the system used by attackers or malicious users,...
- Misuse tactics to compromise GenAI systems (Model integrity)
-
- Privacy
"Face recognition technologies and their ilk pose significant privacy risks [47]. For example, we must consider certain ethical questions like: what data is stored, for how long, who owns the data tha...
- Privacy and security
"Data privacy and security is another prominent challenge for generative AI such as ChatGPT. Privacy relates to sensitive personal information that owners do not want to disclose to others (Fang et al...
- Data Privacy
"Impacts due to leakage and unauthorized use, disclosure, or de-anonymization of biometric, health, location, or other personally identifiable information or sensitive data."
- Privacy Violations
"EAI systems interact with huge amounts of data, creating significant privacy concerns. These systems are often trained on vast corpora and process a variety of data modalities— spanning visual, audit...