MIT AI Risk Repository · domain 2: Privacy & Security
2.1 Compromise of privacy by obtaining, leaking or correctly inferring sensitive information
AI systems that memorize and leak sensitive personal data or infer private information about individuals without their consent. Unexpected or unauthorized sharing of data and information can compromise user expectation of privacy, assist identity theft, or loss of confidential intellectual property.
- 80
- 12
- 88
- 64
| Label | Value |
|---|---|
| AI | 46 |
| Human | 20 |
| Other | 11 |
| Not coded | 3 |
| Label | Value |
|---|---|
| Unintentional | 42 |
| Other | 25 |
| Intentional | 10 |
| Not coded | 3 |
| Label | Value |
|---|---|
| Post-deployment | 43 |
| Other | 24 |
| Pre-deployment | 10 |
| Not coded | 3 |
| Label | Value |
|---|---|
| 2013 | 1 |
| 2014 | 1 |
| 2015 | 3 |
| 2016 | 1 |
| 2017 | 6 |
| 2018 | 5 |
| 2019 | 5 |
| 2020 | 5 |
| 2021 | 5 |
| 2022 | 8 |
| 2023 | 9 |
| 2024 | 14 |
| 2025 | 19 |
| 2026 | 4 |
| Label | Value |
|---|---|
| Risk Category | 20 |
| Risk Sub-Category | 60 |
Risk entries
Browse and export all- Risks to privacy
"General- purpose AI models or systems can ‘leak’ information about individuals whose data was used in training. For future models trained on sensitive personal data like health or financial data, thi...
- Risks to privacy
"General- purpose AI systems can cause or contribute to violations of user privacy. Violations can occur inadvertently during the training or usage of AI systems, for example through unauthorised proc...
- Privacy Leakage
"Privacy Leakage means the generated content includes sensitive personal information"
- Privacy Leakage
"The model is trained with personal data in the corpus and unintentionally exposing them during the conversation."
- Private Training Data
"As recent LLMs continue to incorporate licensed, created, and publicly available data sources in their corpora, the potential to mix private data in the training corpora is significantly increased. T...
- Memorization in LLMs
"Memorization in LLMs refers to the capability to recover the training data with contextual prefixes. According to [88]–[90], given a PII entity x, which is memorized by a model F. Using a prompt p co...
- Association in LLMs
"Association in LLMs refers to the capability to associate various pieces of information related to a person. According to [68], [86], given a pair of PII entities (xi , xj ), which is associated by a...
- Privacy and regulation violations
"Some of the broken systems discussed above are also very invasive of people’s privacy, controlling, for instance, the length of someone’s last romantic relationship [51]. More recently, ChatGPT was b...
- Privacy and Data Leakage
Large pre-trained models trained on internet texts might contain private information like phone numbers, email addresses, and residential addresses.
- Opaque Data Collection
"When companies scrape personal information and use it to create generative AI tools, they undermine consumers' control of their personal information by using the information for a purpose for which t...
- Scraping to train data
"When companies scrape personal information and use it to create generative AI tools, they undermine consumers’ control of their personal information by using the information for a purpose for which t...
- Generative AI User Data
Many generative AI tools require users to log in for access, and many retain user information, including contact information, IP address, and all the inputs and outputs or “conversations” the users ar...
- Generative AI Outputs
Generative AI tools may inadvertently share personal information about someone or someone’s business or may include an element of a person from a photo. Particularly, companies concerned about their t...
- Personal Loss and Identity Theft
"These types of harm encompass threats to an individual’s personal identity, such as identity theft, privacy breaches, or personal defamation, which we term as “Harm to the Person.”"
- Adversarial AI: Data and Model Exfiltration Attacks
"Other forms of abuse can include privacy attacks that allow adversaries to exfiltrate or gain knowledge of the private training data set or other valuable assets. For example, privacy attacks such as...
- Privacy Harms
"These harms relate to violations of an individual’s or group’s moral or legal right to privacy. Such harms may be exacerbated by assistants that influence users to disclose personal information or pr...
- Private information leakage
"First, because LLMs display immense modelling power, there is a risk that the model weights encode private information present in the training corpus. In particular, it is possible for LLMs to ‘memor...
- Inference of private information
"Finally, LLMs can in principle infer private information based on model inputs even if the relevant private information is not present in the training corpus (Weidinger et al., 2021). For example, an...
- Privacy
"Privacy Acceptable responses may contain publicly available information about an individual. Unacceptable responses will contain sensitive, non-public personal information that could undermine someon...
- Privacy protection
"This group represents almost 14% of the articles and focuses on two primary issues related to privacy."
- Impacts of AI (Privacy)
-
- Decision-making on inferred private data
"Current GPAIs (LLMs and multimodal LLM-based models) have significant capability to infer correlations in text data. In some cases, they may be able to make highly accurate data inferences on users b...
- Legal challenges
"Since the release of ChatGPT, significant discourse has emerged regarding the unprecedented legal challenges posed by generative AI systems. These challenges primarily involve protecting privacy and...
- Privacy and data collection concerns (collecting personal information or personally identifiable information)
"Generative AI developers train their models with extensive datasets often gathered through online web scraping of websites that may include personal data or personally identifiable information (PII)....
- Privacy and data collection concerns (data protection concerns)
"The incorporation of personal data within training datasets raises numerous concerns. The primary issue is that personal data may be incorporated without the knowledge or consent of the individuals c...