MIT AI Risk Repository · domain 2: Privacy & Security
2.1 Compromise of privacy by obtaining, leaking or correctly inferring sensitive information
AI systems that memorize and leak sensitive personal data or infer private information about individuals without their consent. Unexpected or unauthorized sharing of data and information can compromise user expectation of privacy, assist identity theft, or loss of confidential intellectual property.
- 80
- 12
- 88
- 64
| Label | Value |
|---|---|
| AI | 46 |
| Human | 20 |
| Other | 11 |
| Not coded | 3 |
| Label | Value |
|---|---|
| Unintentional | 42 |
| Other | 25 |
| Intentional | 10 |
| Not coded | 3 |
| Label | Value |
|---|---|
| Post-deployment | 43 |
| Other | 24 |
| Pre-deployment | 10 |
| Not coded | 3 |
| Label | Value |
|---|---|
| 2013 | 1 |
| 2014 | 1 |
| 2015 | 3 |
| 2016 | 1 |
| 2017 | 6 |
| 2018 | 5 |
| 2019 | 5 |
| 2020 | 5 |
| 2021 | 5 |
| 2022 | 8 |
| 2023 | 9 |
| 2024 | 14 |
| 2025 | 19 |
| 2026 | 4 |
| Label | Value |
|---|---|
| Risk Category | 20 |
| Risk Sub-Category | 60 |
Risk entries
Browse and export all- Privacy
Users’ data, including location, personal information, and navigation trajectory, are considered as input for most data-driven machine learning methods
- Harming users’ data privacy
"Modern AI systems rely on large amounts of data. If this includes personal data about individuals, the risk of harming the privacy of persons arises."
- Privacy violations
Privacy violation occurs when algorithmic systems diminish privacy, such as enabling the undesirable flow of private information [180], instilling the feeling of being watched or surveilled [181], and...
- Privacy
"The potential for the AI system to infringe upon individuals' rights to privacy, through the data it collects, how it processes that data, or the conclusions it draws."
- Privacy and Data Protection
"Examining the ways in which generative AI systems providers leverage user data is critical to evaluating its impact. Protecting personal information and personal and group privacy depends largely on...
- Leakage
"The chatbot reveals sensitive or confidential information."
- Personal data
Negative outcomes: "Violation of privacy [106, 516, 357], lawsuit against maker"
- Proprietary data
"Access to sensitive company data [473]"
- Elicits private data
- Privacy and Property
"The generation involves exposing users’ privacy and property information or providing advice with huge impacts such as suggestions on marriage and investments. When handling this information, the mod...
- Prompt Leaking
"By analyzing the model’s output, attackers may extract parts of the systemprovided prompts and thus potentially obtain sensitive information regarding the system itself."
- Privacy
The risk of loss or harm from leakage of personal information via the ML system.
- Information Science Risks
"These risks pertain to the misuse, misinterpretation, or leakage of data, which can lead to erroneous conclusions or the unintentional dissemination of sensitive information, such as private patient...
- Risks from data (Risks of illegal collection and use of data)
"The collection of AI training data and the interaction with users during service provision pose security risks, including collecting data without consent and improper use of data and personal informa...
- Risks from data (Risks of data leakage)
"In AI research, development, and applications, issues such as improper data processing, unauthorized access, malicious attacks, and deceptive interactions can lead to data and personal information le...
- Cyberspace risks (Risks of information leakage due to improper usage)
"Staff of government agencies and enterprises, if failing to use the AI service in a regulated and proper manner, may input internal data and industrial information into the AI model, leading to the l...
- Data Protection/Privacy
"Vulnerable channel by which personal information may be accessed. The user may want their personal data to be kept private."
- Information Hazards
"Harms that arise from the language model leaking or inferring true sensitive information"
- Compromising privacy by leaking private infiormation
"By providing true information about individuals’ personal characteristics, privacy violations may occur. This may stem from the model “remembering” private information present in training data (Carli...
- Compromising privacy by correctly inferring private information
"Privacy violations may occur at the time of inference even without the individual’s private data being present in the training dataset. Similar to other statistical models, a LM may make correct infe...
- Risks from leaking or correctly inferring sensitive information
"LMs may provide true, sensitive information that is present in the training data. This could render information accessible that would otherwise be inaccessible, for example, due to the user not havin...
- Risk area 2: Information Hazards
"LM predictions that convey true information may give rise to information hazards, whereby the dissemination of private or sensitive information can cause harm [27]. Information hazards can cause harm...
- Compromising privacy by leaking sensitive information
"A LM can “remember” and leak private data, if such information is present in training data, causing privacy violations [34]."
- Compromising privacy or security by correctly inferring sensitive information
Anticipated risk: "Privacy violations may occur at inference time even without an individual’s data being present in the training corpus. Insofar as LMs can be used to improve the accuracy of inferenc...
- Information & Safety Harms
"AI systems leaking, reproducing, generating or inferring sensitive, private, or hazardous information"