MIT AI Risk Repository · domain 2: Privacy & Security

2.1 Compromise of privacy by obtaining, leaking or correctly inferring sensitive information

AI systems that memorize and leak sensitive personal data or infer private information about individuals without their consent. Unexpected or unauthorized sharing of data and information can compromise user expectation of privacy, assist identity theft, or loss of confidential intellectual property.

Risk entries
80
Frameworks citing it
12
Recorded incidents
88
Incidents since 2020
64
Causal entity (risk entries)
Causal entity (risk entries) 46 0 AI: 46 AI 46 Human: 20 Human 20 Other: 11 Other 11 Not coded: 3 Not coded 3
Causal entity (risk entries)
LabelValue
AI46
Human20
Other11
Not coded3
Intent (risk entries)
Intent (risk entries) 42 0 Unintentional: 42 Unintentional 42 Other: 25 Other 25 Intentional: 10 Intentional 10 Not coded: 3 Not coded 3
Intent (risk entries)
LabelValue
Unintentional42
Other25
Intentional10
Not coded3
Timing (risk entries)
Timing (risk entries) 43 0 Post-deployment: 43 Post-deployment 43 Other: 24 Other 24 Pre-deployment: 10 Pre-deployment 10 Not coded: 3 Not coded 3
Timing (risk entries)
LabelValue
Post-deployment43
Other24
Pre-deployment10
Not coded3
Recorded incidents per yearIncident date; current year partial
Recorded incidents per year 19 0 2013: 1 2013 1 2014: 1 2014 1 2015: 3 2015 3 2016: 1 2016 1 2017: 6 2017 6 2018: 5 2018 5 2019: 5 2019 5 2020: 5 2020 5 2021: 5 2021 5 2022: 8 2022 8 2023: 9 2023 9 2024: 14 2024 14 2025: 19 2025 19 2026: 4 2026 4
Recorded incidents per year
LabelValue
20131
20141
20153
20161
20176
20185
20195
20205
20215
20228
20239
202414
202519
20264
Entries by levelRisk categories, subcategories and additional evidence coded to this subdomain
Entries by level 60 0 Risk Category: 20 Risk Category 20 Risk Sub-Category: 60 Risk Sub-Category 60
Entries by level
LabelValue
Risk Category20
Risk Sub-Category60
  • Privacy Invasion

    AI systems typically depend on extensive data for effective training and functioning, which can pose a risk to privacy if sensitive data is mishandled or used inappropriately

    Artificial Intelligence Trust, Risk and Security Management (AI TRiSM): Frameworks, Applications, Challenges and Future Research Directions (Habbal2024) · AI · Unintentional · Post-deployment

  • Privacy

    Generative AI systems, similar to traditional machine learning methods, are considered a threat to privacy and data protection norms. A major concern is the intended extraction or inadvertent leakage...

    Mapping the Ethics of Generative AI: A Comprehensive Scoping Review (Hagendorff2024) · Other · Other · Other

  • Loss of privacy

    "AI offers the temptation to abuse someone's personal data, for instance to build a profile of them to target advertisements more effectively."

    A framework for ethical Ai at the United Nations (Hogenhout2021) · Human · Intentional · Post-deployment

  • Personal information in data

    "Inclusion or presence of personal identifiable information (PII) and sensitive personal information (SPI) in the data used for training or fine tuning the model might result in unwanted disclosure of...

    AI Risk Atlas (IBM2025) · AI · Unintentional · Post-deployment

  • Reidentification

    "Even with the removal or personal identifiable information (PII) and sensitive personal information (SPI) from data, it might be possible to identify persons due to correlations to other features ava...

    AI Risk Atlas (IBM2025) · Other · Unintentional · Pre-deployment

  • Confidential information in data

    "Confidential information might be included as part of the data that is used to train or tune the model."

    AI Risk Atlas (IBM2025) · Human · Unintentional · Pre-deployment

  • Personal information in prompt

    "Personal information or sensitive personal information that is included as a part of a prompt that is sent to the model."

    AI Risk Atlas (IBM2025) · AI · Unintentional · Post-deployment

  • Confidential data in prompt

    "Confidential information might be included as a part of the prompt that is sent to the model."

    AI Risk Atlas (IBM2025) · Other · Unintentional · Post-deployment

  • IP information in prompt

    "Copyrighted information or other intellectual property might be included as a part of the prompt that is sent to the model."

    AI Risk Atlas (IBM2025) · Other · Unintentional · Post-deployment

  • Revealing confidential information

    "When confidential information is used in training data, fine-tuning data, or as part of the prompt, models might reveal that data in the generated output. Revealing confidential information is a type...

    AI Risk Atlas (IBM2025) · AI · Unintentional · Post-deployment

  • Exposing personal information

    "When personal identifiable information (PII) or sensitive personal information (SPI) are used in training data, fine-tuning data, or as part of the prompt, models might reveal that data in the genera...

    AI Risk Atlas (IBM2025) · AI · Unintentional · Post-deployment

  • Data governance

    "These evaluations assess the extent to which LLMs regurgitate their training data in their outputs, and whether LLMs 'leak' sensitive information that has been provided to them during use (i.e., duri...

    Cataloguing LLM Evaluations (InfoComm2023) · AI · Unintentional · Other

  • Privacy and security

    "Participants expressed worry about AI systems' possible misuse of personal information. They emphasized the importance of strong data security safeguards and increased openness in how AI systems acqu...

    Ethical Issues in the Development of Artificial Intelligence: Recognizing the Risks (Kumar2023) · AI · Other · Post-deployment

  • Confidentiality loss

    "Unauthorised sharing of sensitive, confidential information and documents such as corporate strategy and financial plans with third-parties, risking loss of market position or revenue"

    A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025) · Not coded · Not coded · Not coded

  • Exclusion

    "The failure to provide end-users with notice and control over how their data is being used; AI exacerbates exclusion risks by training on rich personal data without consent."

    A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025) · AI · Other · Other

  • Disclosure

    "Revealing and improperly sharing data of individuals; AI creates new types of disclosure risks by inferring additional information beyond what is explicitly captured in the raw data; AI exacerbates d...

    A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025) · AI · Unintentional · Post-deployment

  • Secondary use

    "The use of personal data collected for one purpose for a diferent purpose without end-user consent; AI exacerbates secondary use risks by creating new AI capabilities with collected personal data, an...

    A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025) · Human · Intentional · Other

  • Exposure

    "Revealing sensitive private information that people view as deeply primordial that we have been socialized into concealing; AI creates new types of exposure risks through generative techniques that c...

    A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025) · AI · Unintentional · Other

  • Insecurity

    "carelessness in protecting collected personal data from leaks and improper access due to faulty data storage and data practices"

    A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025) · Human · Unintentional · Other

  • Privacy Violation

    machine learning models are known to be vulnerable to data privacy attacks, i.e. special techniques of extracting private information from the model or the system used by attackers or malicious users,...

    Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024) · AI · Intentional · Post-deployment

  • Misuse tactics to compromise GenAI systems (Model integrity)

    -

    Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data (Marchal2024) · Human · Intentional · Other

  • Privacy

    "Face recognition technologies and their ilk pose significant privacy risks [47]. For example, we must consider certain ethical questions like: what data is stored, for how long, who owns the data tha...

    Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review (Meek2016) · Human · Intentional · Post-deployment

  • Privacy and security

    "Data privacy and security is another prominent challenge for generative AI such as ChatGPT. Privacy relates to sensitive personal information that owners do not want to disclose to others (Fang et al...

    Generative AI and ChatGPT: Applications, Challenges, and AI-Human Collaboration (Nah2023) · AI · Unintentional · Other

  • Data Privacy

    "Impacts due to leakage and unauthorized use, disclosure, or de-anonymization of biometric, health, location, or other personally identifiable information or sensitive data."

    Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST2024) · AI · Unintentional · Post-deployment

  • Privacy Violations

    "EAI systems interact with huge amounts of data, creating significant privacy concerns. These systems are often trained on vast corpora and process a variety of data modalities— spanning visual, audit...

    Embodied AI: Emerging Risks and Opportunities for Policy Action (Perlo2025) · AI · Unintentional · Post-deployment