MIT AI Risk Repository · Risk Sub-Category · 13.01.04

Privacy and Data Protection

Category: Impacts: The Technical Base System

Description

"Examining the ways in which generative AI systems providers leverage user data is critical to evaluating its impact. Protecting personal information and personal and group privacy depends largely on training data, training methods, and security measures."

From Evaluating the Social Impact of Generative AI Systems in Systems and Society (Solaiman2023), as extracted by the MIT AI Risk Repository (CC BY 4.0).

Classification

Causal entity
Human
Intent
Other
Timing
Other

Subdomain definition: AI systems that memorize and leak sensitive personal data or infer private information about individuals without their consent. Unexpected or unauthorized sharing of data and information can compromise user expectation of privacy, assist identity theft, or loss of confidential intellectual property.

Real-world incidents in this subdomain

Browse all incidents in this subdomain

How other frameworks describe this risk

  • Risks to privacy

    International Scientific Report on the Safety of Advanced AI (Bengio2024)

  • Risks to privacy

    International AI Safety Report 2025 (Bengio2025)

  • Privacy Leakage

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Private Training Data

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Memorization in LLMs

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Association in LLMs

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Privacy Leakage

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Privacy and regulation violations

    Navigating the Landscape of AI Ethics and Responsibility (Cunha2023)

Other entries from Solaiman2023