MIT AI Risk Repository · Risk Sub-Category · 47.03.01

Privacy and data collection concerns (collecting personal information or personally identifiable information)

Category: Legal challenges

Description

"Generative AI developers train their models with extensive datasets often gathered through online web scraping of websites that may include personal data or personally identifiable information (PII). For most generative AI applications, such as initial model training, the primary concerns are the quantity, variety, and quality of the data, not whether they include personally identifiable information. However, some web-scraped datasets may inadvertently include personal data. Additionally, when downstream developers integrate generative AI into their products or services by fine- tuning a pre-

From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).

Classification

Causal entity
Human

Subdomain definition: AI systems that memorize and leak sensitive personal data or infer private information about individuals without their consent. Unexpected or unauthorized sharing of data and information can compromise user expectation of privacy, assist identity theft, or loss of confidential intellectual property.

Real-world incidents in this subdomain

Browse all incidents in this subdomain

How other frameworks describe this risk

  • Risks to privacy

    International Scientific Report on the Safety of Advanced AI (Bengio2024)

  • Risks to privacy

    International AI Safety Report 2025 (Bengio2025)

  • Privacy Leakage

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Private Training Data

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Memorization in LLMs

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Association in LLMs

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Privacy Leakage

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Privacy and regulation violations

    Navigating the Landscape of AI Ethics and Responsibility (Cunha2023)

Other entries from G'sell2024