MIT AI Risk Repository · Risk Sub-Category · 71.01.05

Information Science Risks

Category: Scientific Domain of Agents

Description

"These risks pertain to the misuse, misinterpretation, or leakage of data, which can lead to erroneous conclusions or the unintentional dissemination of sensitive information, such as private patient data or proprietary research. Recent research has demonstrated how LLMs can be exploited to generate malicious medical literature that poisons knowledge graphs, potentially manipulating downstream biomedical applications and compromising the integrity of medical knowledge discovery [28]. Such risks are pervasive across all scientific domains."

From Risks of AI Scientists: Prioritizing Safeguarding Over Autonomy (Tang2025), as extracted by the MIT AI Risk Repository (CC BY 4.0).

Classification

Causal entity
Other
Intent
Other

Subdomain definition: AI systems that memorize and leak sensitive personal data or infer private information about individuals without their consent. Unexpected or unauthorized sharing of data and information can compromise user expectation of privacy, assist identity theft, or loss of confidential intellectual property.

Real-world incidents in this subdomain

Browse all incidents in this subdomain

How other frameworks describe this risk

  • Risks to privacy

    International Scientific Report on the Safety of Advanced AI (Bengio2024)

  • Risks to privacy

    International AI Safety Report 2025 (Bengio2025)

  • Privacy Leakage

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Private Training Data

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Memorization in LLMs

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Association in LLMs

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Privacy Leakage

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Privacy and regulation violations

    Navigating the Landscape of AI Ethics and Responsibility (Cunha2023)

Other entries from Tang2025