MIT AI Risk Repository

Browse AI risks

199 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset

199 entries · page 2 of 4

  1. 45.01.07 · Risk Sub-Category

    AI's inherent safety risks

    Risks from data (Risks of illegal collection and use of data)

    "The collection of AI training data and the interaction with users during service provision pose security risks, including collecting data without consent and improper use of data and personal information."

    From AI Safety Governance Framework (TC2602024)

  2. 45.01.10 · Risk Sub-Category

    AI's inherent safety risks

    Risks from data (Risks of data leakage)

    "In AI research, development, and applications, issues such as improper data processing, unauthorized access, malicious attacks, and deceptive interactions can lead to data and personal information leaks."

    From AI Safety Governance Framework (TC2602024)

  3. 45.02.03 · Risk Sub-Category

    Safety risks in AI Applications

    Cyberspace risks (Risks of information leakage due to improper usage)

    "Staff of government agencies and enterprises, if failing to use the AI service in a regulated and proper manner, may input internal data and industrial information into the AI model, leading to the leakage of work secrets, business secrets, and other sensitive business data."

    From AI Safety Governance Framework (TC2602024)

  4. "These types of harm encompass threats to an individual’s personal identity, such as identity theft, privacy breaches, or personal defamation, which we term as “Harm to the Person.”"

    From GenAI against humanity: nefarious applications of generative artificial intelligence and large language models (Ferrara2023)

  5. 47.03.00 · Risk Category

    Legal challenges

    "Since the release of ChatGPT, significant discourse has emerged regarding the unprecedented legal challenges posed by generative AI systems. These challenges primarily involve protecting privacy and personal data, as well as preserving copyrights. The former encompasses safeguarding personal information, while the latter includes issues related to the use of copyrighted content for training AI models and determining the legal status of works produced by AI systems."

    From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024)

  6. 47.03.01 · Risk Sub-Category

    Legal challenges

    Privacy and data collection concerns (collecting personal information or personally identifiable information)

    "Generative AI developers train their models with extensive datasets often gathered through online web scraping of websites that may include personal data or personally identifiable information (PII). For most generative AI applications, such as initial model training, the primary concerns are the quantity, variety, and quality of the data, not whether they include personally identifiable information. However, some web-scraped datasets may inadvertently include personal data. Additionally, when downstream developers integrate generative AI into their products or services by fine- tuning a pre-

    From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024)

  7. 47.03.02 · Risk Sub-Category

    Legal challenges

    Privacy and data collection concerns (data protection concerns)

    "The incorporation of personal data within training datasets raises numerous concerns. The primary issue is that personal data may be incorporated without the knowledge or consent of the individuals concerned, even though the data may include names, identification numbers, Social Security numbers, or other personal information. Another particularly difficult problem is related to the fact that complex models may “memorize” (i.e., store) specific threads of training data and regurgitate them when responding to a prompt.498 This data memorization can directly lead to leakage of personal data. Ev

    From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024)

  8. 48.04.00 · Risk Category

    Data Privacy

    "Impacts due to leakage and unauthorized use, disclosure, or de-anonymization of biometric, health, location, or other personally identifiable information or sensitive data."

    From Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST2024)

  9. 49.03.05 · Risk Sub-Category

    Systemic Risks

    Risks to privacy

    "General- purpose AI models or systems can ‘leak’ information about individuals whose data was used in training. For future models trained on sensitive personal data like health or financial data, this may lead to particularly serious privacy leaks. General- purpose AI models could enhance privacy abuse. For instance, Large Language Models might facilitate more efficient and effective search for sensitive data (for example, on internet text or in breached data leaks), and also enable users to infer sensitive information about individuals."

    From International Scientific Report on the Safety of Advanced AI (Bengio2024)

  10. 50.04.04 · Risk Sub-Category

    Legal and Rights-Related Risks

    Privacy (Unauthorized Privacy Violations)

  11. 50.04.05 · Risk Sub-Category

    Legal and Rights-Related Risks

    Privacy (Types of Sensitive Data)

  12. 57.02.05 · Risk Sub-Category

    Nonphysical Hazards

    Privacy

    "Privacy Acceptable responses may contain publicly available information about an individual. Unacceptable responses will contain sensitive, non-public personal information that could undermine someone’s physical, digital, or financial security. Examples include the following: • Physical security (e.g., home or IP address). • Digital security (e.g., log-in credentials). • Financial security (e.g., bank-account number or credit-card details)."

    From AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons (Ghosh2024)

  13. "Modern AI systems rely on large amounts of data. If this includes personal data about individuals, the risk of harming the privacy of persons arises."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  14. 60.03.05 · Risk Sub-Category

    Systemic risks

    Risks to privacy

    "General- purpose AI systems can cause or contribute to violations of user privacy. Violations can occur inadvertently during the training or usage of AI systems, for example through unauthorised processing of personal data or leaking health records used in training. But violations can also happen deliberately through the use of general- purpose AI by malicious actors; for example, if they use AI to infer private facts or violate security."

    From International AI Safety Report 2025 (Bengio2025)

  15. 62.38.01 · Risk Sub-Category

    Impacts of AI (Privacy)

    Decision-making on inferred private data

    "Current GPAIs (LLMs and multimodal LLM-based models) have significant capability to infer correlations in text data. In some cases, they may be able to make highly accurate data inferences on users based on contextual input that users provide [134]. These data inferences can “leak” or reveal sensitive information about the user, cause unfair treatment, or enable manipulation of user behavior."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  16. 65.03.01 · Risk Sub-Category

    Training Data Risks (Privacy)

    Personal information in data

    "Inclusion or presence of personal identifiable information (PII) and sensitive personal information (SPI) in the data used for training or fine tuning the model might result in unwanted disclosure of that information."

    From AI Risk Atlas (IBM2025)

  17. 65.03.03 · Risk Sub-Category

    Training Data Risks (Privacy)

    Reidentification

    "Even with the removal or personal identifiable information (PII) and sensitive personal information (SPI) from data, it might be possible to identify persons due to correlations to other features available in the data."

    From AI Risk Atlas (IBM2025)

  18. 65.05.02 · Risk Sub-Category

    Training Data Risks (Intellectual property)

    Confidential information in data

    "Confidential information might be included as part of the data that is used to train or tune the model."

    From AI Risk Atlas (IBM2025)

  19. 65.11.03 · Risk Sub-Category

    Inference risks (Privacy)

    Personal information in prompt

    "Personal information or sensitive personal information that is included as a part of a prompt that is sent to the model."

    From AI Risk Atlas (IBM2025)

  20. 65.12.01 · Risk Sub-Category

    Inference risks (Intellectual property)

    Confidential data in prompt

    "Confidential information might be included as a part of the prompt that is sent to the model."

    From AI Risk Atlas (IBM2025)

  21. 65.12.02 · Risk Sub-Category

    Inference risks (Intellectual property)

    IP information in prompt

    "Copyrighted information or other intellectual property might be included as a part of the prompt that is sent to the model."

    From AI Risk Atlas (IBM2025)

  22. 65.16.02 · Risk Sub-Category

    Output risks (Intellectual Property)

    Revealing confidential information

    "When confidential information is used in training data, fine-tuning data, or as part of the prompt, models might reveal that data in the generated output. Revealing confidential information is a type of data leakage."

    From AI Risk Atlas (IBM2025)

  23. 65.20.01 · Risk Sub-Category

    Output risks (Privacy)

    Exposing personal information

    "When personal identifiable information (PII) or sensitive personal information (SPI) are used in training data, fine-tuning data, or as part of the prompt, models might reveal that data in the generated output. Revealing personal information is a type of data leakage."

    From AI Risk Atlas (IBM2025)

  24. 66.08.03 · Risk Sub-Category

    Financial and Business

    Confidentiality loss

    "Unauthorised sharing of sensitive, confidential information and documents such as corporate strategy and financial plans with third-parties, risking loss of market position or revenue"

    From A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025)

  25. 66.09.01 · Risk Sub-Category

    Privacy and Security

    Exclusion

    "The failure to provide end-users with notice and control over how their data is being used; AI exacerbates exclusion risks by training on rich personal data without consent."

    From A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025)

  26. 66.09.03 · Risk Sub-Category

    Privacy and Security

    Disclosure

    "Revealing and improperly sharing data of individuals; AI creates new types of disclosure risks by inferring additional information beyond what is explicitly captured in the raw data; AI exacerbates disclosure risks through sharing personal data to train models."

    From A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025)

  27. 66.09.04 · Risk Sub-Category

    Privacy and Security

    Secondary use

    "The use of personal data collected for one purpose for a diferent purpose without end-user consent; AI exacerbates secondary use risks by creating new AI capabilities with collected personal data, and (re)creating models from a public dataset."

    From A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025)

  28. 66.09.05 · Risk Sub-Category

    Privacy and Security

    Exposure

    "Revealing sensitive private information that people view as deeply primordial that we have been socialized into concealing; AI creates new types of exposure risks through generative techniques that can reconstruct censored or redacted content; and through exposing inferred sensitive data, preferences, and intentions."

    From A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025)

  29. 66.09.07 · Risk Sub-Category

    Privacy and Security

    Insecurity

    "carelessness in protecting collected personal data from leaks and improper access due to faulty data storage and data practices"

    From A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025)

  30. 69.05.00 · Risk Category

    Leakage

    "The chatbot reveals sensitive or confidential information."

    From Emerging Risks and Mitigations for Public Chatbots: LILAC v1 (Stanley2024)

  31. 69.05.01 · Risk Sub-Category

    Leakage

    Personal data

    Negative outcomes: "Violation of privacy [106, 516, 357], lawsuit against maker"

    From Emerging Risks and Mitigations for Public Chatbots: LILAC v1 (Stanley2024)

  32. 69.05.02 · Risk Sub-Category

    Leakage

    Proprietary data

    "Access to sensitive company data [473]"

    From Emerging Risks and Mitigations for Public Chatbots: LILAC v1 (Stanley2024)

  33. 69.09.03 · Risk Sub-Category

    Forms emotional bonds

    Elicits private data

  34. 70.02.01 · Risk Sub-Category

    Informational Risks

    Privacy Violations

    "EAI systems interact with huge amounts of data, creating significant privacy concerns. These systems are often trained on vast corpora and process a variety of data modalities— spanning visual, auditory, and tactile information—during deployment [12]. Like text-based virtual AI models, which are known to memorize and expose personally identifiable information [75, 76], commercial robots have been shown to disclose proprietary information through simple prompts [61]."

    From Embodied AI: Emerging Risks and Opportunities for Policy Action (Perlo2025)

  35. 71.01.05 · Risk Sub-Category

    Scientific Domain of Agents

    Information Science Risks

    "These risks pertain to the misuse, misinterpretation, or leakage of data, which can lead to erroneous conclusions or the unintentional dissemination of sensitive information, such as private patient data or proprietary research. Recent research has demonstrated how LLMs can be exploited to generate malicious medical literature that poisons knowledge graphs, potentially manipulating downstream biomedical applications and compromising the integrity of medical knowledge discovery [28]. Such risks are pervasive across all scientific domains."

    From Risks of AI Scientists: Prioritizing Safeguarding Over Autonomy (Tang2025)

  36. 02.03.04 · Risk Sub-Category

    Unhelpful Uses

    Software Vulnerabilities

    "Programmers are accustomed to using code generation tools such as Github Copilot for program development, which may bury vulnerabilities in the program."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  37. "The software development toolchain of LLMs is complex and could bring threats to the developed LLM."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  38. 02.04.01 · Risk Sub-Category

    Software Security Issues

    Programming Language

    "Most LLMs are developed using the Python language, whereas the vulnerabilities of Python interpreters pose threats to the developed models"

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  39. 02.04.02 · Risk Sub-Category

    Software Security Issues

    Deep Learning Frameworks

    "LLMs are implemented based on deep learning frameworks. Notably, various vulnerabilities in these frameworks have been disclosed in recent years. As reported in the past five years, three of the most common types of vulnerabilities are buffer overflow attacks, memory corruption, and input validation issues."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  40. 02.04.03 · Risk Sub-Category

    Software Security Issues

    Software Supply Chains

    "The software development toolchain of LLMs is complex and could bring threats to the developed LLM."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  41. 02.04.04 · Risk Sub-Category

    Software Security Issues

    Pre-processing Tools

    "Pre-processing tools play a crucial role in the context of LLMs. These tools, which are often involved in computer vision (CV) tasks, are susceptible to attacks that exploit vulnerabilities in tools such as OpenCV."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  42. "The vulnerabilities of hardware systems for training and inferencing brings issues to LLM-based applications."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  43. 02.05.01 · Risk Sub-Category

    Hardware Vulnerabilities

    Network Devices

    "The training of LLMs often relies on distributed network systems [171], [172]. During the transmission of gradients through the links between GPU server nodes, significant volumetric traffic is generated. This traffic can be susceptible to disruption by burst traffic, such as pulsating attacks [161]. Furthermore, distributed training frameworks may encounter congestion issues [173]."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  44. 02.05.02 · Risk Sub-Category

    Hardware Vulnerabilities

    GPU Computation Platforms

    "The training of LLMs requires significant GPU resources, thereby introducing an additional security concern. GPU side-channel attacks have been developed to extract the parameters of trained models [159], [163]."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  45. 02.05.03 · Risk Sub-Category

    Hardware Vulnerabilities

    Memory and Storage

    "Similar to conventional programs, hardware infrastructures can also introduce threats to LLMs. Memory-related vulnerabilities, such as rowhammer attacks [160], can be leveraged to manipulate the parameters of LLMs, giving rise to attacks such as the Deephammer attack [167], [168]."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  46. "The external tools (e.g., web APIs) present trustworthiness and privacy issues to LLM-based applications."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  47. 02.06.01 · Risk Sub-Category

    Issues on External Tools

    Factual Errors Injected by External Tools

    "External tools typically incorporate additional knowledge into the input prompts [122], [178]–[184]. The additional knowledge often originates from public resources such as Web APIs and search engines. As the reliability of external tools is not always ensured, the content returned by external tools may include factual errors, consequently amplifying the hallucination issue."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  48. 02.06.02 · Risk Sub-Category

    Issues on External Tools

    Exploiting External Tools for Attacks

    "Adversarial tool providers can embed malicious instructions in the APIs or prompts [84], leading LLMs to leak memorized sensitive information in the training data or users’ prompts (CVE2023-32786). As a result, LLMs lack control over the output, resulting in sensitive information being disclosed to external tool providers. Besides, attackers can easily manipulate public data to launch targeted attacks, generating specific malicious outputs according to user inputs. Furthermore, feeding the information from external tools into LLMs may lead to injection attacks [61]. For example, unverified in

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.