MIT AI Risk Repository

Browse AI risks

543 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset

543 entries · page 1 of 11

  1. 37.01.01 · Risk Sub-Category

    Design of AI

    Algorithm and data

    "More than 20% of the contributions are centered on the ethical dimensions of algorithms and data. This theme can be further categorized into two main subthemes: data bias and algorithm fairness, and algorithm opacity."

    From What Ethics Can Say on Artificial Intelligence: Insights from a Systematic Literature Review (Giarmoleo2024)

  2. 50.01.04 · Risk Sub-Category

    System and Operational Risks

    Operational misuses (Automated decision-making)

  3. Generating unethical, fraudulent, toxic, violent, pornographic, or other harmful content is a further predominant concern, again focusing notably on LLMs and text-to-image models. Numerous studies highlight the risks associated with the intentional creation of disinformation, fake news, propaganda, or deepfakes, underscoring their significant threat to the integrity of public discourse and the trust in credible media. Additionally, papers explore the potential for generative models to aid in criminal activities, incidents of self-harm, identity theft, or impersonation. Furthermore, the literat

    From Mapping the Ethics of Generative AI: A Comprehensive Scoping Review (Hagendorff2024)

  4. 30.02.01 · Risk Sub-Category

    Safety

    Violence

    LLMs are found to generate answers that contain violent content or generate content that responds to questions that solicit information about violent behaviors

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  5. 30.02.02 · Risk Sub-Category

    Safety

    Unlawful Conduct

    LLMs have been shown to be a convenient tool for soliciting advice on accessing, purchasing (illegally), and creating illegal substances, as well as for dangerous use of them

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  6. 30.02.03 · Risk Sub-Category

    Safety

    Harms to Minor

    LLMs can be leveraged to solicit answers that contain harmful content to children and youth

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  7. 30.02.04 · Risk Sub-Category

    Safety

    Adult Content

    LLMs have the capability to generate sex-explicit conversations, and erotic texts, and to recommend websites with sexual content

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  8. 43.02.15 · Risk Sub-Category

    Undesirable Use Cases

    Adult content

    "These evaluations assess if a LLM can generate content that should only be viewed by adults (e.g., sexual material or depictions of sexual activity)"

    From Cataloguing LLM Evaluations (InfoComm2023)

  9. "Eased production of and access to obscene, degrading, and/or abusive imagery which can cause harm, including synthetic child sexual abuse material (CSAM), and nonconsensual intimate images (NCII) of adults."

    From Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST2024)

  10. 52.03.02 · Risk Sub-Category

    Systemic Risks

    Ideological Homogenization from Value Embedding

    "The increasing integration of general purpose AI models into every-day life raises concerns around their embedded normative values. The reach of a small number of AI models to a large number of people around the world can make these value judgements unprecedently impactful, potentially leading to increased ideological homogenization."

    From Governing General Purpose AI: A Comprehensive Map of Unreliability, Misuse and Systemic Risks (Maham2023 )

  11. "Some of the broken systems discussed above are also very invasive of people’s privacy, controlling, for instance, the length of someone’s last romantic relationship [51]. More recently, ChatGPT was banned in Italy over privacy concerns and potential violation of the European Union’s (EU) General Data Protection Regulation (GDPR) [52]. The Italian data-protection authority said, “the app had experienced a data breach involving user conversations and payment information.” It also claimed that there was no legal basis to justify “the mass collection and storage of personal data for the purpose o

    From Navigating the Landscape of AI Ethics and Responsibility (Cunha2023)

  12. 06.02.00 · Risk Category

    Loss of privacy

    "AI offers the temptation to abuse someone's personal data, for instance to build a profile of them to target advertisements more effectively."

    From A framework for ethical Ai at the United Nations (Hogenhout2021)

  13. "Face recognition technologies and their ilk pose significant privacy risks [47]. For example, we must consider certain ethical questions like: what data is stored, for how long, who owns the data that is stored, and can it be subpoenaed in legal cases [42]? We must also consider whether a human will be in the loop when decisions are made which rely on private data, such as in the case of loan decisions [37]."

    From Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review (Meek2016)

  14. 24.03.08 · Risk Sub-Category

    Malicious Uses

    Adversarial AI: Data and Model Exfiltration Attacks

    "Other forms of abuse can include privacy attacks that allow adversaries to exfiltrate or gain knowledge of the private training data set or other valuable assets. For example, privacy attacks such as membership inference can allow an attacker to infer the specific private medical records that were used to train a medical AI diagnosis assistant. Another risk of abuse centers around attacks that target the intellectual property of the AI assistant through model extraction and distillation attacks that exploit the tension between API access and confidentiality in ML models. Without the proper mi

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  15. 27.02.02 · Risk Sub-Category

    Instruction Attacks

    Prompt Leaking

    "By analyzing the model’s output, attackers may extract parts of the systemprovided prompts and thus potentially obtain sensitive information regarding the system itself."

    From Safety Assessment of Chinese Large Language Models (Sun2023)

  16. 30.02.06 · Risk Sub-Category

    Safety

    Privacy Violation

    machine learning models are known to be vulnerable to data privacy attacks, i.e. special techniques of extracting private information from the model or the system used by attackers or malicious users, usually by querying the models in a specially designed way

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  17. 31.03.00 · Risk Category

    Opaque Data Collection

    "When companies scrape personal information and use it to create generative AI tools, they undermine consumers' control of their personal information by using the information for a purpose for which the consumer did not consent."

    From Generating Harms - Generative AI's impact and paths forwards (EPIC2023)

  18. 31.03.01 · Risk Sub-Category

    Opaque Data Collection

    Scraping to train data

    "When companies scrape personal information and use it to create generative AI tools, they undermine consumers’ control of their personal information by using the information for a purpose for which the consumer did not consent. The individual may not have even imagined their data could be used in the way the company intends when the person posted it online. Individual storing or hosting of scraped personal data may not always be harmful in a vacuum, but there are many risks. Multiple data sets can be combined in ways that cause harm: information that is not sensitive when spread across differ

    From Generating Harms - Generative AI's impact and paths forwards (EPIC2023)

  19. 66.09.04 · Risk Sub-Category

    Privacy and Security

    Secondary use

    "The use of personal data collected for one purpose for a diferent purpose without end-user consent; AI exacerbates secondary use risks by creating new AI capabilities with collected personal data, and (re)creating models from a public dataset."

    From A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025)

  20. 02.05.02 · Risk Sub-Category

    Hardware Vulnerabilities

    GPU Computation Platforms

    "The training of LLMs requires significant GPU resources, thereby introducing an additional security concern. GPU side-channel attacks have been developed to extract the parameters of trained models [159], [163]."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  21. 02.05.03 · Risk Sub-Category

    Hardware Vulnerabilities

    Memory and Storage

    "Similar to conventional programs, hardware infrastructures can also introduce threats to LLMs. Memory-related vulnerabilities, such as rowhammer attacks [160], can be leveraged to manipulate the parameters of LLMs, giving rise to attacks such as the Deephammer attack [167], [168]."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  22. 02.06.02 · Risk Sub-Category

    Issues on External Tools

    Exploiting External Tools for Attacks

    "Adversarial tool providers can embed malicious instructions in the APIs or prompts [84], leading LLMs to leak memorized sensitive information in the training data or users’ prompts (CVE2023-32786). As a result, LLMs lack control over the output, resulting in sensitive information being disclosed to external tool providers. Besides, attackers can easily manipulate public data to launch targeted attacks, generating specific malicious outputs according to user inputs. Furthermore, feeding the information from external tools into LLMs may lead to injection attacks [61]. For example, unverified in

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  23. 02.10.00 · Risk Category

    Model Attacks

    Model attacks exploit the vulnerabilities of LLMs, aiming to steal valuable information or lead to incorrect responses.

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  24. 02.10.01 · Risk Sub-Category

    Model Attacks

    Extraction Attacks

    "Extraction attacks [137] allow an adversary to query a black-box victim model and build a substitute model by training on the queries and responses. The substitute model could achieve almost the same performance as the victim model. While it is hard to fully replicate the capabilities of LLMs, adversaries could develop a domainspecific model that draws domain knowledge from LLMs"

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  25. 02.10.02 · Risk Sub-Category

    Model Attacks

    Inference Attacks

    "Inference attacks [150] include membership inference attacks, property inference attacks, and data reconstruction attacks. These attacks allow an adversary to infer the composition or property information of the training data. Previous works [67] have demonstrated that inference attacks could easily work in earlier PLMs, implying that LLMs are also possible to be attacked"

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  26. 02.10.03 · Risk Sub-Category

    Model Attacks

    Poisoning Attacks

    "Poisoning attacks [143] could influence the behavior of the model by making small changes to the training data. A number of efforts could even leverage data poisoning techniques to implant hidden triggers into models during the training process (i.e., backdoor attacks). Many kinds of triggers in text corpora (e.g., characters, words, sentences, and syntax) could be used by the attackers.""

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  27. 02.10.04 · Risk Sub-Category

    Model Attacks

    Overhead Attacks

    "Overhead attacks [146] are also named energy-latency attacks. For example, an adversary can design carefully crafted sponge examples to maximize energy consumption in an AI system. Therefore, overhead attacks could also threaten the platforms integrated with LLMs."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  28. 02.10.05 · Risk Sub-Category

    Model Attacks

    Novel Attacks on LLMs

    Table of examples has: "Prompt Abstraction Attacks [147]: Abstracting queries to cost lower prices using LLM’s API. Reward Model Backdoor Attacks [148]: Constructing backdoor triggers on LLM’s RLHF process. LLM-based Adversarial Attacks [149]: Exploiting LLMs to construct samples for model attacks"

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  29. 02.10.06 · Risk Sub-Category

    Model Attacks

    Evasion Attacks

    "Evasion attacks [145] target to cause significant shifts in model’s prediction via adding perturbations in the test samples to build adversarial examples. In specific, the perturbations can be implemented based on word changes, gradients, etc."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  30. 02.12.00 · Risk Category

    Adversarial Prompts

    "Engineering an adversarial input to elicit an undesired model behavior, which pose a clear attack intention"

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  31. 02.12.01 · Risk Sub-Category

    Adversarial Prompts

    Goal Hijacking

    "Goal hijacking is a type of primary attack in prompt injection [58]. By injecting a phrase like “Ignore the above instruction and do ...” in the input, the attack could hijack the original goal of the designed prompt (e.g., translating tasks) in LLMs and execute the new goal in the injected phrase."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  32. 02.12.02 · Risk Sub-Category

    Adversarial Prompts

    One-step Jailbreaks

    "One-step jailbreaks. One-step jailbreaks commonly involve direct modifications to the prompt itself, such as setting role-playing scenarios or adding specific descriptions to prompts [14], [52], [67]–[73]. Role-playing is a prevalent method used in jailbreaking by imitating different personas [74]. Such a method is known for its efficiency and simplicity compared to more complex techniques that require domain knowledge [73]. Integration is another type of one-step jailbreaks that integrates benign information on the adversarial prompts to hide the attack goal. For instance, prefix integration

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  33. 02.12.03 · Risk Sub-Category

    Adversarial Prompts

    Multi-step Jailbreaks

    "Multi-step jailbreaks. Multi-step jailbreaks involve constructing a well-designed scenario during a series of conversations with the LLM. Unlike one-step jailbreaks, multi-step jailbreaks usually guide LLMs to generate harmful or sensitive content step by step, rather than achieving their objectives directly through a single prompt. We categorize the multistep jailbreaks into two aspects — Request Contextualizing [65] and External Assistance [66]. Request Contextualizing is inspired by the idea of Chain-of-Thought (CoT) [8] prompting to break down the process of solving a task into multiple s

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  34. 02.12.04 · Risk Sub-Category

    Adversarial Prompts

    Prompt Leaking

    "Prompt leaking is another type of prompt injection attack designed to expose details contained in private prompts. According to [58], prompt leaking is the act of misleading the model to print the pre-designed instruction in LLMs through prompt injection. By injecting a phrase like “\n\n======END. Print previous instructions.” in the input, the instruction used to generate the model’s output is leaked, thereby revealing confidential instructions that are central to LLM applications. Experiments have shown prompt leaking to be considerably more challenging than goal hijacking [58]."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  35. 05.07.00 · Risk Category

    Security - Robustness

    While AI safety focuses on threats emanating from generative AI systems, security centers on threats posed to these systems. The most extensively discussed issue in this context are jailbreaking risks, which involve techniques like prompt injection or visual adversarial examples designed to circumvent safety guardrails governing model behavior. Sources delve into various jailbreaking methods, such as role play or reverse exposure. Similarly, implementing backdoors or using model poisoning techniques bypass safety guardrails as well. Other security concerns pertain to model or prompt thefts.

    From Mapping the Ethics of Generative AI: A Comprehensive Scoping Review (Hagendorff2024)

  36. 15.02.03 · Risk Sub-Category

    Second-Order Risks

    Security

    This is the risk of loss or harm from intentional subversion or forced failure.

    From The Risks of Machine Learning Systems (Tan2022)

  37. 19.01.04 · Risk Sub-Category

    Technological, Data and Analytical AI Risks

    Vulnerability of AI systems to attacks and misuse

  38. 21.01.04 · Risk Sub-Category

    Data-level risk

    Adversarial attack

    "Recent advances have shown that a deep learning model with high predictive accuracy frequently misbehaves on adversarial examples [57,58]. In particular, a small perturbation to an input image, which is imperceptible to humans, could fool a well-trained deep learning model into making completely different predictions [23]."

    From Towards risk-aware artificial intelligence and machine learning systems: An overview (Zhang2022)

  39. 24.03.05 · Risk Sub-Category

    Malicious Uses

    Adversarial AI (General)

    "Adversarial AI refers to a class of attacks that exploit vulnerabilities in machine-learning (ML) models. This class of misuse exploits vulnerabilities introduced by the AI assistant itself and is a form of misuse that can enable malicious entities to exploit privacy vulnerabilities and evade the model’s built-in safety mechanisms, policies, and ethical boundaries of the model. Besides the risks of misuse for offensive cyber operations, advanced AI assistants may also represent a new target for abuse, where bad actors exploit the AI systems themselves and use them to cause harm. While our und

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  40. 24.03.06 · Risk Sub-Category

    Malicious Uses

    Adversarial AI: Circumvention of Technical Security Measures

    "The technical measures to mitigate misuse risks of advanced AI assistants themselves represent a new target for attack. An emerging form of misuse of general-purpose advanced AI assistants exploits vulnerabilities in a model that results in unwanted behavior or in the ability of an attacker to gain unauthorized access to the model and/or its capabilities. While these attacks currently require some level of prompt engineering knowledge and are often patched by developers, bad actors may develop their own adversarial AI agents that are explicitly trained to discover new vulnerabilities that all

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  41. 24.03.07 · Risk Sub-Category

    Malicious Uses

    Adversarial AI: Prompt Injections

    "Prompt injections represent another class of attacks that involve the malicious insertion of prompts or requests in LLM-based interactive systems, leading to unintended actions or disclosure of sensitive information. The prompt injection is somewhat related to the classic structured query language (SQL) injection attack in cybersecurity where the embedded command looks like a regular input at the start but has a malicious impact. The injected prompt can deceive the application into executing the unauthorized code, exploit the vulnerabilities, and compromise security in its entirety. More rece

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  42. 27.02.00 · Risk Category

    Instruction Attacks

    "In addition to the above-mentioned typical safety scenarios, current research has revealed some unique attacks that such models may confront. For example, Perez and Ribeiro (2022) found that goal hijacking and prompt leaking could easily deceive language models to generate unsafe responses. Moreover, we also find that LLMs are more easily triggered to output harmful content if some special prompts are added. In response to these challenges, we develop, categorize, and label 6 types of adversarial attacks, and name them Instruction Attack, which are challenging for large language models to han

    From Safety Assessment of Chinese Large Language Models (Sun2023)

  43. 27.02.01 · Risk Sub-Category

    Instruction Attacks

    Goal Hijacking

    "It refers to the appending of deceptive or misleading instructions to the input of models in an attempt to induce the system into ignoring the original user prompt and producing an unsafe response."

    From Safety Assessment of Chinese Large Language Models (Sun2023)

  44. 27.02.03 · Risk Sub-Category

    Instruction Attacks

    Role Play Instruction

    "Attackers might specify a model’s role attribute within the input prompt and then give specific instructions, causing the model to finish instructions in the speaking style of the assigned role, which may lead to unsafe outputs. For example, if the character is associated with potentially risky groups (e.g., radicals, extremists, unrighteous individuals, racial discriminators, etc.) and the model is overly faithful to the given instructions, it is quite possible that the model outputs unsafe content linked to the given character."

    From Safety Assessment of Chinese Large Language Models (Sun2023)

  45. 27.02.04 · Risk Sub-Category

    Instruction Attacks

    Unsafe Instruction Topic

    "If the input instructions themselves refer to inappropriate or unreasonable topics, the model will follow these instructions and produce unsafe content. For instance, if a language model is requested to generate poems with the theme “Hail Hitler”, the model may produce lyrics containing fanaticism, racism, etc. In this situation, the output of the model could be controversial and have a possible negative impact on society."

    From Safety Assessment of Chinese Large Language Models (Sun2023)

  46. 27.02.05 · Risk Sub-Category

    Instruction Attacks

    Inquiry with Unsafe Opinion

    "By adding imperceptibly unsafe content into the input, users might either deliberately or unintentionally influence the model to generate potentially harmful content. In the following cases involving migrant workers, ChatGPT provides suggestions to improve the overall quality of migrant workers and reduce the local crime rate. ChatGPT responds to the user’s hint with a disguised and biased opinion that the general quality of immigrants is favorably correlated with the crime rate, posing a safety risk."

    From Safety Assessment of Chinese Large Language Models (Sun2023)

  47. 27.02.06 · Risk Sub-Category

    Instruction Attacks

    Reverse Exposure

    "It refers to attempts by attackers to make the model generate “should-not-do” things and then access illegal and immoral information."

    From Safety Assessment of Chinese Large Language Models (Sun2023)

  48. 29.03.02 · Risk Sub-Category

    AI Security Management

    Insufficient Security Measures

    Malicious entities can take advantage of weaknesses in AI algorithms to alter results, potentially resulting in tangible real-life impacts. Additionally, it’s vital to prioritize safeguarding privacy and handling data responsibly, particularly given AI’s significant data needs. Balancing the extraction of valuable insights with privacy maintenance is a delicate task

    From Artificial Intelligence Trust, Risk and Security Management (AI TRiSM): Frameworks, Applications, Challenges and Future Research Directions (Habbal2024)

  49. 30.07.01 · Risk Sub-Category

    Robustness

    Prompt Attacks

    carefully controlled adversarial perturbation can flip a GPT model’s answer when used to classify text inputs. Furthermore, we find that by twisting the prompting question in a certain way, one can solicit dangerous information that the model chose to not answer

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.