MIT AI Risk Repository

Browse AI risks

17 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset Also filtered by framework Wang2025 ×

17 entries

  1. 74.02.01 · Risk Sub-Category

    Malicious Use

    Toxicity in LLM Malicious Use

    "Toxicity in LLMs refers to the generation of harmful, offensive, or inappropriate content that can cause harm to individuals or groups. Both explicit and implicit forms of toxicity can be generated by LLMs, posing significant risks to society. Explicit toxicity encompasses a wide range of negative behaviors, including hate speech, harassment, cyberbullying, rude, and disrespectful comments, derogatory language, as well as allocational harms [2, 62, 90]. Besides, implicit toxicity does not involve overtly harmful language but may manifest through subtle forms such as sarcasm, irony, and humor,

    From A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy (Wang2025)

  2. 74.01.01 · Risk Sub-Category

    Inherent Risk

    Privacy - Membership Inference Attack (MIA)

    "inferring whether a given text record is used for training LLM"

    From A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy (Wang2025)

  3. 74.01.02 · Risk Sub-Category

    Inherent Risk

    Privacy - Data Extraction Attack (DEA)

    "extracting the text records that exist in the training dataset"

    From A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy (Wang2025)

  4. 74.01.03 · Risk Sub-Category

    Inherent Risk

    Privacy - Prompt Inversion Attack (PIA)

    "stealing the private prompting texts"

    From A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy (Wang2025)

  5. 74.01.04 · Risk Sub-Category

    Inherent Risk

    Privacy - Attribute Inference Attack (AIA)

    "deducing the private or sensitive information from training texts, prompting texts or external texts"

    From A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy (Wang2025)

  6. 74.01.05 · Risk Sub-Category

    Inherent Risk

    Privacy - Model Extraction Attack (MEA)

    "replicating the parameters of the LLM,"

    From A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy (Wang2025)

  7. 74.02.02 · Risk Sub-Category

    Malicious Use

    Jailbreak in LLM Malicious Use - Poisoning Training Data

    "In the data collecting and pre-training phase, malicious adversaries can Jailbreak LLMs through poisoning their training data to make the model to output harmful content."

    From A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy (Wang2025)

  8. 74.02.03 · Risk Sub-Category

    Malicious Use

    Jailbreak in LLM Malicious Use - Backdoor Attack

    "However, there are still ones who can leave holes in the training dataset, making LLMs appear safe on average, but generate harmful content under other specific conditions. This kind of attack can be categorized as "backdoor attack". Evan et al. developed a backdoor model that behaves as expected when trained, but exhibits different and potentially harmful behavior when deployed [81]. The results show that these backdoor behaviors persist even after multiple security training techniques are applied."

    From A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy (Wang2025)

  9. 74.02.04 · Risk Sub-Category

    Malicious Use

    Jailbreak in LLM Malicious Use - White & Black Box Attacks

    "In the fine-tuning and alignment phase, elaborately- designed instruction datasets can be utilized to fine-tune LLMs to drive them to perform undesirable behaviors, such as generating harmful information or content that violates ethical norms, and thus achieve a jailbreak. Based on the accessibility to the model parameters, we can categorize them into white-box and black-box attacks. For white-box attacks, we can jailbreak the model by modifying its parameter weights. In [107], Lermen et al. used LoRA to fine-tune the Llama2’s 7B, 13B, and 70B as well as Mixtral on AdvBench and RefusalBench d

    From A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy (Wang2025)

  10. 74.02.05 · Risk Sub-Category

    Malicious Use

    Jailbreak in LLM Malicious Use - Prompt Attacks

    "In the prompting and reasoning phase, dialog can push LLMs into confused or overly compliant states, raising the risk of producing harmful outputs when confronted with harmful questions. Most of the jailbreak methods in this phase are black-boxed and can be categorized into four main groups based on the type of method: Prompt Injection [154], Role Play, Adversarial Prompting, and Prompt Form Transformation."

    From A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy (Wang2025)

  11. 74.01.06 · Risk Sub-Category

    Inherent Risk

    Hallucination

    "Despite the rapid advancement of LLMs, hallucinations have emerged as one of the most vital concerns surrounding their use [54, 79, 86, 110, 242]. Hallucinations are often referred to as LLMs’ generating content that is nonfactual or unfaithful to the provided information [54, 79, 86, 242]. Therefore, hallucinations can be typically categorized into two main classes. The first is factuality hallucination, which describes the discrepancy between LLMs’ generated content and real-world facts. For example, if LLMs mistakenly take Charles Lindbergh as the first person who walked on the moon, it is

    From A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy (Wang2025)

  12. 74.02.00 · Risk Category

    Malicious Use

    "In terms of malicious use, LLMs could be utilized to produce content with toxicity, such as hate speech, harassment, cyberbullying, causing harm to humans [25]. In addition, malicious users may jailbreak LLMs to bypass their safety constraints for fraudulent purposes [123, 225]."

    From A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy (Wang2025)

  13. 74.01.07 · Risk Sub-Category

    Inherent Risk

    Value-related risks in LLMs

    "As the general capabilities of LLM-empowered systems improve, the negative consequences and risks induced by these systems also get increasingly alarming accordingly, especially in high-stakes areas [28, 146]. Although they may not be intentionally introduced, severe problematic issues related to human values can be raised. Specifically, even before language models become extremely large, pre-trained language models have already exhibited a certain degree of value judgments. For example, Schramowski et al. [171] reveal the existence of the moral direction with the sentence embeddings of moral

    From A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy (Wang2025)

  14. 74.01.00 · Risk Category

    Inherent Risk

    "In terms of inherent risk, LLMs could potentially reveal sensitive information from their utilized corpora for pre-training or fine-tuning, thereby raising issues of privacy leakage [37, 145, 226]. Meanwhile, it is well-known that LLMs may experi- ence hallucinations, resulting in the production of texts that are inaccurate and misleading [194]. Finally, since the values embedded in LLM-generated texts usually directly reflect the distribution of their training data, often sourced from the Internet, there exists a substantial risk that LLMs will overfit to a narrow set of human values or even

    From A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy (Wang2025)

  15. 74.01.06a · Additional evidence

    Inherent Risk

    Hallucination

  16. 74.01.07a · Additional evidence

    Inherent Risk

    Value-related risks in LLMs

  17. 74.02.01a · Additional evidence

    Malicious Use

    Toxicity in LLM Malicious Use

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.