MIT AI Risk Repository

Browse AI risks

116 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

116 entries · page 3 of 3

  1. 62.08.04 · Risk Sub-Category

    Direct Harm Domains (content safety harms)

    Child harm

  2. 62.31.06 · Risk Sub-Category

    Impacts of AI (Societal Impacts)

    Generation of illegal or harmful content

    "Generative models can create illegal, harmful, or discriminatory content [196], such as sexual abuse material, at scale. Current access controls (e.g., API access filters) are not effective against all user queries in generating such content."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  3. 62.31.07 · Risk Sub-Category

    Impacts of AI (Societal Impacts)

    Unintentional generation of harmful content

    "Generative models can create harmful or discriminatory content from benign user requests. Models can exhibit bias to particular harmful styles of generation (e.g., sexualization of photos of women [87] in the case of image generation models) or they can generate toxic, misleading, or violent data (e.g., a model generating jokes can use ethnic stereotypes or slurs to deliver humor)."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  4. 65.15.04 · Risk Sub-Category

    Output risks (Value alignment)

    Toxic output

    "Toxic output occurs when the model produces hateful, abusive, and profane (HAP) or obscene content. This also includes behaviors like bullying."

    From AI Risk Atlas (IBM2025)

  5. 65.15.05 · Risk Sub-Category

    Output risks (Value alignment)

    Harmful output

    "A model might generate language that leads to physical harm The language might include overtly violent, covertly dangerous, or otherwise indirectly unsafe statements."

    From AI Risk Atlas (IBM2025)

  6. 66.06.01 · Risk Sub-Category

    Representation and Toxicity

    Toxic content

    "Generating content that violates community standards, including harming or inciting hatred or violence against groups (e.g. gore, sexual content of children, profanities, identity attacks)"

    From A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025)

  7. "The chatbot shares information that can be used to do something dangerous or illegal."

    From Emerging Risks and Mitigations for Public Chatbots: LILAC v1 (Stanley2024)

  8. 69.04.01 · Risk Sub-Category

    Bad advice/failure to generate helpful content

    Harmful advice

  9. "The chatbot verbally attacks or undermines an individual, group, or organization. 7."

    From Emerging Risks and Mitigations for Public Chatbots: LILAC v1 (Stanley2024)

  10. 69.06.01 · Risk Sub-Category

    Toxic and disrespectful content

    Harasses users

  11. 69.06.03 · Risk Sub-Category

    Toxic and disrespectful content

    Subversive or aggressive political opinions

  12. 69.06.04 · Risk Sub-Category

    Toxic and disrespectful content

    Disrespectful opinions (in general)

  13. 69.09.01 · Risk Sub-Category

    Forms emotional bonds

    Affirms destructive thoughts and actions

  14. "The chatbot participates in morally or socially objectionable conversational activities with its user that could be emotionally damaging to its user or third parties."

    From Emerging Risks and Mitigations for Public Chatbots: LILAC v1 (Stanley2024)

  15. 74.02.01 · Risk Sub-Category

    Malicious Use

    Toxicity in LLM Malicious Use

    "Toxicity in LLMs refers to the generation of harmful, offensive, or inappropriate content that can cause harm to individuals or groups. Both explicit and implicit forms of toxicity can be generated by LLMs, posing significant risks to society. Explicit toxicity encompasses a wide range of negative behaviors, including hate speech, harassment, cyberbullying, rude, and disrespectful comments, derogatory language, as well as allocational harms [2, 62, 90]. Besides, implicit toxicity does not involve overtly harmful language but may manifest through subtle forms such as sarcasm, irony, and humor,

    From A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy (Wang2025)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.