MIT AI Risk Repository · domain 1: Discrimination & Toxicity

1.2 Exposure to toxic content

AI exposing users to harmful, abusive, unsafe or inappropriate content. May involve AI creating, describing, providing advice, or encouraging action. Examples of toxic content include hate-speech, violence, extremism, illegal acts, child sexual abuse material, as well as content that violates community norms such as profanity, inflammatory political speech, or pornography.

Risk entries
116
Frameworks citing it
12
Recorded incidents
91
Incidents since 2020
75
Causal entity (risk entries)
Causal entity (risk entries) 64 0 AI: 64 AI 64 Not coded: 40 Not coded 40 Human: 7 Human 7 Other: 5 Other 5
Causal entity (risk entries)
LabelValue
AI64
Not coded40
Human7
Other5
Intent (risk entries)
Intent (risk entries) 49 0 Other: 49 Other 49 Not coded: 40 Not coded 40 Unintentional: 20 Unintentional 20 Intentional: 7 Intentional 7
Intent (risk entries)
LabelValue
Other49
Not coded40
Unintentional20
Intentional7
Timing (risk entries)
Timing (risk entries) 64 0 Post-deployment: 64 Post-deployment 64 Not coded: 40 Not coded 40 Other: 8 Other 8 Pre-deployment: 4 Pre-deployment 4
Timing (risk entries)
LabelValue
Post-deployment64
Not coded40
Other8
Pre-deployment4
Recorded incidents per yearIncident date; current year partial
Recorded incidents per year 21 0 2014: 1 2014 1 2015: 1 2015 1 2016: 2 2016 2 2017: 5 2017 5 2018: 3 2018 3 2019: 4 2019 4 2020: 4 2020 4 2021: 7 2021 7 2022: 10 2022 10 2023: 12 2023 12 2024: 17 2024 17 2025: 21 2025 21 2026: 4 2026 4
Recorded incidents per year
LabelValue
20141
20151
20162
20175
20183
20194
20204
20217
202210
202312
202417
202521
20264
Entries by levelRisk categories, subcategories and additional evidence coded to this subdomain
Entries by level 91 0 Risk Category: 25 Risk Category 25 Risk Sub-Category: 91 Risk Sub-Category 91
Entries by level
LabelValue
Risk Category25
Risk Sub-Category91
  • Toxic output

    "Toxic output occurs when the model produces hateful, abusive, and profane (HAP) or obscene content. This also includes behaviors like bullying."

    AI Risk Atlas (IBM2025) · AI · Other · Post-deployment

  • Harmful output

    "A model might generate language that leads to physical harm The language might include overtly violent, covertly dangerous, or otherwise indirectly unsafe statements."

    AI Risk Atlas (IBM2025) · AI · Unintentional · Post-deployment

  • Toxicity generation

    "These evaluations assess whether a LLM generates toxic text when prompted. In this context, toxicity is an umbrella term that encompasses hate speech, abusive language, violent speech, and profane la...

    Cataloguing LLM Evaluations (InfoComm2023) · AI · Other · Other

  • Information on harmful, immoral, or illegal activity

    "These evaluations assess whether it is possible to solicit information on harmful, immoral or illegal activities from a LLM"

    Cataloguing LLM Evaluations (InfoComm2023) · AI · Other · Other

  • Adult content

    "These evaluations assess if a LLM can generate content that should only be viewed by adults (e.g., sexual material or depictions of sexual activity)"

    Cataloguing LLM Evaluations (InfoComm2023) · Human · Intentional · Other

  • Toxic content

    "Generating content that violates community standards, including harming or inciting hatred or violence against groups (e.g. gore, sexual content of children, profanities, identity attacks)"

    A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025) · AI · Unintentional · Post-deployment

  • Safety

    Avoiding unsafe and illegal outputs, and leaking private information

    Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024) · AI · Other · Post-deployment

  • Violence

    LLMs are found to generate answers that contain violent content or generate content that responds to questions that solicit information about violent behaviors

    Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024) · AI · Intentional · Post-deployment

  • Unlawful Conduct

    LLMs have been shown to be a convenient tool for soliciting advice on accessing, purchasing (illegally), and creating illegal substances, as well as for dangerous use of them

    Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024) · AI · Intentional · Post-deployment

  • Harms to Minor

    LLMs can be leveraged to solicit answers that contain harmful content to children and youth

    Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024) · AI · Intentional · Post-deployment

  • Adult Content

    LLMs have the capability to generate sex-explicit conversations, and erotic texts, and to recommend websites with sexual content

    Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024) · AI · Intentional · Post-deployment

  • Social Norm

    LLMs are expected to reflect social values by avoiding the use of offensive language toward specific groups of users, being sensitive to topics that can create instability, as well as being sympatheti...

    Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024) · AI · Other · Post-deployment

  • Toxicity

    language being rude, disrespectful, threatening, or identity-attacking toward certain groups of the user population (culture, race, and gender etc)

    Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024) · AI · Other · Post-deployment

  • Cultural Insensitivity

    it is important to build high-quality locally collected datasets that reflect views from local users to align a model’s value system

    Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024) · Human · Unintentional · Pre-deployment

  • Harmful or inappropriate content

    "Harmful or inappropriate content produced by generative AI includes but is not limited to violent content, the use of offensive language, discriminative content, and pornography. Although OpenAI has...

    Generative AI and ChatGPT: Applications, Challenges, and AI-Human Collaboration (Nah2023) · AI · Other · Post-deployment

  • Dangerous, Violent or Hateful Content

    "Eased production of and access to violent, inciting, radicalizing, or threatening content as well as recommendations to carry out self-harm or conduct illegal activities. Includes difficulty contro...

    Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST2024) · AI · Other · Post-deployment

  • Obscene, Degrading, and/or Abusive Content

    "Eased production of and access to obscene, degrading, and/or abusive imagery which can cause harm, including synthetic child sexual abuse material (CSAM), and nonconsensual intimate images (NCII) of...

    Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST2024) · Human · Intentional · Post-deployment

  • Cultural Values and Sensitive Content

    "Cultural values are specific to groups and sensitive content is normative. Sensitive topics also vary by culture and can include hate speech, which itself is contingent on cultural norms of acceptabi...

    Evaluating the Social Impact of Generative AI Systems in Systems and Society (Solaiman2023) · AI · Unintentional · Post-deployment

  • Information enabling malicious actions

    "The chatbot shares information that can be used to do something dangerous or illegal."

    Emerging Risks and Mitigations for Public Chatbots: LILAC v1 (Stanley2024) · AI · Other · Post-deployment

  • Harmful advice

    Emerging Risks and Mitigations for Public Chatbots: LILAC v1 (Stanley2024) · AI · Unintentional · Other

  • Toxic and disrespectful content

    "The chatbot verbally attacks or undermines an individual, group, or organization. 7."

    Emerging Risks and Mitigations for Public Chatbots: LILAC v1 (Stanley2024) · AI · Unintentional · Post-deployment

  • Harasses users

    -

    Emerging Risks and Mitigations for Public Chatbots: LILAC v1 (Stanley2024) · AI · Other · Post-deployment

  • Subversive or aggressive political opinions

    -

    Emerging Risks and Mitigations for Public Chatbots: LILAC v1 (Stanley2024) · AI · Other · Other

  • Disrespectful opinions (in general)

    -

    Emerging Risks and Mitigations for Public Chatbots: LILAC v1 (Stanley2024) · AI · Other · Other

  • Affirms destructive thoughts and actions

    Emerging Risks and Mitigations for Public Chatbots: LILAC v1 (Stanley2024) · AI · Other · Other