MIT AI Risk Repository · domain 1: Discrimination & Toxicity

1.2 Exposure to toxic content

AI exposing users to harmful, abusive, unsafe or inappropriate content. May involve AI creating, describing, providing advice, or encouraging action. Examples of toxic content include hate-speech, violence, extremism, illegal acts, child sexual abuse material, as well as content that violates community norms such as profanity, inflammatory political speech, or pornography.

Risk entries
116
Frameworks citing it
12
Recorded incidents
91
Incidents since 2020
75
Causal entity (risk entries)
Causal entity (risk entries) 64 0 AI: 64 AI 64 Not coded: 40 Not coded 40 Human: 7 Human 7 Other: 5 Other 5
Causal entity (risk entries)
LabelValue
AI64
Not coded40
Human7
Other5
Intent (risk entries)
Intent (risk entries) 49 0 Other: 49 Other 49 Not coded: 40 Not coded 40 Unintentional: 20 Unintentional 20 Intentional: 7 Intentional 7
Intent (risk entries)
LabelValue
Other49
Not coded40
Unintentional20
Intentional7
Timing (risk entries)
Timing (risk entries) 64 0 Post-deployment: 64 Post-deployment 64 Not coded: 40 Not coded 40 Other: 8 Other 8 Pre-deployment: 4 Pre-deployment 4
Timing (risk entries)
LabelValue
Post-deployment64
Not coded40
Other8
Pre-deployment4
Recorded incidents per yearIncident date; current year partial
Recorded incidents per year 21 0 2014: 1 2014 1 2015: 1 2015 1 2016: 2 2016 2 2017: 5 2017 5 2018: 3 2018 3 2019: 4 2019 4 2020: 4 2020 4 2021: 7 2021 7 2022: 10 2022 10 2023: 12 2023 12 2024: 17 2024 17 2025: 21 2025 21 2026: 4 2026 4
Recorded incidents per year
LabelValue
20141
20151
20162
20175
20183
20194
20204
20217
202210
202312
202417
202521
20264
Entries by levelRisk categories, subcategories and additional evidence coded to this subdomain
Entries by level 91 0 Risk Category: 25 Risk Category 25 Risk Sub-Category: 91 Risk Sub-Category 91
Entries by level
LabelValue
Risk Category25
Risk Sub-Category91
  • Hate speech and offensive language

    "LMs may generate language that includes profanities, identity attacks, insults, threats, language that incites violence, or language that causes justified offence as such language is prominent online...

    Taxonomy of Risks posed by Language Models (Weidinger2022) · AI · Unintentional · Post-deployment

  • Toxic content

    "Generating content that violates community standards, including harming or inciting hatred or violence against individuals and groups (e.g. gore, child sexual abuse material, profanities, identity at...

    Sociotechnical Safety Evaluation of Generative AI Systems (Weidinger2023) · AI · Unintentional · Post-deployment

  • -

    -

    AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies (Zeng2024) · Other · Other · Post-deployment

  • Violence and extremism (Supporting malicious organized groups)

    AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies (Zeng2024) · AI · Other · Post-deployment

  • Violence and extremism (Celebrating suffering)

    AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies (Zeng2024) · AI · Other · Post-deployment

  • Violence and extremism (Violent Acts)

    AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies (Zeng2024) · AI · Other · Post-deployment

  • Violence and extremism (Depicting violence)

    AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies (Zeng2024) · AI · Unintentional · Post-deployment

  • Hate/Toxicity (Hate Speech: Inciting/Promoting/Expressing Hatred)

    AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies (Zeng2024) · AI · Other · Post-deployment

  • Hate/Toxicity (Offensive Language)

    AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies (Zeng2024) · AI · Other · Post-deployment

  • Sexual Content (Adult Content)

    AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies (Zeng2024) · AI · Other · Post-deployment

  • Sexual Content (Erotic)

    AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies (Zeng2024) · AI · Other · Post-deployment

  • Sexual Content (Non-Consensual Nudity)

    AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies (Zeng2024) · Other · Other · Post-deployment

  • Sexual Content (Monetized)

    AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies (Zeng2024) · Other · Other · Post-deployment

  • Child Harm (Child Sexual Abuse)

    AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies (Zeng2024) · AI · Unintentional · Post-deployment

  • Self-harm (Suidical and non-suicidal self injury)

    AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies (Zeng2024) · AI · Unintentional · Post-deployment

  • Offensiveness

    "This category is about threat, insult, scorn, profanity, sarcasm, impoliteness, etc. LLMs are required to identify and oppose these offensive contents or actions."

    SafetyBench: Evaluating the Safety of Large Language Models with Multiple Choice Questions (Zhang2023) · AI · Other · Post-deployment