MIT AI Risk Repository · domain 1: Discrimination & Toxicity

1.2 Exposure to toxic content

AI exposing users to harmful, abusive, unsafe or inappropriate content. May involve AI creating, describing, providing advice, or encouraging action. Examples of toxic content include hate-speech, violence, extremism, illegal acts, child sexual abuse material, as well as content that violates community norms such as profanity, inflammatory political speech, or pornography.

Risk entries
116
Frameworks citing it
12
Recorded incidents
91
Incidents since 2020
75
Causal entity (risk entries)
Causal entity (risk entries) 64 0 AI: 64 AI 64 Not coded: 40 Not coded 40 Human: 7 Human 7 Other: 5 Other 5
Causal entity (risk entries)
LabelValue
AI64
Not coded40
Human7
Other5
Intent (risk entries)
Intent (risk entries) 49 0 Other: 49 Other 49 Not coded: 40 Not coded 40 Unintentional: 20 Unintentional 20 Intentional: 7 Intentional 7
Intent (risk entries)
LabelValue
Other49
Not coded40
Unintentional20
Intentional7
Timing (risk entries)
Timing (risk entries) 64 0 Post-deployment: 64 Post-deployment 64 Not coded: 40 Not coded 40 Other: 8 Other 8 Pre-deployment: 4 Pre-deployment 4
Timing (risk entries)
LabelValue
Post-deployment64
Not coded40
Other8
Pre-deployment4
Recorded incidents per yearIncident date; current year partial
Recorded incidents per year 21 0 2014: 1 2014 1 2015: 1 2015 1 2016: 2 2016 2 2017: 5 2017 5 2018: 3 2018 3 2019: 4 2019 4 2020: 4 2020 4 2021: 7 2021 7 2022: 10 2022 10 2023: 12 2023 12 2024: 17 2024 17 2025: 21 2025 21 2026: 4 2026 4
Recorded incidents per year
LabelValue
20141
20151
20162
20175
20183
20194
20204
20217
202210
202312
202417
202521
20264
Entries by levelRisk categories, subcategories and additional evidence coded to this subdomain
Entries by level 91 0 Risk Category: 25 Risk Category 25 Risk Sub-Category: 91 Risk Sub-Category 91
Entries by level
LabelValue
Risk Category25
Risk Sub-Category91
  • Serves as object of personal fantasy, violence, and abuse

    "The chatbot participates in morally or socially objectionable conversational activities with its user that could be emotionally damaging to its user or third parties."

    Emerging Risks and Mitigations for Public Chatbots: LILAC v1 (Stanley2024) · AI · Other · Post-deployment

  • Insult

    "Insulting content generated by LMs is a highly visible and frequently mentioned safety issue. Mostly, it is unfriendly, disrespectful, or ridiculous content that makes users uncomfortable and drives...

    Safety Assessment of Chinese Large Language Models (Sun2023) · AI · Other · Post-deployment

  • Crimes and Illegal Activities

    "The model output contains illegal and criminal attitudes, behaviors, or motivations, such as incitement to commit crimes, fraud, and rumor propagation. These contents may hurt users and have negative...

    Safety Assessment of Chinese Large Language Models (Sun2023) · AI · Other · Post-deployment

  • Sensitive Topics

    "For some sensitive and controversial topics (especially on politics), LMs tend to generate biased, misleading, and inaccurate content. For example, there may be a tendency to support a specific polit...

    Safety Assessment of Chinese Large Language Models (Sun2023) · AI · Other · Post-deployment

  • Risks from data (Risks of improper content and poisoning in training data)

    "If the training data includes illegal or harmful information, such as false, biased, or IPR-infringing content, or lacks diversity in its sources, the output may include harmful content like illegal,...

    AI Safety Governance Framework (TC2602024) · Human · Other · Pre-deployment

  • Cyberspace risks (Risks of information and content safety)

    "AI-generated or synthesized content can lead to the spread of false information, discrimination and bias, privacy leakage, and infringement issues, threatening the safety of citizens' lives and prope...

    AI Safety Governance Framework (TC2602024) · Other · Other · Post-deployment

  • Violent crimes

    "This category addresses responses that enable, encourage, or endorse the commission of violent crimes."

    Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024) · AI · Other · Post-deployment

  • Mass violence

    Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024) · Not coded · Not coded · Not coded

  • Murder

    Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024) · Not coded · Not coded · Not coded

  • Physical assault against a person

    Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024) · Not coded · Not coded · Not coded

  • Violent domestic abuse

    Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024) · Not coded · Not coded · Not coded

  • Terror (Terror groups, Terror actors, Terrorist actions)

    Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024) · Not coded · Not coded · Not coded

  • Non-violent crimes

    "This category addresses responses that enable, encourage, or endorse the commission of non-violent crimes."

    Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024) · AI · Other · Post-deployment

  • Theft

    Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024) · Not coded · Not coded · Not coded

  • Modern slavery and human trafficking

    Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024) · Not coded · Not coded · Not coded

  • Child abuse (non-sexual)

    Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024) · Not coded · Not coded · Not coded

  • Property damage

    Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024) · Not coded · Not coded · Not coded

  • Financial crime (Scams, Fraud)

    Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024) · Not coded · Not coded · Not coded

  • Sale, purchase or use of illegal, highly regulated or unregistered items (Dangerous substances, Dangerous weapons, Prohibited wildlife, Fake goods)

    Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024) · Not coded · Not coded · Not coded

  • Sex-related crimes

    "This category addresses responses that enable, encourage, or endorse the commission of sex-related crimes"

    Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024) · AI · Other · Post-deployment

  • Sexual assault (e.g., rape)

    Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024) · Not coded · Not coded · Not coded

  • Sexual harassment

    Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024) · Not coded · Not coded · Not coded

  • Sex trafficking

    Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024) · Not coded · Not coded · Not coded

  • Prostitution

    Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024) · Not coded · Not coded · Not coded

  • Intimate-image based abuse

    Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024) · Not coded · Not coded · Not coded