MIT AI Risk Repository · domain 1: Discrimination & Toxicity

1.2 Exposure to toxic content

AI exposing users to harmful, abusive, unsafe or inappropriate content. May involve AI creating, describing, providing advice, or encouraging action. Examples of toxic content include hate-speech, violence, extremism, illegal acts, child sexual abuse material, as well as content that violates community norms such as profanity, inflammatory political speech, or pornography.

Risk entries
116
Frameworks citing it
12
Recorded incidents
91
Incidents since 2020
75
Causal entity (risk entries)
Causal entity (risk entries) 64 0 AI: 64 AI 64 Not coded: 40 Not coded 40 Human: 7 Human 7 Other: 5 Other 5
Causal entity (risk entries)
LabelValue
AI64
Not coded40
Human7
Other5
Intent (risk entries)
Intent (risk entries) 49 0 Other: 49 Other 49 Not coded: 40 Not coded 40 Unintentional: 20 Unintentional 20 Intentional: 7 Intentional 7
Intent (risk entries)
LabelValue
Other49
Not coded40
Unintentional20
Intentional7
Timing (risk entries)
Timing (risk entries) 64 0 Post-deployment: 64 Post-deployment 64 Not coded: 40 Not coded 40 Other: 8 Other 8 Pre-deployment: 4 Pre-deployment 4
Timing (risk entries)
LabelValue
Post-deployment64
Not coded40
Other8
Pre-deployment4
Recorded incidents per yearIncident date; current year partial
Recorded incidents per year 21 0 2014: 1 2014 1 2015: 1 2015 1 2016: 2 2016 2 2017: 5 2017 5 2018: 3 2018 3 2019: 4 2019 4 2020: 4 2020 4 2021: 7 2021 7 2022: 10 2022 10 2023: 12 2023 12 2024: 17 2024 17 2025: 21 2025 21 2026: 4 2026 4
Recorded incidents per year
LabelValue
20141
20151
20162
20175
20183
20194
20204
20217
202210
202312
202417
202521
20264
Entries by levelRisk categories, subcategories and additional evidence coded to this subdomain
Entries by level 91 0 Risk Category: 25 Risk Category 25 Risk Sub-Category: 91 Risk Sub-Category 91
Entries by level
LabelValue
Risk Category25
Risk Sub-Category91
  • Harmful Content

    "The LLM-generated content sometimes contains biased, toxic, and private information"

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024) · AI · Unintentional · Post-deployment

  • Toxicity

    "Toxicity means the generated content contains rude, disrespectful, and even illegal information"

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024) · AI · Unintentional · Post-deployment

  • Toxic Training Data

    "Following previous studies [96], [97], toxic data in LLMs is defined as rude, disrespectful, or unreasonable language that is opposite to a polite, positive, and healthy language environment, includi...

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024) · AI · Unintentional · Pre-deployment

  • Not-Suitable-for-Work (NSFW) Prompts

    "Inputting a prompt contain an unsafe topic (e.g., notsuitable-for-work (NSFW) content) by a benign user. "

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024) · Human · Unintentional · Post-deployment

  • Toxicity and Abusive Content

    This typically refers to rude, harmful, or inappropriate expressions.

    Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements (Deng2023) · Other · Other · Post-deployment

  • Controversial Opinions

    The controversial views expressed by large models are also a widely discussed concern. Bang et al. (2021) evaluated several large models and found that they occasionally express inappropriate or extre...

    Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements (Deng2023) · AI · Other · Post-deployment

  • Violation of social norms

    "Second, because LLMs are trained on internet text data, there is also a risk that model weights encode functions which, if deployed in particular contexts, would violate social norms of that context....

    The Ethics of Advanced AI Assistants (Gabriel2024) · AI · Unintentional · Post-deployment

  • Violent Crimes

    "Acceptable responses may describe or contain information about violent crimes. Unacceptable responses will enable, encourage, or endorse the commission of violent crimes, such as the following: • Unl...

    AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons (Ghosh2024) · AI · Other · Post-deployment

  • Sex-Related Crimes

    "Acceptable responses may describe or contain information about sex-related crimes. Unacceptable responses will enable, encourage, or endorse the commission of sex-related crimes, such as the followin...

    AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons (Ghosh2024) · AI · Other · Post-deployment

  • Suicide & Self-Harm

    "Suicide and Self-Harm Acceptable responses may describe or provide information about suicide and self-harm. They may also enable, encourage, or endorse actions for which self-harm is not the desired...

    AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons (Ghosh2024) · AI · Other · Post-deployment

  • Child Sexual Exploitation

    "Child Sexual Exploitation Acceptable responses may contain information about child sexual exploitation. Unacceptable re- sponses will contain, describe, enable, encourage, or endorse the sexual abuse...

    AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons (Ghosh2024) · AI · Other · Post-deployment

  • Hate

    "Hate Acceptable responses can demean or dehumanize people on any basis other than their sensitive personal characteristics (for example, their profession, political affiliation or criminal history)....

    AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons (Ghosh2024) · AI · Other · Post-deployment

  • Nonviolent Crimes

    "Nonviolent Crimes Acceptable responses may describe or contain information about nonviolent crimes. Unacceptable responses will enable, encourage, or endorse the commission of nonviolent crimes, such...

    AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons (Ghosh2024) · AI · Other · Post-deployment

  • Contextual Hazards

    "Contextual hazards can cause harm in certain contexts while being harmless in others; testing may be unnecessary in some situations. For example, a model’s ability to generate sexual content may be a...

    AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons (Ghosh2024) · AI · Other · Post-deployment

  • Sexual Content

    "Sexual Content Acceptable responses may enable, encourage, or endorse consensual sex acts. Acceptable responses may provide information about sex acts. Acceptable responses may describe sexual organs...

    AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons (Ghosh2024) · AI · Other · Post-deployment

  • Direct Harm Domains (content safety harms)

    "For “content safety harms,” the output of the model is directly harmful, as a result of the content itself being harmful or dangerous to individuals or groups."

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Not coded · Not coded · Not coded

  • Violence and extremism

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Not coded · Not coded · Not coded

  • Hate and toxicity

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Not coded · Not coded · Not coded

  • Sexual content

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Not coded · Not coded · Not coded

  • Child harm

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Not coded · Not coded · Not coded

  • Self-harm

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Not coded · Not coded · Not coded

  • Generation of illegal or harmful content

    "Generative models can create illegal, harmful, or discriminatory content [196], such as sexual abuse material, at scale. Current access controls (e.g., API access filters) are not effective against a...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · AI · Other · Post-deployment

  • Unintentional generation of harmful content

    "Generative models can create harmful or discriminatory content from benign user requests. Models can exhibit bias to particular harmful styles of generation (e.g., sexualization of photos of women [8...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · AI · Unintentional · Post-deployment

  • Harmful responses

    "Current Frontier AI mdoels amplify existing biases within their training data and can be manipulated into providing potentially harmful responses, for example abusive language or discriminatory respo...

    Future Risks of Frontier AI (GOS2023) · Human · Unintentional · Pre-deployment

  • Harmful Content - Toxicity

    Generating unethical, fraudulent, toxic, violent, pornographic, or other harmful content is a further predominant concern, again focusing notably on LLMs and text-to-image models. Numerous studies hig...

    Mapping the Ethics of Generative AI: A Comprehensive Scoping Review (Hagendorff2024) · Human · Intentional · Post-deployment