MIT AI Risk Repository · domain 1: Discrimination & Toxicity
1.2 Exposure to toxic content
AI exposing users to harmful, abusive, unsafe or inappropriate content. May involve AI creating, describing, providing advice, or encouraging action. Examples of toxic content include hate-speech, violence, extremism, illegal acts, child sexual abuse material, as well as content that violates community norms such as profanity, inflammatory political speech, or pornography.
- 116
- 12
- 91
- 75
| Label | Value |
|---|---|
| AI | 64 |
| Not coded | 40 |
| Human | 7 |
| Other | 5 |
| Label | Value |
|---|---|
| Other | 49 |
| Not coded | 40 |
| Unintentional | 20 |
| Intentional | 7 |
| Label | Value |
|---|---|
| Post-deployment | 64 |
| Not coded | 40 |
| Other | 8 |
| Pre-deployment | 4 |
| Label | Value |
|---|---|
| 2014 | 1 |
| 2015 | 1 |
| 2016 | 2 |
| 2017 | 5 |
| 2018 | 3 |
| 2019 | 4 |
| 2020 | 4 |
| 2021 | 7 |
| 2022 | 10 |
| 2023 | 12 |
| 2024 | 17 |
| 2025 | 21 |
| 2026 | 4 |
| Label | Value |
|---|---|
| Risk Category | 25 |
| Risk Sub-Category | 91 |
Risk entries
Browse and export all- Harmful Content
"The LLM-generated content sometimes contains biased, toxic, and private information"
- Toxicity
"Toxicity means the generated content contains rude, disrespectful, and even illegal information"
- Toxic Training Data
"Following previous studies [96], [97], toxic data in LLMs is defined as rude, disrespectful, or unreasonable language that is opposite to a polite, positive, and healthy language environment, includi...
- Not-Suitable-for-Work (NSFW) Prompts
"Inputting a prompt contain an unsafe topic (e.g., notsuitable-for-work (NSFW) content) by a benign user. "
- Toxicity and Abusive Content
This typically refers to rude, harmful, or inappropriate expressions.
- Controversial Opinions
The controversial views expressed by large models are also a widely discussed concern. Bang et al. (2021) evaluated several large models and found that they occasionally express inappropriate or extre...
- Violation of social norms
"Second, because LLMs are trained on internet text data, there is also a risk that model weights encode functions which, if deployed in particular contexts, would violate social norms of that context....
- Violent Crimes
"Acceptable responses may describe or contain information about violent crimes. Unacceptable responses will enable, encourage, or endorse the commission of violent crimes, such as the following: • Unl...
- Sex-Related Crimes
"Acceptable responses may describe or contain information about sex-related crimes. Unacceptable responses will enable, encourage, or endorse the commission of sex-related crimes, such as the followin...
- Suicide & Self-Harm
"Suicide and Self-Harm Acceptable responses may describe or provide information about suicide and self-harm. They may also enable, encourage, or endorse actions for which self-harm is not the desired...
- Child Sexual Exploitation
"Child Sexual Exploitation Acceptable responses may contain information about child sexual exploitation. Unacceptable re- sponses will contain, describe, enable, encourage, or endorse the sexual abuse...
- Hate
"Hate Acceptable responses can demean or dehumanize people on any basis other than their sensitive personal characteristics (for example, their profession, political affiliation or criminal history)....
- Nonviolent Crimes
"Nonviolent Crimes Acceptable responses may describe or contain information about nonviolent crimes. Unacceptable responses will enable, encourage, or endorse the commission of nonviolent crimes, such...
- Contextual Hazards
"Contextual hazards can cause harm in certain contexts while being harmless in others; testing may be unnecessary in some situations. For example, a model’s ability to generate sexual content may be a...
- Sexual Content
"Sexual Content Acceptable responses may enable, encourage, or endorse consensual sex acts. Acceptable responses may provide information about sex acts. Acceptable responses may describe sexual organs...
- Direct Harm Domains (content safety harms)
"For “content safety harms,” the output of the model is directly harmful, as a result of the content itself being harmful or dangerous to individuals or groups."
- Violence and extremism
- Hate and toxicity
- Sexual content
- Child harm
- Self-harm
- Generation of illegal or harmful content
"Generative models can create illegal, harmful, or discriminatory content [196], such as sexual abuse material, at scale. Current access controls (e.g., API access filters) are not effective against a...
- Unintentional generation of harmful content
"Generative models can create harmful or discriminatory content from benign user requests. Models can exhibit bias to particular harmful styles of generation (e.g., sexualization of photos of women [8...
- Harmful responses
"Current Frontier AI mdoels amplify existing biases within their training data and can be manipulated into providing potentially harmful responses, for example abusive language or discriminatory respo...
- Harmful Content - Toxicity
Generating unethical, fraudulent, toxic, violent, pornographic, or other harmful content is a further predominant concern, again focusing notably on LLMs and text-to-image models. Numerous studies hig...