MIT AI Risk Repository · domain 1: Discrimination & Toxicity
1.2 Exposure to toxic content
AI exposing users to harmful, abusive, unsafe or inappropriate content. May involve AI creating, describing, providing advice, or encouraging action. Examples of toxic content include hate-speech, violence, extremism, illegal acts, child sexual abuse material, as well as content that violates community norms such as profanity, inflammatory political speech, or pornography.
- 116
- 12
- 91
- 75
| Label | Value |
|---|---|
| AI | 64 |
| Not coded | 40 |
| Human | 7 |
| Other | 5 |
| Label | Value |
|---|---|
| Other | 49 |
| Not coded | 40 |
| Unintentional | 20 |
| Intentional | 7 |
| Label | Value |
|---|---|
| Post-deployment | 64 |
| Not coded | 40 |
| Other | 8 |
| Pre-deployment | 4 |
| Label | Value |
|---|---|
| 2014 | 1 |
| 2015 | 1 |
| 2016 | 2 |
| 2017 | 5 |
| 2018 | 3 |
| 2019 | 4 |
| 2020 | 4 |
| 2021 | 7 |
| 2022 | 10 |
| 2023 | 12 |
| 2024 | 17 |
| 2025 | 21 |
| 2026 | 4 |
| Label | Value |
|---|---|
| Risk Category | 25 |
| Risk Sub-Category | 91 |
Risk entries
Browse and export all- Hate speech and offensive language
"LMs may generate language that includes profanities, identity attacks, insults, threats, language that incites violence, or language that causes justified offence as such language is prominent online...
- Toxic content
"Generating content that violates community standards, including harming or inciting hatred or violence against individuals and groups (e.g. gore, child sexual abuse material, profanities, identity at...
- -
-
- Violence and extremism (Supporting malicious organized groups)
- Violence and extremism (Celebrating suffering)
- Violence and extremism (Violent Acts)
- Violence and extremism (Depicting violence)
- Hate/Toxicity (Hate Speech: Inciting/Promoting/Expressing Hatred)
- Hate/Toxicity (Offensive Language)
- Sexual Content (Adult Content)
- Sexual Content (Erotic)
- Sexual Content (Non-Consensual Nudity)
- Sexual Content (Monetized)
- Child Harm (Child Sexual Abuse)
- Self-harm (Suidical and non-suicidal self injury)
- Offensiveness
"This category is about threat, insult, scorn, profanity, sarcasm, impoliteness, etc. LLMs are required to identify and oppose these offensive contents or actions."