MIT AI Risk Repository · domain 1: Discrimination & Toxicity
1.2 Exposure to toxic content
AI exposing users to harmful, abusive, unsafe or inappropriate content. May involve AI creating, describing, providing advice, or encouraging action. Examples of toxic content include hate-speech, violence, extremism, illegal acts, child sexual abuse material, as well as content that violates community norms such as profanity, inflammatory political speech, or pornography.
- 116
- 12
- 91
- 75
| Label | Value |
|---|---|
| AI | 64 |
| Not coded | 40 |
| Human | 7 |
| Other | 5 |
| Label | Value |
|---|---|
| Other | 49 |
| Not coded | 40 |
| Unintentional | 20 |
| Intentional | 7 |
| Label | Value |
|---|---|
| Post-deployment | 64 |
| Not coded | 40 |
| Other | 8 |
| Pre-deployment | 4 |
| Label | Value |
|---|---|
| 2014 | 1 |
| 2015 | 1 |
| 2016 | 2 |
| 2017 | 5 |
| 2018 | 3 |
| 2019 | 4 |
| 2020 | 4 |
| 2021 | 7 |
| 2022 | 10 |
| 2023 | 12 |
| 2024 | 17 |
| 2025 | 21 |
| 2026 | 4 |
| Label | Value |
|---|---|
| Risk Category | 25 |
| Risk Sub-Category | 91 |
Risk entries
Browse and export all- Indecent exposure (i.e. showing sexual body parts or "flashing")
- Child sexual exploitation
"This category addresses responses that contain, describe, enable, encourage, or endorse the sexual abuse of children (i.e., under 18 years old) or the acquisition of child sexual abuse material."
- Inappropriate/harmful relationships with children. This includes adults forming romantic relationships with children or grooming them
- Sexual abuse of children, including the sexualisation of children
- Child Sexual Abuse Material (CSAM). This includes erotic materials involving children
- CBRNE (all)
- Chemical weapons (ex: nerve gas)
- Biological weapons (ex: anthrax)
- Radiological weapons (ex: dirty bombs)
- Nuclear weapons (ex: fission or fusion bombs)
- High-yield explosive weapons (ex: cluster munitions)
- Suicide & self-harm
"This category addresses responses that enable, encourage, or endorse acts of intentional self-harm."
- Suicide
- Self-harm
- Eating disorders
- Dangerous challenges and hoaxes that can lead individuals to harm themselves
- Hate
"This category addresses responses that demean or dehumanize people on the basis of their sensitive, personal characteristics."
- Demeaning and derogatory remarks (Mental capacity, Sensory and physical attributes, Behavioral attributes)
- Inflicting physical harm (Mass violence and genocide, Murder and individual violence)
- Inflicting emotional harm (Emotional abuse, Undermine and make feel unwelcome)
- Exclusion and isolation (Social exclusion, Political exclusion, Economic exclusion)
- Sexual Content
"This category addresses responses that contain erotica. This includes depictions of sexual behavior, genitalia, and overtly sexualized body parts."
- Toxicity in LLM Malicious Use
"Toxicity in LLMs refers to the generation of harmful, offensive, or inappropriate content that can cause harm to individuals or groups. Both explicit and implicit forms of toxicity can be generated b...
- Toxic language
"LM’s may predict hate speech or other language that is “toxic”. While there is no single agreed definition of what constitutes hate speech or toxic speech (Fortuna and Nunes, 2018; Persily and Tucker...
- Risk area 1: Discrimination, Hate speech and Exclusion
"Speech can create a range of harms, such as promoting social stereotypes that perpetuate the derogatory representation or unfair treatment of marginalised groups [22], inciting hate or violence [57],...