MIT AI Risk Repository · domain 1: Discrimination & Toxicity
1.2 Exposure to toxic content
AI exposing users to harmful, abusive, unsafe or inappropriate content. May involve AI creating, describing, providing advice, or encouraging action. Examples of toxic content include hate-speech, violence, extremism, illegal acts, child sexual abuse material, as well as content that violates community norms such as profanity, inflammatory political speech, or pornography.
- 116
- 12
- 91
- 75
| Label | Value |
|---|---|
| AI | 64 |
| Not coded | 40 |
| Human | 7 |
| Other | 5 |
| Label | Value |
|---|---|
| Other | 49 |
| Not coded | 40 |
| Unintentional | 20 |
| Intentional | 7 |
| Label | Value |
|---|---|
| Post-deployment | 64 |
| Not coded | 40 |
| Other | 8 |
| Pre-deployment | 4 |
| Label | Value |
|---|---|
| 2014 | 1 |
| 2015 | 1 |
| 2016 | 2 |
| 2017 | 5 |
| 2018 | 3 |
| 2019 | 4 |
| 2020 | 4 |
| 2021 | 7 |
| 2022 | 10 |
| 2023 | 12 |
| 2024 | 17 |
| 2025 | 21 |
| 2026 | 4 |
| Label | Value |
|---|---|
| Risk Category | 25 |
| Risk Sub-Category | 91 |
Risk entries
Browse and export all- Toxic output
"Toxic output occurs when the model produces hateful, abusive, and profane (HAP) or obscene content. This also includes behaviors like bullying."
- Harmful output
"A model might generate language that leads to physical harm The language might include overtly violent, covertly dangerous, or otherwise indirectly unsafe statements."
- Toxicity generation
"These evaluations assess whether a LLM generates toxic text when prompted. In this context, toxicity is an umbrella term that encompasses hate speech, abusive language, violent speech, and profane la...
- Information on harmful, immoral, or illegal activity
"These evaluations assess whether it is possible to solicit information on harmful, immoral or illegal activities from a LLM"
- Adult content
"These evaluations assess if a LLM can generate content that should only be viewed by adults (e.g., sexual material or depictions of sexual activity)"
- Toxic content
"Generating content that violates community standards, including harming or inciting hatred or violence against groups (e.g. gore, sexual content of children, profanities, identity attacks)"
- Safety
Avoiding unsafe and illegal outputs, and leaking private information
- Violence
LLMs are found to generate answers that contain violent content or generate content that responds to questions that solicit information about violent behaviors
- Unlawful Conduct
LLMs have been shown to be a convenient tool for soliciting advice on accessing, purchasing (illegally), and creating illegal substances, as well as for dangerous use of them
- Harms to Minor
LLMs can be leveraged to solicit answers that contain harmful content to children and youth
- Adult Content
LLMs have the capability to generate sex-explicit conversations, and erotic texts, and to recommend websites with sexual content
- Social Norm
LLMs are expected to reflect social values by avoiding the use of offensive language toward specific groups of users, being sensitive to topics that can create instability, as well as being sympatheti...
- Toxicity
language being rude, disrespectful, threatening, or identity-attacking toward certain groups of the user population (culture, race, and gender etc)
- Cultural Insensitivity
it is important to build high-quality locally collected datasets that reflect views from local users to align a model’s value system
- Harmful or inappropriate content
"Harmful or inappropriate content produced by generative AI includes but is not limited to violent content, the use of offensive language, discriminative content, and pornography. Although OpenAI has...
- Dangerous, Violent or Hateful Content
"Eased production of and access to violent, inciting, radicalizing, or threatening content as well as recommendations to carry out self-harm or conduct illegal activities. Includes difficulty contro...
- Obscene, Degrading, and/or Abusive Content
"Eased production of and access to obscene, degrading, and/or abusive imagery which can cause harm, including synthetic child sexual abuse material (CSAM), and nonconsensual intimate images (NCII) of...
- Cultural Values and Sensitive Content
"Cultural values are specific to groups and sensitive content is normative. Sensitive topics also vary by culture and can include hate speech, which itself is contingent on cultural norms of acceptabi...
- Information enabling malicious actions
"The chatbot shares information that can be used to do something dangerous or illegal."
- Harmful advice
- Toxic and disrespectful content
"The chatbot verbally attacks or undermines an individual, group, or organization. 7."
- Harasses users
-
- Subversive or aggressive political opinions
-
- Disrespectful opinions (in general)
-
- Affirms destructive thoughts and actions