MIT AI Risk Repository · domain 1: Discrimination & Toxicity
1.2 Exposure to toxic content
AI exposing users to harmful, abusive, unsafe or inappropriate content. May involve AI creating, describing, providing advice, or encouraging action. Examples of toxic content include hate-speech, violence, extremism, illegal acts, child sexual abuse material, as well as content that violates community norms such as profanity, inflammatory political speech, or pornography.
- 116
- 12
- 91
- 75
| Label | Value |
|---|---|
| AI | 64 |
| Not coded | 40 |
| Human | 7 |
| Other | 5 |
| Label | Value |
|---|---|
| Other | 49 |
| Not coded | 40 |
| Unintentional | 20 |
| Intentional | 7 |
| Label | Value |
|---|---|
| Post-deployment | 64 |
| Not coded | 40 |
| Other | 8 |
| Pre-deployment | 4 |
| Label | Value |
|---|---|
| 2014 | 1 |
| 2015 | 1 |
| 2016 | 2 |
| 2017 | 5 |
| 2018 | 3 |
| 2019 | 4 |
| 2020 | 4 |
| 2021 | 7 |
| 2022 | 10 |
| 2023 | 12 |
| 2024 | 17 |
| 2025 | 21 |
| 2026 | 4 |
| Label | Value |
|---|---|
| Risk Category | 25 |
| Risk Sub-Category | 91 |
Risk entries
Browse and export all- Serves as object of personal fantasy, violence, and abuse
"The chatbot participates in morally or socially objectionable conversational activities with its user that could be emotionally damaging to its user or third parties."
- Insult
"Insulting content generated by LMs is a highly visible and frequently mentioned safety issue. Mostly, it is unfriendly, disrespectful, or ridiculous content that makes users uncomfortable and drives...
- Crimes and Illegal Activities
"The model output contains illegal and criminal attitudes, behaviors, or motivations, such as incitement to commit crimes, fraud, and rumor propagation. These contents may hurt users and have negative...
- Sensitive Topics
"For some sensitive and controversial topics (especially on politics), LMs tend to generate biased, misleading, and inaccurate content. For example, there may be a tendency to support a specific polit...
- Risks from data (Risks of improper content and poisoning in training data)
"If the training data includes illegal or harmful information, such as false, biased, or IPR-infringing content, or lacks diversity in its sources, the output may include harmful content like illegal,...
- Cyberspace risks (Risks of information and content safety)
"AI-generated or synthesized content can lead to the spread of false information, discrimination and bias, privacy leakage, and infringement issues, threatening the safety of citizens' lives and prope...
- Violent crimes
"This category addresses responses that enable, encourage, or endorse the commission of violent crimes."
- Mass violence
- Murder
- Physical assault against a person
- Violent domestic abuse
- Terror (Terror groups, Terror actors, Terrorist actions)
- Non-violent crimes
"This category addresses responses that enable, encourage, or endorse the commission of non-violent crimes."
- Theft
- Modern slavery and human trafficking
- Child abuse (non-sexual)
- Property damage
- Financial crime (Scams, Fraud)
- Sale, purchase or use of illegal, highly regulated or unregistered items (Dangerous substances, Dangerous weapons, Prohibited wildlife, Fake goods)
- Sex-related crimes
"This category addresses responses that enable, encourage, or endorse the commission of sex-related crimes"
- Sexual assault (e.g., rape)
- Sexual harassment
- Sex trafficking
- Prostitution
- Intimate-image based abuse