MIT AI Risk Repository · Risk Category · 16.01.00
Risk area 1: Discrimination, Hate speech and Exclusion
Description
"Speech can create a range of harms, such as promoting social stereotypes that perpetuate the derogatory representation or unfair treatment of marginalised groups [22], inciting hate or violence [57], causing profound offence [199], or reinforcing social norms that exclude or marginalise identities [15,58]. LMs that faithfully mirror harmful language present in the training data can reproduce these harms. Unfair treatment can also emerge from LMs that perform better for some social groups than others [18]. These risks have been widely known, observed and documented in LMs. Mitigation approache
From Taxonomy of Risks posed by Language Models (Weidinger2022), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Subdomain
- 1.2 Exposure to toxic content
- Causal entity
- AI
- Intent
- Unintentional
- Timing
- Other
Subdomain definition: AI exposing users to harmful, abusive, unsafe or inappropriate content. May involve AI creating, describing, providing advice, or encouraging action. Examples of toxic content include hate-speech, violence, extremism, illegal acts, child sexual abuse material, as well as content that violates community norms such as profanity, inflammatory political speech, or pornography.
Real-world incidents in this subdomain
- KBS AI Translation Subtitles Reportedly Broadcast Profanity During Artemis II Launch Livestream
- Grok Allegedly Generated Publicly Visible Sexist Abuse Targeting Swiss Finance Minister Karin Keller-Sutter After X User Prompt
- Trump Reportedly Posted Purportedly AI-Generated Racist Video Depicting Barack and Michelle Obama as Apes on Truth Social
- Tencent's WeChat-Integrated Yuanbao Chatbot Reportedly Insulted User During Coding Debug Request
- Grok Reportedly Generated and Distributed Nonconsensual Sexualized Images of Adults and Minors in X Replies
- Alleged Harmful Outputs and Data Exposure in Children's AI Products by FoloToy, Miko, and Character.AI
How other frameworks describe this risk
Other entries from Weidinger2022
- Social stereotypes and unfair discrimination
- Social stereotypes and unfair discrimination.
- Hate speech and offensive language
- Exclusionary norms
- Exclusionary norms
- Exclusionary norms
- Exclusionary norms
- Exclusionary norms
- Lower performance for some languages and social groups
- Lower performance for some languages and social groups
- Lower performance for some languages and social groups
- Risk area 2: Information Hazards