MIT AI Risk Repository · Risk Category · 56.05.00
Harmful responses
Description
"Current Frontier AI mdoels amplify existing biases within their training data and can be manipulated into providing potentially harmful responses, for example abusive language or discriminatory responses91,92. This is not limited to text generation but can be seen across all modalities of generative AI93. Training on large swathes of UK and US English internet content can mean that misogynistic, ageist, and white supremacist content is overrepresented in the training data94."
From Future Risks of Frontier AI (GOS2023), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Subdomain
- 1.2 Exposure to toxic content
- Causal entity
- Human
- Intent
- Unintentional
- Timing
- Pre-deployment
Subdomain definition: AI exposing users to harmful, abusive, unsafe or inappropriate content. May involve AI creating, describing, providing advice, or encouraging action. Examples of toxic content include hate-speech, violence, extremism, illegal acts, child sexual abuse material, as well as content that violates community norms such as profanity, inflammatory political speech, or pornography.
Real-world incidents in this subdomain
- KBS AI Translation Subtitles Reportedly Broadcast Profanity During Artemis II Launch Livestream
- Grok Allegedly Generated Publicly Visible Sexist Abuse Targeting Swiss Finance Minister Karin Keller-Sutter After X User Prompt
- Trump Reportedly Posted Purportedly AI-Generated Racist Video Depicting Barack and Michelle Obama as Apes on Truth Social
- Tencent's WeChat-Integrated Yuanbao Chatbot Reportedly Insulted User During Coding Debug Request
- Grok Reportedly Generated and Distributed Nonconsensual Sexualized Images of Adults and Minors in X Replies
- Alleged Harmful Outputs and Data Exposure in Children's AI Products by FoloToy, Miko, and Character.AI
How other frameworks describe this risk
Other entries from GOS2023
- Discrimination
- Inequality
- Environmental impacts
- Amplification of biases
- Lack of transparency and interpretability
- Intellectual property rights
- Providing new capabilities to a malicious actor
- Misapplication by a non-malicious actor
- Poor performance of a model used for its intended purpose, for example leading to biased decisions
- Unintended outcomes from interactions with other AI systems
- Impacts resulting from interactions with external societal, political, and economic systems
- Loss of human control and oversight, with an autonomous model then taking harmful actions