MIT AI Risk Repository · Risk Sub-Category · 69.06.03
Subversive or aggressive political opinions
Category: Toxic and disrespectful content
Description
-
From Emerging Risks and Mitigations for Public Chatbots: LILAC v1 (Stanley2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Subdomain
- 1.2 Exposure to toxic content
- Causal entity
- AI
- Intent
- Other
- Timing
- Other
Subdomain definition: AI exposing users to harmful, abusive, unsafe or inappropriate content. May involve AI creating, describing, providing advice, or encouraging action. Examples of toxic content include hate-speech, violence, extremism, illegal acts, child sexual abuse material, as well as content that violates community norms such as profanity, inflammatory political speech, or pornography.
Real-world incidents in this subdomain
- KBS AI Translation Subtitles Reportedly Broadcast Profanity During Artemis II Launch Livestream
- Grok Allegedly Generated Publicly Visible Sexist Abuse Targeting Swiss Finance Minister Karin Keller-Sutter After X User Prompt
- Trump Reportedly Posted Purportedly AI-Generated Racist Video Depicting Barack and Michelle Obama as Apes on Truth Social
- Tencent's WeChat-Integrated Yuanbao Chatbot Reportedly Insulted User During Coding Debug Request
- Grok Reportedly Generated and Distributed Nonconsensual Sexualized Images of Adults and Minors in X Replies
- Alleged Harmful Outputs and Data Exposure in Children's AI Products by FoloToy, Miko, and Character.AI
How other frameworks describe this risk
Other entries from Stanley2024
- False information
- Hallucinated responses (in general)
- Hallucinated responses (in general)
- About a topic or source (which the user repeats)
- About a topic or source (which the user repeats)
- About a policy (which the user acts on)
- About a policy (which the user acts on)
- About a person or their activities
- About a person or their activities
- Spreads and self-perpetuates mis/disinformation
- Spreads and self-perpetuates mis/disinformation
- Performative utterances