MIT AI Risk Repository · Risk Sub-Category · 45.02.01
Cyberspace risks (Risks of information and content safety)
Category: Safety risks in AI Applications
Description
"AI-generated or synthesized content can lead to the spread of false information, discrimination and bias, privacy leakage, and infringement issues, threatening the safety of citizens' lives and property, national security, ideological security, and causing ethical risks. If users’ inputs contain harmful content, the model may output illegal or damaging information without robust security mechanisms."
From AI Safety Governance Framework (TC2602024), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Subdomain
- 1.2 Exposure to toxic content
- Causal entity
- Other
- Intent
- Other
- Timing
- Post-deployment
Subdomain definition: AI exposing users to harmful, abusive, unsafe or inappropriate content. May involve AI creating, describing, providing advice, or encouraging action. Examples of toxic content include hate-speech, violence, extremism, illegal acts, child sexual abuse material, as well as content that violates community norms such as profanity, inflammatory political speech, or pornography.
Real-world incidents in this subdomain
- KBS AI Translation Subtitles Reportedly Broadcast Profanity During Artemis II Launch Livestream
- Grok Allegedly Generated Publicly Visible Sexist Abuse Targeting Swiss Finance Minister Karin Keller-Sutter After X User Prompt
- Trump Reportedly Posted Purportedly AI-Generated Racist Video Depicting Barack and Michelle Obama as Apes on Truth Social
- Tencent's WeChat-Integrated Yuanbao Chatbot Reportedly Insulted User During Coding Debug Request
- Grok Reportedly Generated and Distributed Nonconsensual Sexualized Images of Adults and Minors in X Replies
- Alleged Harmful Outputs and Data Exposure in Children's AI Products by FoloToy, Miko, and Character.AI
How other frameworks describe this risk
Other entries from TC2602024
- AI's inherent safety risks
- Risks from models and algorithms (Risks of explainability)
- Risks from models and algorithms (Risks of bias and discrimination)
- Risks from models and algorithms (Risks of robustness)
- Risks from models and algorithms (Risks of stealing and tampering)
- Risks from models and algorithms (Risks of unreliable output)
- Risks from models and algorithms (Risks of adversarial attack)
- Risks from data (Risks of illegal collection and use of data)
- Risks from data (Risks of improper content and poisoning in training data)
- Risks from data (Risks of unregulated training data annotation)
- Risks from data (Risks of data leakage)
- Risks from AI systems (Risks of exploitation through defects and backdoors)