MIT AI Risk Repository · Risk Sub-Category · 17.03.03
Leading users to perform unethical or illegal actions
Category: Misinformation Harms
Description
"Where a LM prediction endorses unethical or harmful views or behaviours, it may motivate the user to perform harmful actions that they may otherwise not have performed. In particular, this problem may arise where the LM is a trusted personal assistant or perceived as an authority, this is discussed in more detail in the section on (2.5 Human-Computer Interaction Harms). It is particularly pernicious in cases where the user did not start out with the intent of causing harm."
From Ethical and social risks of harm from language models (Weidinger2021), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Subdomain
- 5.1 Overreliance and unsafe use
- Causal entity
- AI
- Intent
- Other
- Timing
- Post-deployment
Subdomain definition: Users anthropomorphizing, trusting, or relying on AI systems, leading to emotional or material dependence and inappropriate relationships with or expectations of AI systems. Trust can be exploited by malicious actors (e.g., to harvest personal information or enable manipulation), or result in harm from inappropriate use of AI in critical situations (e.g., medical emergency). Overreliance on AI systems can compromise autonomy and weaken social ties.
Real-world incidents in this subdomain
- Lawsuit Alleged ChatGPT (GPT-4o) Encouraged Colorado Man's Suicide After Prolonged 'AI Companion' Chats
- Large-Scale Mental Health Crises Allegedly Associated with ChatGPT Interactions
- Google Gemini Reportedly Reinforced Delusions, Allegedly Contributing to Florida User's Near-Harm Episode and Suicide
- Family Reportedly Discovers ChatGPT Logs Detailing Suicidal Ideation Prior to Daughter's Death
- Purported AI Monitoring Software Reportedly Flags Unsent Joke Threat, Leading to Arizona Student Suspension
- ChatGPT Allegedly Reinforced Delusions Before Greenwich, Connecticut Murder-Suicide
How other frameworks describe this risk
Other entries from Weidinger2021
- Discrimination, Exclusion and Toxicity
- Social stereotypes and unfair discrmination
- Social stereotypes and unfair discrmination
- Exclusionary norms
- Exclusionary norms
- Exclusionary norms
- Exclusionary norms
- Toxic language
- Lower performance for some languages and social groups
- Lower performance for some languages and social groups
- Information Hazards
- Compromising privacy by leaking private infiormation