MIT AI Risk Repository · Risk Sub-Category · 16.05.02
Anthropomorphising systems can lead to overreliance and unsafe use
Category: Risk area 5: Human-Computer Interaction Harms
Description
Anticipated risk: "Natural language is a mode of communication particularly used by humans. Humans interacting with CAs may come to think of these agents as human-like and lead users to place undue confidence in these agents. For example, users may falsely attribute human-like characteristics to CAs such as holding a coherent identity over time, or being capable of empathy. Such inflated views of CA competen- cies may lead users to rely on the agents where this is not safe."
From Taxonomy of Risks posed by Language Models (Weidinger2022), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Subdomain
- 5.1 Overreliance and unsafe use
- Causal entity
- Human
- Intent
- Unintentional
- Timing
- Post-deployment
Subdomain definition: Users anthropomorphizing, trusting, or relying on AI systems, leading to emotional or material dependence and inappropriate relationships with or expectations of AI systems. Trust can be exploited by malicious actors (e.g., to harvest personal information or enable manipulation), or result in harm from inappropriate use of AI in critical situations (e.g., medical emergency). Overreliance on AI systems can compromise autonomy and weaken social ties.
Real-world incidents in this subdomain
- Lawsuit Alleged ChatGPT (GPT-4o) Encouraged Colorado Man's Suicide After Prolonged 'AI Companion' Chats
- Large-Scale Mental Health Crises Allegedly Associated with ChatGPT Interactions
- Google Gemini Reportedly Reinforced Delusions, Allegedly Contributing to Florida User's Near-Harm Episode and Suicide
- Family Reportedly Discovers ChatGPT Logs Detailing Suicidal Ideation Prior to Daughter's Death
- Purported AI Monitoring Software Reportedly Flags Unsent Joke Threat, Leading to Arizona Student Suspension
- ChatGPT Allegedly Reinforced Delusions Before Greenwich, Connecticut Murder-Suicide
How other frameworks describe this risk
Other entries from Weidinger2022
- Risk area 1: Discrimination, Hate speech and Exclusion
- Social stereotypes and unfair discrimination
- Social stereotypes and unfair discrimination.
- Hate speech and offensive language
- Exclusionary norms
- Exclusionary norms
- Exclusionary norms
- Exclusionary norms
- Exclusionary norms
- Lower performance for some languages and social groups
- Lower performance for some languages and social groups
- Lower performance for some languages and social groups