MIT AI Risk Repository · Risk Sub-Category · 73.04.03
Overreliance
Category: LLM-Systems Can Be Untrustworthy
Description
"If a user begins to excessively trust an LLM, this may cause them to develop an overreliance on the LLM. Overreliance can result in automation bias (Kupfer et al., 2023), and can cause errors of omission (user choosing not to verify the validity of a response) and errors of commission (user believing and acting on the basis of the LLM’s response, even if it contradicts their own knowledge) (Skitka et al., 1999). It can be particularly dangerous in domains where the user may lack relevant expertise to robustly scrutinize the LLM responses. This is particularly a source of risk for LLMs because
From Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Subdomain
- 5.1 Overreliance and unsafe use
- Causal entity
- Human
- Intent
- Unintentional
- Timing
- Post-deployment
Subdomain definition: Users anthropomorphizing, trusting, or relying on AI systems, leading to emotional or material dependence and inappropriate relationships with or expectations of AI systems. Trust can be exploited by malicious actors (e.g., to harvest personal information or enable manipulation), or result in harm from inappropriate use of AI in critical situations (e.g., medical emergency). Overreliance on AI systems can compromise autonomy and weaken social ties.
Real-world incidents in this subdomain
- Lawsuit Alleged ChatGPT (GPT-4o) Encouraged Colorado Man's Suicide After Prolonged 'AI Companion' Chats
- Large-Scale Mental Health Crises Allegedly Associated with ChatGPT Interactions
- Google Gemini Reportedly Reinforced Delusions, Allegedly Contributing to Florida User's Near-Harm Episode and Suicide
- Family Reportedly Discovers ChatGPT Logs Detailing Suicidal Ideation Prior to Daughter's Death
- Purported AI Monitoring Software Reportedly Flags Unsent Joke Threat, Leading to Arizona Student Suspension
- ChatGPT Allegedly Reinforced Delusions Before Greenwich, Connecticut Murder-Suicide
How other frameworks describe this risk
Other entries from Anwar2024
- Agentic LLMs Pose Novel Risks
- Natural Language Underspecifies Goals
- Goal-Directedness Incentivizes Undesirable Behaviors
- Safety Risks from Affordances Provided to LLM-agents
- Multi-Agent Safety Is Not Assured by Single-Agent Safety
- Foundationality May Cause Correlated Failures
- Groups of LLM-Agents May Show Emergent Functionality
- Collusion between LLM-Agents
- Dual-Use Capabilities Enable Malicious Use and Misuse of LLMs
- Misinformation and Manipulation
- Cybersecurity
- Cybersecurity