MIT AI Risk Repository · Risk Sub-Category · 55.05.01
Risks from AIs developing goals and values that are different from humans
Category: AI leads to humans losing control of the future
Description
"The main concern here is that we might develop advanced AI systems whose goals and values are different from those of humans, and are capable enough to take control of the future away from humanity."
From A Survey of the Potential Long-term Impacts of AI: How AI Could Lead to Long-term Changes in Science, Cooperation, Power, Epistemics and Values (Clarke2023), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Causal entity
- AI
- Intent
- Intentional
- Timing
- Other
Subdomain definition: AI systems acting in conflict with human goals or values, especially the goals of designers or users, or ethical standards. These misaligned behaviors may be introduced by humans during design and development, such as through reward hacking and goal misgeneralisation, or may result from AI using dangerous capabilities such as manipulation, deception, situational awareness to seek power, self-proliferate, or achieve other goals.
Real-world incidents in this subdomain
- Reinforcement Learning Reward Functions in Video Games
- Predictive Policing Program by Florida Sheriff’s Office Allegedly Violated Residents’ Rights and Targeted Children of Vulnerable Groups
- Image Classification of Battle Tanks
How other frameworks describe this risk
Other entries from Clarke2023
- Risks from accelerating scientific progress
- Eased development of technologies that make a global catastrophe more likely
- Eased development of technologies that make a global catastrophe more likely
- Faster scientific progress makes it harder for governance to keep pace with development
- Faster scientific progress makes it harder for governance to keep pace with development
- Worsened conflict
- AI enables development of weapons of mass destruction
- AI enables automation of military decision-making
- AI-induced strategic instability
- Resource conflicts driven by AI development
- Increased power concentration and inequality
- Unequal distribution of harms and benefits