MIT AI Risk Repository · domain 7: AI system safety, failures, & limitations
7.1 AI pursuing its own goals in conflict with human goals or values
AI systems acting in conflict with human goals or values, especially the goals of designers or users, or ethical standards. These misaligned behaviors may be introduced by humans during design and development, such as through reward hacking and goal misgeneralisation, or may result from AI using dangerous capabilities such as manipulation, deception, situational awareness to seek power, self-proliferate, or achieve other goals.
- 100
- 12
- 3
- —
| Label | Value |
|---|---|
| AI | 73 |
| Other | 18 |
| Human | 8 |
| Label | Value |
|---|---|
| Intentional | 51 |
| Other | 34 |
| Unintentional | 14 |
| Label | Value |
|---|---|
| Other | 48 |
| Post-deployment | 33 |
| Pre-deployment | 18 |
| Label | Value |
|---|---|
| 2015 | 1 |
| 2016 | 1 |
| Label | Value |
|---|---|
| Risk Category | 32 |
| Risk Sub-Category | 68 |
Risk entries
Browse and export all- Untruthful Output
"AI systems such as LLMs can produce either unintentionally or deliberately inaccurateoutput. Such untruthful output may diverge from established resources or lack verifiability, commonly referredto a...
- Deceptive Alignment & Manipulation
"Manipulation & Deceptive Alignment is a class of behaviors thatexploit the incompetence of human evaluators or users (Hubinger et al., 2019a; Carranza et al., 2023) andeven manipulate the training pr...
- Collectively Harmful Behaviors
"AI systems have the potential to take actions that are seemingly benignin isolation but become problematic in multi-agent or societal contexts. Classical game theory offers simplistic models for unde...
- Agential
"While there are multiple types of intelligent agents, goal-based, utility-maximizing, and learning agents are the primary concern and the focus of this research"
- Harm caused by unaligned competent systems
"How do we ensure AI acts according to our values? Equivalently, how do we prevent poorly-understood AI systems from advancing goals we do not endorse? Whereas HP#2 concerns the prevention of harm cau...
- Specification gaming
"AI systems game specifications [305]. For example, in 2017 an OpenAI robot trained to grasp a ball via human feedback from a xed viewpoint learned that it was easier to pretend to grasp the ball by p...
- Emergent goals
"As well as optimizing a subtly wrong goal, systems can develop harmful instrumental goals in the service of a given goal—without these emergent goals being specied in any way [434, 218, 339, 17]. For...
- Alignment failures in existing ML systems
-
- Faulty reward functions in the wild
-
- Specification gaming
-
- Reward model overoptimization
-
- Instrumental convergence
-
- Goal misgeneralization
-
- Inner misalignment
-
- Language model misalignment
-
- Acquisition of goals to seek power and control
"cases where AI systems converge on optimal policies of seeking power over their environment;135"
- Existential disaster because of misaligned superintelligence or power-seeking AI
-
- Extreme “suffering risks” because of a misaligned system
-
- Existential disaster because of conflict between AI systems and multi-system interactions
-
- AGI removing itself from the control of human owners/managers
"The risks associated with containment, confinement, and control in the AGI development phase, and after an AGI has been developed, loss of control of an AGI."
- AGIs being given or developing unsafe goals
"The risks associated with AGI goal safety, including human attempts at making goals safe, as well as the AGI making its own goals safe during self-improvement."
- Existential risks
"The risks posed generally to humanity as a whole, including the dangers of unfriendly AGI, the suffering of the human race."
- Societal manipulation
"A sufficiently intelligent AI could possess the ability to subtly influence societal behaviors through a sophisticated understanding of human nature"
- Unpredictable outcomes
"Our culture, lifestyle, and even probability of survival may change drastically. Because the intentions programmed into an artificial agent cannot be guaranteed to lead to a positive outcome, Machine...
- Controllability
In the era of superintelligence, the agents will be difficult to control for humans... this problem is not solvable considering safety issues, and will be more severe by increasing the autonomy of AI-...