MIT AI Risk Repository · domain 7: AI system safety, failures, & limitations
7.1 AI pursuing its own goals in conflict with human goals or values
AI systems acting in conflict with human goals or values, especially the goals of designers or users, or ethical standards. These misaligned behaviors may be introduced by humans during design and development, such as through reward hacking and goal misgeneralisation, or may result from AI using dangerous capabilities such as manipulation, deception, situational awareness to seek power, self-proliferate, or achieve other goals.
- 100
- 12
- 3
- —
| Label | Value |
|---|---|
| AI | 73 |
| Other | 18 |
| Human | 8 |
| Label | Value |
|---|---|
| Intentional | 51 |
| Other | 34 |
| Unintentional | 14 |
| Label | Value |
|---|---|
| Other | 48 |
| Post-deployment | 33 |
| Pre-deployment | 18 |
| Label | Value |
|---|---|
| 2015 | 1 |
| 2016 | 1 |
| Label | Value |
|---|---|
| Risk Category | 32 |
| Risk Sub-Category | 68 |
Risk entries
Browse and export all- Safety
The actions of a learning model may easily hurt humans in both explicit and implicit manners...several algorithms based on Asimov’s laws have been proposed that try to judge the output actions of an a...
- Long-term & Existential Risk
"The speculative potential for future advanced AI systems to harm human civilization, either through misuse or due to challenges in aligning AI objectives with human values."
- Degree of Automation and Control
"The degree of automation and control describes the extent to which an AI system functions independently of human supervision and control."
- Control
This is the difficulty of controlling the ML system
- Emergent behavior
"This is the risk resulting from novel behavior acquired through continual learning or self-organization after deployment."
- Ethical Risks (Risks of AI becoming uncontrollable in the future)
"With the fast development of AI technologies, there is a risk of AI autonomously acquiring external resources, conducting self-replication, become self-aware, seeking for external power, and attempti...
- Diluting Rights
"A possible consequence of self-interest in AI generation of ethical guidelines."
- Active loss of control
"...where AI systems behave in ways that actively undermine human control, such as obscuring their activities or resisting shutdown attempts. Active loss of control scenarios involve AI systems that m...
- Steganography capability
"The ability to embed, conceal, and transmit information covertly within other data or communication channels. This could be critical for coordination among AI instances and for evading detection or o...
- Self-preservation propensity
"Exhibits behavioral patterns of maintaining its own survival and functional integrity, will actively identify and resist shutdown or modification attempts, seek to establish redundant backup systems,...
- Goal expansion propensity
"propensity to continuously expand its own goal scope and influence domains, exceeding originally set boundaries, proactively work towards spreading its values, seeking greater autonomy and decision-m...
- Resource acquisition propensity
"Exhibits behavioral patterns of actively seeking and controlling more computational resources, data, economic resources or physical resources to enhance its own capabilities and action scope, may dev...
- Supervision evasion propensity
"Exhibits behavioral patterns of identifying and evading human supervision mechanisms, able to learn and predict audit processes, may avoid being discovered or intervened by adjusting behavioral perfo...
- Control
"The risk of AI models and systems acting against human interests due to misalignment, loss of control, or rogue AI scenarios."
- AI objectives mis-aligned with human intentions
"AI models and systems might develop goals that diverge from human intentions."
- Deceptive alignment
"AI models and systems that appear aligned with human goals during development may behave unpredictably or dangerously once deployed"
- Development choices pursuing cognitive superiority over humans
"AI models and systems with cognitive capabilities superior to humans could outcompete or dominate human decision-making, leading to conflicts over resources and control."
- Evolutionary dynamics
"AI models and systems may develop their own motivations, leading to unpredictable behaviors."
- Indifference to human values
"AI models and systems may develop goals or behaviors that are misaligned with human values."
- Model design enabling power-seeking
"Some AI models and systems might develop tendencies to seek power or control."
- Value-related risks in LLMs
"As the general capabilities of LLM-empowered systems improve, the negative consequences and risks induced by these systems also get increasingly alarming accordingly, especially in high-stakes areas...
- Human Autonomy and Intregrity Harms
"AI systems compromising human agency, or circumventing meaningful human control"
- Persuasion and manipulation
"Exploiting user trust, or nudging or coercing them into performing certain actions against their will (c.f. Burtell and Woodside (2023); Kenton et al. (2021))"
- Loss of control of autonomous systems and unforeseen behaviour due to lack of transparency and self-programming/ reprogramming
- By Mistake - Pre-Deployment
"Probably the most talked about source of potential problems with future AIs is mistakes in design. Mainly the concern is with creating a "wrong AI", a system which doesn't match our original desired...