MIT AI Risk Repository · Risk Sub-Category · 24.09.05
Runaway processes
Category: Cooperation
Description
The 2010 flash crash is an example of a runaway process caused by interacting algorithms. Runaway processes are characterised by feedback loops that accelerate the process itself. Typically, these feedback loops arise from the interaction of multiple agents in a population... Within highly complex systems, the emergence of runaway processes may be hard to predict, because the conditions under which positive feedback loops occur may be non-obvious. The system of interacting AI assistants, their human principals, other humans and other algorithms will certainly be highly complex. Therefore, ther
From The Ethics of Advanced AI Assistants (Gabriel2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
Subdomain definition: AI systems acting in conflict with human goals or values, especially the goals of designers or users, or ethical standards. These misaligned behaviors may be introduced by humans during design and development, such as through reward hacking and goal misgeneralisation, or may result from AI using dangerous capabilities such as manipulation, deception, situational awareness to seek power, self-proliferate, or achieve other goals.
Real-world incidents in this subdomain
- Reinforcement Learning Reward Functions in Video Games
- Predictive Policing Program by Florida Sheriff’s Office Allegedly Violated Residents’ Rights and Targeted Children of Vulnerable Groups
- Image Classification of Battle Tanks
How other frameworks describe this risk
- Natural Language Underspecifies Goals
- Loss of control
- Loss of control
- Sudden loss of control
- AI leads to humans losing control of the future
- Risks from delegating decision-making power to misaligned AIs
- Risks from AIs developing goals and values that are different from humans
- Future AI systems might actively reduce human control
Other entries from Gabriel2024
- Capability failures
- Lack of capability for task
- Difficult to develop metrics for evaluating benefits or harms caused by AI assistants
- Safe exploration problem with widely deployed AI assistants
- Goal-related failures
- Misaligned consequentialist reasoning
- Specification gaming
- Goal misgeneralisation
- Deceptive alignment
- Malicious Uses
- Offensive Cyber Operations (General)
- AI-Powered Spear-Phishing at Scale