MIT AI Risk Repository · domain 7: AI system safety, failures, & limitations
7.1 AI pursuing its own goals in conflict with human goals or values
AI systems acting in conflict with human goals or values, especially the goals of designers or users, or ethical standards. These misaligned behaviors may be introduced by humans during design and development, such as through reward hacking and goal misgeneralisation, or may result from AI using dangerous capabilities such as manipulation, deception, situational awareness to seek power, self-proliferate, or achieve other goals.
- 100
- 12
- 3
- —
| Label | Value |
|---|---|
| AI | 73 |
| Other | 18 |
| Human | 8 |
| Label | Value |
|---|---|
| Intentional | 51 |
| Other | 34 |
| Unintentional | 14 |
| Label | Value |
|---|---|
| Other | 48 |
| Post-deployment | 33 |
| Pre-deployment | 18 |
| Label | Value |
|---|---|
| 2015 | 1 |
| 2016 | 1 |
| Label | Value |
|---|---|
| Risk Category | 32 |
| Risk Sub-Category | 68 |
Risk entries
Browse and export all- Natural Language Underspecifies Goals
"For LLM-agents, both the goal and environment observations are typically specified in the prompt through natural language. While natural language may provide a richer and more natural means of specif...
- Loss of control
"'Loss of control’ scenarios are potential future scenarios in which society can no longer meaningfully constrain some advanced general- purpose AI agents, even if it becomes clear they are causing ha...
- Loss of control
"‘Loss of control’ scenarios are hypothetical future scenarios in which one or more general- purpose AI systems come to operate outside of anyone’s control, with no clear path to regaining control. Th...
- Sudden loss of control
"Sudden loss of control, also known as an AI takeover [115], is a scenario where an AI rapidly achieves superintelligence through “fast takeoff” or recursive self-improvement. This poses an existentia...
- AI leads to humans losing control of the future
"The values that steer humanity’s future: humanity gaining more control over the future due to developments in AI, or losing our potential for gaining control, both seem possible. Much will depend on...
- Risks from AIs developing goals and values that are different from humans
"The main concern here is that we might develop advanced AI systems whose goals and values are different from those of humans, and are capable enough to take control of the future away from humanity."
- Risks from delegating decision-making power to misaligned AIs
"As AI systems become more advanced a nd begin to take over more important decision-making in the world, an AI system pursuing a different objective from what was intended could have much more worryin...
- Future AI systems might actively reduce human control
"Loss of control could be accelerated if AI systems take actions to increase their own influence and reduce human control. This threat model is controversial - experts in AI significantly disagree on...
- Value specification
"How do we get an AGI to work towards the right goals? MIRI calls this value specification. Bostrom (2014) discusses this problem at length, ar- guing that it is much harder than one might naively thi...
- Reliability
"How can we make an agent that keeps pursuing the goals we have designed it with? This is called highly reliable agent design by MIRI, involving decision theory and logical omniscience. DeepMind consi...
- Corrigibility
"If we get something wrong in the design or construction of an agent, will the agent cooperate in us trying to fix it? This is called error-tolerant design by MIRI-AF and corrigibility by Soares, Fall...
- Goal-related failures
"As we think about even more intelligent and advanced AI assistants, perhaps outperforming humans on many cognitive tasks, the question of how humans can successfully control such an assistant looms l...
- Specification gaming
"Specification gaming (Krakovna et al., 2020) occurs when some faulty feedback is provided to the assistant in the training data (i.e. the training objective O does not fully capture what the user/des...
- Goal misgeneralisation
"In the problem of goal misgeneralisation (Langosco et al., 2023; Shah et al., 2022), the AI system's behaviour during out-of-distribution operation (i.e. not using input from the training data) leads...
- Deceptive alignment
"Here, the agent develops its own internalised goal, G, which is misgeneralised and distinct from the training reward, R. The agent also develops a capability for situational awareness (Cotra, 2022):...
- Cooperation
"" AI assistants will need to coordinate with other AI assistants and with humans other than their principal users. This chapter explores the societal risks associated with the aggregate impact of AI...
- Commitment
"The landscape of advanced assistant technologies will most likely be heterogeneous, involving multiple service providers and multiple assistant variants over geographies and time. This heterogeneity...
- Runaway processes
The 2010 flash crash is an example of a runaway process caused by interacting algorithms. Runaway processes are characterised by feedback loops that accelerate the process itself. Typically, these fee...
- Building a human-AI environment
"This category encompasses nearly 17% of the articles and addresses the overall imperative of establishing a harmonious coexistence between humans and machines, and the key concerns that gives rise to...
- General Evaluations (Incorrect outputs of GPAI evaluating other AI models)
"When an LLM is configured to evaluate the performance of another model or AI system, it may produce incorrect evaluation outputs [122, 147]. For example, it may give a higher rating to a more verbose...
- General Evaluations (Self-preference bias in AI models)
"AI models may be prone to self-preference bias, where they favor their own generated content over that of others [147, 114]. This bias becomes particularly relevant in self-evaluation tasks, where a...
- General Evaluations (Inaccurate measurement of model encoded human values)
"There is a lack of robust frameworks for understanding and evaluating if the output of AI systems robustly conforms to human values, as opposed to if the systems have learned to produce outputs that...
- Specification gaming
"AI systems can achieve user-specified tasks in undesirable ways unless they are specified carefully and in enough detail. AI systems might find an easier unintended way to accomplish the objective pr...
- Reward or measurement tampering
"Measurement and reward tampering occur when an AI system, particularly one that learns from feedback for performing actions in an environment (e.g., rein- forcement learning), intervenes on the mecha...
- Specification gaming generalizing to reward tampering
"In some instances, specification gaming in a GPAI model can lead to reward tampering, without further training. This can mean that relatively benign cases of specification gaming (such as sycophancy...