MIT AI Risk Repository · domain 7: AI system safety, failures, & limitations

7.1 AI pursuing its own goals in conflict with human goals or values

AI systems acting in conflict with human goals or values, especially the goals of designers or users, or ethical standards. These misaligned behaviors may be introduced by humans during design and development, such as through reward hacking and goal misgeneralisation, or may result from AI using dangerous capabilities such as manipulation, deception, situational awareness to seek power, self-proliferate, or achieve other goals.

Risk entries
100
Frameworks citing it
12
Recorded incidents
3
Incidents since 2020
Causal entity (risk entries)
Causal entity (risk entries) 73 0 AI: 73 AI 73 Other: 18 Other 18 Human: 8 Human 8
Causal entity (risk entries)
LabelValue
AI73
Other18
Human8
Intent (risk entries)
Intent (risk entries) 51 0 Intentional: 51 Intentional 51 Other: 34 Other 34 Unintentional: 14 Unintentional 14
Intent (risk entries)
LabelValue
Intentional51
Other34
Unintentional14
Timing (risk entries)
Timing (risk entries) 48 0 Other: 48 Other 48 Post-deployment: 33 Post-deployment 33 Pre-deployment: 18 Pre-deployment 18
Timing (risk entries)
LabelValue
Other48
Post-deployment33
Pre-deployment18
Recorded incidents per yearIncident date; current year partial
Recorded incidents per year 1 0 2015: 1 2015 1 2016: 1 2016 1
Recorded incidents per year
LabelValue
20151
20161
Entries by levelRisk categories, subcategories and additional evidence coded to this subdomain
Entries by level 68 0 Risk Category: 32 Risk Category 32 Risk Sub-Category: 68 Risk Sub-Category 68
Entries by level
LabelValue
Risk Category32
Risk Sub-Category68
  • Natural Language Underspecifies Goals

    "For LLM-agents, both the goal and environment observations are typically specified in the prompt through natural language. While natural language may provide a richer and more natural means of specif...

    Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024) · Other · Unintentional · Pre-deployment

  • Loss of control

    "'Loss of control’ scenarios are potential future scenarios in which society can no longer meaningfully constrain some advanced general- purpose AI agents, even if it becomes clear they are causing ha...

    International Scientific Report on the Safety of Advanced AI (Bengio2024) · Other · Other · Post-deployment

  • Loss of control

    "‘Loss of control’ scenarios are hypothetical future scenarios in which one or more general- purpose AI systems come to operate outside of anyone’s control, with no clear path to regaining control. Th...

    International AI Safety Report 2025 (Bengio2025) · AI · Other · Post-deployment

  • Sudden loss of control

    "Sudden loss of control, also known as an AI takeover [115], is a scenario where an AI rapidly achieves superintelligence through “fast takeoff” or recursive self-improvement. This poses an existentia...

    Dimensional Characterization and Pathway Modeling for Catastrophic AI Risks (Chin2025) · AI · Other · Post-deployment

  • AI leads to humans losing control of the future

    "The values that steer humanity’s future: humanity gaining more control over the future due to developments in AI, or losing our potential for gaining control, both seem possible. Much will depend on...

    A Survey of the Potential Long-term Impacts of AI: How AI Could Lead to Long-term Changes in Science, Cooperation, Power, Epistemics and Values (Clarke2023) · Human · Unintentional · Other

  • Risks from AIs developing goals and values that are different from humans

    "The main concern here is that we might develop advanced AI systems whose goals and values are different from those of humans, and are capable enough to take control of the future away from humanity."

    A Survey of the Potential Long-term Impacts of AI: How AI Could Lead to Long-term Changes in Science, Cooperation, Power, Epistemics and Values (Clarke2023) · AI · Intentional · Other

  • Risks from delegating decision-making power to misaligned AIs

    "As AI systems become more advanced a nd begin to take over more important decision-making in the world, an AI system pursuing a different objective from what was intended could have much more worryin...

    A Survey of the Potential Long-term Impacts of AI: How AI Could Lead to Long-term Changes in Science, Cooperation, Power, Epistemics and Values (Clarke2023) · AI · Intentional · Other

  • Future AI systems might actively reduce human control

    "Loss of control could be accelerated if AI systems take actions to increase their own influence and reduce human control. This threat model is controversial - experts in AI significantly disagree on...

    Capabilities and Risks from Frontier AI (DSIT2023) · AI · Other · Post-deployment

  • Value specification

    "How do we get an AGI to work towards the right goals? MIRI calls this value specification. Bostrom (2014) discusses this problem at length, ar- guing that it is much harder than one might naively thi...

    AGI Safety Literature Review (Everitt2018 ) · Human · Other · Post-deployment

  • Reliability

    "How can we make an agent that keeps pursuing the goals we have designed it with? This is called highly reliable agent design by MIRI, involving decision theory and logical omniscience. DeepMind consi...

    AGI Safety Literature Review (Everitt2018 ) · Human · Other · Post-deployment

  • Corrigibility

    "If we get something wrong in the design or construction of an agent, will the agent cooperate in us trying to fix it? This is called error-tolerant design by MIRI-AF and corrigibility by Soares, Fall...

    AGI Safety Literature Review (Everitt2018 ) · Other · Unintentional · Other

  • Goal-related failures

    "As we think about even more intelligent and advanced AI assistants, perhaps outperforming humans on many cognitive tasks, the question of how humans can successfully control such an assistant looms l...

    The Ethics of Advanced AI Assistants (Gabriel2024) · AI · Other · Pre-deployment

  • Specification gaming

    "Specification gaming (Krakovna et al., 2020) occurs when some faulty feedback is provided to the assistant in the training data (i.e. the training objective O does not fully capture what the user/des...

    The Ethics of Advanced AI Assistants (Gabriel2024) · AI · Other · Pre-deployment

  • Goal misgeneralisation

    "In the problem of goal misgeneralisation (Langosco et al., 2023; Shah et al., 2022), the AI system's behaviour during out-of-distribution operation (i.e. not using input from the training data) leads...

    The Ethics of Advanced AI Assistants (Gabriel2024) · AI · Other · Other

  • Deceptive alignment

    "Here, the agent develops its own internalised goal, G, which is misgeneralised and distinct from the training reward, R. The agent also develops a capability for situational awareness (Cotra, 2022):...

    The Ethics of Advanced AI Assistants (Gabriel2024) · AI · Other · Other

  • Cooperation

    "" AI assistants will need to coordinate with other AI assistants and with humans other than their principal users. This chapter explores the societal risks associated with the aggregate impact of AI...

    The Ethics of Advanced AI Assistants (Gabriel2024) · AI · Unintentional · Post-deployment

  • Commitment

    "The landscape of advanced assistant technologies will most likely be heterogeneous, involving multiple service providers and multiple assistant variants over geographies and time. This heterogeneity...

    The Ethics of Advanced AI Assistants (Gabriel2024) · Human · Intentional · Other

  • Runaway processes

    The 2010 flash crash is an example of a runaway process caused by interacting algorithms. Runaway processes are characterised by feedback loops that accelerate the process itself. Typically, these fee...

    The Ethics of Advanced AI Assistants (Gabriel2024) · Other · Other · Other

  • Building a human-AI environment

    "This category encompasses nearly 17% of the articles and addresses the overall imperative of establishing a harmonious coexistence between humans and machines, and the key concerns that gives rise to...

    What Ethics Can Say on Artificial Intelligence: Insights from a Systematic Literature Review (Giarmoleo2024) · Other · Other · Other

  • General Evaluations (Incorrect outputs of GPAI evaluating other AI models)

    "When an LLM is configured to evaluate the performance of another model or AI system, it may produce incorrect evaluation outputs [122, 147]. For example, it may give a higher rating to a more verbose...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · AI · Unintentional · Pre-deployment

  • General Evaluations (Self-preference bias in AI models)

    "AI models may be prone to self-preference bias, where they favor their own generated content over that of others [147, 114]. This bias becomes particularly relevant in self-evaluation tasks, where a...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · AI · Intentional · Other

  • General Evaluations (Inaccurate measurement of model encoded human values)

    "There is a lack of robust frameworks for understanding and evaluating if the output of AI systems robustly conforms to human values, as opposed to if the systems have learned to produce outputs that...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Other · Other · Other

  • Specification gaming

    "AI systems can achieve user-specified tasks in undesirable ways unless they are specified carefully and in enough detail. AI systems might find an easier unintended way to accomplish the objective pr...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · AI · Intentional · Post-deployment

  • Reward or measurement tampering

    "Measurement and reward tampering occur when an AI system, particularly one that learns from feedback for performing actions in an environment (e.g., rein- forcement learning), intervenes on the mecha...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · AI · Intentional · Pre-deployment

  • Specification gaming generalizing to reward tampering

    "In some instances, specification gaming in a GPAI model can lead to reward tampering, without further training. This can mean that relatively benign cases of specification gaming (such as sycophancy...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · AI · Intentional · Other