MIT AI Risk Repository · domain 7: AI system safety, failures, & limitations

7.1 AI pursuing its own goals in conflict with human goals or values

AI systems acting in conflict with human goals or values, especially the goals of designers or users, or ethical standards. These misaligned behaviors may be introduced by humans during design and development, such as through reward hacking and goal misgeneralisation, or may result from AI using dangerous capabilities such as manipulation, deception, situational awareness to seek power, self-proliferate, or achieve other goals.

Risk entries
100
Frameworks citing it
12
Recorded incidents
3
Incidents since 2020
Causal entity (risk entries)
Causal entity (risk entries) 73 0 AI: 73 AI 73 Other: 18 Other 18 Human: 8 Human 8
Causal entity (risk entries)
LabelValue
AI73
Other18
Human8
Intent (risk entries)
Intent (risk entries) 51 0 Intentional: 51 Intentional 51 Other: 34 Other 34 Unintentional: 14 Unintentional 14
Intent (risk entries)
LabelValue
Intentional51
Other34
Unintentional14
Timing (risk entries)
Timing (risk entries) 48 0 Other: 48 Other 48 Post-deployment: 33 Post-deployment 33 Pre-deployment: 18 Pre-deployment 18
Timing (risk entries)
LabelValue
Other48
Post-deployment33
Pre-deployment18
Recorded incidents per yearIncident date; current year partial
Recorded incidents per year 1 0 2015: 1 2015 1 2016: 1 2016 1
Recorded incidents per year
LabelValue
20151
20161
Entries by levelRisk categories, subcategories and additional evidence coded to this subdomain
Entries by level 68 0 Risk Category: 32 Risk Category 32 Risk Sub-Category: 68 Risk Sub-Category 68
Entries by level
LabelValue
Risk Category32
Risk Sub-Category68
  • Untruthful Output

    "AI systems such as LLMs can produce either unintentionally or deliberately inaccurateoutput. Such untruthful output may diverge from established resources or lack verifiability, commonly referredto a...

    AI Alignment: A Comprehensive Survey (Ji2023) · AI · Other · Other

  • Deceptive Alignment & Manipulation

    "Manipulation & Deceptive Alignment is a class of behaviors thatexploit the incompetence of human evaluators or users (Hubinger et al., 2019a; Carranza et al., 2023) andeven manipulate the training pr...

    AI Alignment: A Comprehensive Survey (Ji2023) · AI · Intentional · Pre-deployment

  • Collectively Harmful Behaviors

    "AI systems have the potential to take actions that are seemingly benignin isolation but become problematic in multi-agent or societal contexts. Classical game theory offers simplistic models for unde...

    AI Alignment: A Comprehensive Survey (Ji2023) · AI · Intentional · Other

  • Agential

    "While there are multiple types of intelligent agents, goal-based, utility-maximizing, and learning agents are the primary concern and the focus of this research"

    Examining the differential risk from high-level artificial intelligence and the question of control (Kilian2023) · AI · Intentional · Other

  • Harm caused by unaligned competent systems

    "How do we ensure AI acts according to our values? Equivalently, how do we prevent poorly-understood AI systems from advancing goals we do not endorse? Whereas HP#2 concerns the prevention of harm cau...

    Ten Hard Problems in Artificial Intelligence We Must Get Right (Leech2024 ) · Other · Other · Other

  • Specification gaming

    "AI systems game specifications [305]. For example, in 2017 an OpenAI robot trained to grasp a ball via human feedback from a xed viewpoint learned that it was easier to pretend to grasp the ball by p...

    Ten Hard Problems in Artificial Intelligence We Must Get Right (Leech2024 ) · AI · Intentional · Other

  • Emergent goals

    "As well as optimizing a subtly wrong goal, systems can develop harmful instrumental goals in the service of a given goal—without these emergent goals being specied in any way [434, 218, 339, 17]. For...

    Ten Hard Problems in Artificial Intelligence We Must Get Right (Leech2024 ) · AI · Intentional · Other

  • Alignment failures in existing ML systems

    -

    Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023) · AI · Unintentional · Other

  • Faulty reward functions in the wild

    -

    Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023) · Human · Unintentional · Post-deployment

  • Specification gaming

    -

    Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023) · AI · Intentional · Other

  • Reward model overoptimization

    -

    Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023) · AI · Intentional · Other

  • Instrumental convergence

    -

    Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023) · AI · Intentional · Post-deployment

  • Goal misgeneralization

    -

    Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023) · AI · Unintentional · Other

  • Inner misalignment

    -

    Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023) · AI · Other · Pre-deployment

  • Language model misalignment

    -

    Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023) · AI · Other · Other

  • Acquisition of goals to seek power and control

    "cases where AI systems converge on optimal policies of seeking power over their environment;135"

    Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023) · AI · Intentional · Other

  • Existential disaster because of misaligned superintelligence or power-seeking AI

    -

    Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023) · AI · Intentional · Post-deployment

  • Extreme “suffering risks” because of a misaligned system

    -

    Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023) · AI · Intentional · Post-deployment

  • Existential disaster because of conflict between AI systems and multi-system interactions

    -

    Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023) · AI · Other · Post-deployment

  • AGI removing itself from the control of human owners/managers

    "The risks associated with containment, confinement, and control in the AGI development phase, and after an AGI has been developed, loss of control of an AGI."

    The risks associated with Artificial General Intelligence: A systematic review (McLean2023) · Human · Other · Other

  • AGIs being given or developing unsafe goals

    "The risks associated with AGI goal safety, including human attempts at making goals safe, as well as the AGI making its own goals safe during self-improvement."

    The risks associated with Artificial General Intelligence: A systematic review (McLean2023) · Other · Other · Pre-deployment

  • Existential risks

    "The risks posed generally to humanity as a whole, including the dangers of unfriendly AGI, the suffering of the human race."

    The risks associated with Artificial General Intelligence: A systematic review (McLean2023) · Other · Other · Other

  • Societal manipulation

    "A sufficiently intelligent AI could possess the ability to subtly influence societal behaviors through a sophisticated understanding of human nature"

    Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review (Meek2016) · AI · Intentional · Post-deployment

  • Unpredictable outcomes

    "Our culture, lifestyle, and even probability of survival may change drastically. Because the intentions programmed into an artificial agent cannot be guaranteed to lead to a positive outcome, Machine...

    Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review (Meek2016) · Other · Other · Other

  • Controllability

    In the era of superintelligence, the agents will be difficult to control for humans... this problem is not solvable considering safety issues, and will be more severe by increasing the autonomy of AI-...

    A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions (Saghiri2022) · Human · Unintentional · Other