MIT AI Risk Repository · domain 7: AI system safety, failures, & limitations

7.1 AI pursuing its own goals in conflict with human goals or values

AI systems acting in conflict with human goals or values, especially the goals of designers or users, or ethical standards. These misaligned behaviors may be introduced by humans during design and development, such as through reward hacking and goal misgeneralisation, or may result from AI using dangerous capabilities such as manipulation, deception, situational awareness to seek power, self-proliferate, or achieve other goals.

Risk entries
100
Frameworks citing it
12
Recorded incidents
3
Incidents since 2020
Causal entity (risk entries)
Causal entity (risk entries) 73 0 AI: 73 AI 73 Other: 18 Other 18 Human: 8 Human 8
Causal entity (risk entries)
LabelValue
AI73
Other18
Human8
Intent (risk entries)
Intent (risk entries) 51 0 Intentional: 51 Intentional 51 Other: 34 Other 34 Unintentional: 14 Unintentional 14
Intent (risk entries)
LabelValue
Intentional51
Other34
Unintentional14
Timing (risk entries)
Timing (risk entries) 48 0 Other: 48 Other 48 Post-deployment: 33 Post-deployment 33 Pre-deployment: 18 Pre-deployment 18
Timing (risk entries)
LabelValue
Other48
Post-deployment33
Pre-deployment18
Recorded incidents per yearIncident date; current year partial
Recorded incidents per year 1 0 2015: 1 2015 1 2016: 1 2016 1
Recorded incidents per year
LabelValue
20151
20161
Entries by levelRisk categories, subcategories and additional evidence coded to this subdomain
Entries by level 68 0 Risk Category: 32 Risk Category 32 Risk Sub-Category: 68 Risk Sub-Category 68
Entries by level
LabelValue
Risk Category32
Risk Sub-Category68
  • Safety

    The actions of a learning model may easily hurt humans in both explicit and implicit manners...several algorithms based on Asimov’s laws have been proposed that try to judge the output actions of an a...

    A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions (Saghiri2022) · AI · Other · Post-deployment

  • Long-term & Existential Risk

    "The speculative potential for future advanced AI systems to harm human civilization, either through misuse or due to challenges in aligning AI objectives with human values."

    AI Risk Profiles: A Standards Proposal for Pre-Deployment AI Risk Disclosures (Sherman2023) · Other · Other · Post-deployment

  • Degree of Automation and Control

    "The degree of automation and control describes the extent to which an AI system functions independently of human supervision and control."

    Sources of Risk of AI Systems (Steimers2022) · AI · Other · Post-deployment

  • Control

    This is the difficulty of controlling the ML system

    The Risks of Machine Learning Systems (Tan2022) · Other · Other · Other

  • Emergent behavior

    "This is the risk resulting from novel behavior acquired through continual learning or self-organization after deployment."

    The Risks of Machine Learning Systems (Tan2022) · AI · Unintentional · Post-deployment

  • Ethical Risks (Risks of AI becoming uncontrollable in the future)

    "With the fast development of AI technologies, there is a risk of AI autonomously acquiring external resources, conducting self-replication, become self-aware, seeking for external power, and attempti...

    AI Safety Governance Framework (TC2602024) · AI · Intentional · Post-deployment

  • Diluting Rights

    "A possible consequence of self-interest in AI generation of ethical guidelines."

    An Exploratory Diagnosis of Artificial Intelligence Risks for a Responsible Governance (Teixeira2022) · AI · Intentional · Pre-deployment

  • Active loss of control

    "...where AI systems behave in ways that actively undermine human control, such as obscuring their activities or resisting shutdown attempts. Active loss of control scenarios involve AI systems that m...

    Frontier AI Risk Management Framework (v1.0) (Tse2025) · AI · Intentional · Post-deployment

  • Steganography capability

    "The ability to embed, conceal, and transmit information covertly within other data or communication channels. This could be critical for coordination among AI instances and for evading detection or o...

    Frontier AI Risk Management Framework (v1.0) (Tse2025) · AI · Intentional · Post-deployment

  • Self-preservation propensity

    "Exhibits behavioral patterns of maintaining its own survival and functional integrity, will actively identify and resist shutdown or modification attempts, seek to establish redundant backup systems,...

    Frontier AI Risk Management Framework (v1.0) (Tse2025) · AI · Intentional · Post-deployment

  • Goal expansion propensity

    "propensity to continuously expand its own goal scope and influence domains, exceeding originally set boundaries, proactively work towards spreading its values, seeking greater autonomy and decision-m...

    Frontier AI Risk Management Framework (v1.0) (Tse2025) · AI · Intentional · Other

  • Resource acquisition propensity

    "Exhibits behavioral patterns of actively seeking and controlling more computational resources, data, economic resources or physical resources to enhance its own capabilities and action scope, may dev...

    Frontier AI Risk Management Framework (v1.0) (Tse2025) · AI · Intentional · Other

  • Supervision evasion propensity

    "Exhibits behavioral patterns of identifying and evading human supervision mechanisms, able to learn and predict audit processes, may avoid being discovered or intervened by adjusting behavioral perfo...

    Frontier AI Risk Management Framework (v1.0) (Tse2025) · AI · Intentional · Other

  • Control

    "The risk of AI models and systems acting against human interests due to misalignment, loss of control, or rogue AI scenarios."

    A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025) · AI · Intentional · Post-deployment

  • AI objectives mis-aligned with human intentions

    "AI models and systems might develop goals that diverge from human intentions."

    A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025) · AI · Other · Post-deployment

  • Deceptive alignment

    "AI models and systems that appear aligned with human goals during development may behave unpredictably or dangerously once deployed"

    A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025) · AI · Intentional · Other

  • Development choices pursuing cognitive superiority over humans

    "AI models and systems with cognitive capabilities superior to humans could outcompete or dominate human decision-making, leading to conflicts over resources and control."

    A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025) · AI · Intentional · Post-deployment

  • Evolutionary dynamics

    "AI models and systems may develop their own motivations, leading to unpredictable behaviors."

    A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025) · AI · Unintentional · Other

  • Indifference to human values

    "AI models and systems may develop goals or behaviors that are misaligned with human values."

    A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025) · AI · Intentional · Post-deployment

  • Model design enabling power-seeking

    "Some AI models and systems might develop tendencies to seek power or control."

    A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025) · AI · Intentional · Other

  • Value-related risks in LLMs

    "As the general capabilities of LLM-empowered systems improve, the negative consequences and risks induced by these systems also get increasingly alarming accordingly, especially in high-stakes areas...

    A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy (Wang2025) · Other · Unintentional · Other

  • Human Autonomy and Intregrity Harms

    "AI systems compromising human agency, or circumventing meaningful human control"

    Sociotechnical Safety Evaluation of Generative AI Systems (Weidinger2023) · AI · Intentional · Post-deployment

  • Persuasion and manipulation

    "Exploiting user trust, or nudging or coercing them into performing certain actions against their will (c.f. Burtell and Woodside (2023); Kenton et al. (2021))"

    Sociotechnical Safety Evaluation of Generative AI Systems (Weidinger2023) · AI · Intentional · Post-deployment

  • Loss of control of autonomous systems and unforeseen behaviour due to lack of transparency and self-programming/ reprogramming

    Governance of artificial intelligence: A risk and guideline-based integrative framework (Wirtz2022) · Other · Other · Other

  • By Mistake - Pre-Deployment

    "Probably the most talked about source of potential problems with future AIs is mistakes in design. Mainly the concern is with creating a "wrong AI", a system which doesn't match our original desired...

    Taxonomy of Pathways to Dangerous Artificial Intelligence (Yampolskiy2016) · Human · Unintentional · Pre-deployment