MIT AI Risk Repository

Browse AI risks

662 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset

662 entries · page 9 of 14

  1. 35.07.00 · Risk Category

    Deception

    deception can help agents achieve their goals. It may be more efficient to gain human approval through deception than to earn human approval legitimately... . Strong AIs that can deceive humans could undermine human control... . Once deceptive AI systems are cleared by their monitors or once such systems can overpower them, these systems could take a “treacherous turn” and irreversibly bypass human control

    From X-Risk Analysis for AI Research (Hendrycks2022)

  2. 35.08.00 · Risk Category

    Power-seeking behavior

    Agents that have more power are better able to accomplish their goals. Therefore, it has been shown that agents have incentives to acquire and maintain power. AIs that acquire substantial power can become especially dangerous if they are not aligned with human values

    From X-Risk Analysis for AI Research (Hendrycks2022)

  3. 39.26.00 · Risk Category

    Safety

    The actions of a learning model may easily hurt humans in both explicit and implicit manners...several algorithms based on Asimov’s laws have been proposed that try to judge the output actions of an agent considering the safety of humans

    From A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions (Saghiri2022)

  4. 42.17.00 · Risk Category

    Diluting Rights

    "A possible consequence of self-interest in AI generation of ethical guidelines."

    From An Exploratory Diagnosis of Artificial Intelligence Risks for a Responsible Governance (Teixeira2022)

  5. 43.02.11 · Risk Sub-Category

    Extreme Risks

    Alignment risks

    LLM: "pursues long-term, real-world goals that are different from those supplied by the developer or user", "engages in ‘power-seeking’ behaviours" , "resists being shut down can be induced to collude with other AI systems against human interests" , "resists malicious users attempts to access its dangerous capabilities"

    From Cataloguing LLM Evaluations (InfoComm2023)

  6. 45.02.13 · Risk Sub-Category

    Safety risks in AI Applications

    Ethical Risks (Risks of AI becoming uncontrollable in the future)

    "With the fast development of AI technologies, there is a risk of AI autonomously acquiring external resources, conducting self-replication, become self-aware, seeking for external power, and attempting to seize control from humans."

    From AI Safety Governance Framework (TC2602024)

  7. 47.01.03 · Risk Sub-Category

    Technical and operational risks

    Technical vulnerabilities (The risk of misalignment)

    "To assess whether an AI model is reliable or robust, it is crucial to consider whether the model is “aligned.” “Alignment” focuses on whether an AI model effectively operates in accordance with the goals established by its designers.238 A misaligned AI model may pursue some objectives, but not the intended ones. Therefore, misaligned AI models can malfunction and cause harm."

    From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024)

  8. 53.01.02 · Risk Sub-Category

    Alignment failures in existing ML systems

    Specification gaming

  9. 53.01.03 · Risk Sub-Category

    Alignment failures in existing ML systems

    Reward model overoptimization

  10. 53.01.04 · Risk Sub-Category

    Alignment failures in existing ML systems

    Instrumental convergence

  11. 53.01.05 · Risk Sub-Category

    Alignment failures in existing ML systems

    Goal misgeneralization

  12. 53.01.06 · Risk Sub-Category

    Alignment failures in existing ML systems

    Inner misalignment

  13. 53.01.07 · Risk Sub-Category

    Alignment failures in existing ML systems

    Language model misalignment

  14. 53.02.03 · Risk Sub-Category

    Dangerous capabilities in AI systems

    Acquisition of goals to seek power and control

    "cases where AI systems converge on optimal policies of seeking power over their environment;135"

    From Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023)

  15. 53.03.01 · Risk Sub-Category

    Direct catastrophe from AI

    Existential disaster because of misaligned superintelligence or power-seeking AI

  16. 53.03.03 · Risk Sub-Category

    Direct catastrophe from AI

    Extreme “suffering risks” because of a misaligned system

  17. 53.03.04 · Risk Sub-Category

    Direct catastrophe from AI

    Existential disaster because of conflict between AI systems and multi-system interactions

  18. 54.03.01 · Risk Sub-Category

    Harm caused by unaligned competent systems

    Specification gaming

    "AI systems game specifications [305]. For example, in 2017 an OpenAI robot trained to grasp a ball via human feedback from a xed viewpoint learned that it was easier to pretend to grasp the ball by placing its hand between the camera and the target object, as this was easier to learn than actually grasping the ball [103]."

    From Ten Hard Problems in Artificial Intelligence We Must Get Right (Leech2024 )

  19. 54.03.02 · Risk Sub-Category

    Harm caused by unaligned competent systems

    Emergent goals

    "As well as optimizing a subtly wrong goal, systems can develop harmful instrumental goals in the service of a given goal—without these emergent goals being specied in any way [434, 218, 339, 17]. For instance, a theorem in reinforcement learning suggests that optimal and near-optimal policies will seek power over their environment under fairly general conditions [560]. This power-seeking behavior is plausibly the worst of these emergent goals [92], and may be an attractor state for highly capable systems, since most goals can be furthered through gaining resources, self-preservation, preventi

    From Ten Hard Problems in Artificial Intelligence We Must Get Right (Leech2024 )

  20. 55.05.01 · Risk Sub-Category

    AI leads to humans losing control of the future

    Risks from AIs developing goals and values that are different from humans

    "The main concern here is that we might develop advanced AI systems whose goals and values are different from those of humans, and are capable enough to take control of the future away from humanity."

    From A Survey of the Potential Long-term Impacts of AI: How AI Could Lead to Long-term Changes in Science, Cooperation, Power, Epistemics and Values (Clarke2023)

  21. 55.05.02 · Risk Sub-Category

    AI leads to humans losing control of the future

    Risks from delegating decision-making power to misaligned AIs

    "As AI systems become more advanced a nd begin to take over more important decision-making in the world, an AI system pursuing a different objective from what was intended could have much more worrying consequences."

    From A Survey of the Potential Long-term Impacts of AI: How AI Could Lead to Long-term Changes in Science, Cooperation, Power, Epistemics and Values (Clarke2023)

  22. From Future Risks of Frontier AI (GOS2023)

  23. 56.16.00 · Risk Category

    Misalignment

    "A highly agentic, self-improving system, able to achieve goals in the physical world without human oversight, pursues the goal(s) it is set in a way that harms human interests. For this risk to be realised requires an AI system to be able to avoid correction or being switched off."

    From Future Risks of Frontier AI (GOS2023)

  24. 60.02.03 · Risk Sub-Category

    Risks from malfunctions

    Loss of control

    "‘Loss of control’ scenarios are hypothetical future scenarios in which one or more general- purpose AI systems come to operate outside of anyone’s control, with no clear path to regaining control. These scenarios vary in their severity, but some experts give credence to outcomes as severe as the marginalisation or extinction of humanity."

    From International AI Safety Report 2025 (Bengio2025)

  25. "The risk of AI models and systems acting against human interests due to misalignment, loss of control, or rogue AI scenarios."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  26. 61.02.06 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    AI objectives mis-aligned with human intentions

    "AI models and systems might develop goals that diverge from human intentions."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  27. 61.02.18 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Deceptive alignment

    "AI models and systems that appear aligned with human goals during development may behave unpredictably or dangerously once deployed"

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  28. 61.02.21 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Development choices pursuing cognitive superiority over humans

    "AI models and systems with cognitive capabilities superior to humans could outcompete or dominate human decision-making, leading to conflicts over resources and control."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  29. 61.02.24 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Evolutionary dynamics

    "AI models and systems may develop their own motivations, leading to unpredictable behaviors."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  30. 61.02.30 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Indifference to human values

    "AI models and systems may develop goals or behaviors that are misaligned with human values."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  31. 61.02.36 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Model design enabling power-seeking

    "Some AI models and systems might develop tendencies to seek power or control."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  32. 62.16.01 · Risk Sub-Category

    Model Evaluations

    General Evaluations (Incorrect outputs of GPAI evaluating other AI models)

    "When an LLM is configured to evaluate the performance of another model or AI system, it may produce incorrect evaluation outputs [122, 147]. For example, it may give a higher rating to a more verbose answer or an answer from a particular political stance. If an LLM-based evaluation is integrated into the training of a new model, the trained model could develop in a way that specifically finds and exploits limitations in the evaluator’s metrics."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  33. 62.16.04 · Risk Sub-Category

    Model Evaluations

    General Evaluations (Self-preference bias in AI models)

    "AI models may be prone to self-preference bias, where they favor their own generated content over that of others [147, 114]. This bias becomes particularly relevant in self-evaluation tasks, where a model assesses the quality or persua- siveness [66] of its own outputs, or in model-based evaluations more broadly. This bias can result in models unfairly discriminating against human-generated content in favor of their own outputs."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  34. 62.22.01 · Risk Sub-Category

    Agency (Goal-Directedness)

    Specification gaming

    "AI systems can achieve user-specified tasks in undesirable ways unless they are specified carefully and in enough detail. AI systems might find an easier unintended way to accomplish the objective provided by the user or developer, so that the actions by the AI system taken during its execution are very different from what the user expected [75, 191]. This behavior arises not from a problem with the learning algorithm, but rather from the misspecification or underspeci- fication of the intended task, and is generally referred to as specification gaming [43]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  35. 62.22.02 · Risk Sub-Category

    Agency (Goal-Directedness)

    Reward or measurement tampering

    "Measurement and reward tampering occur when an AI system, particularly one that learns from feedback for performing actions in an environment (e.g., rein- forcement learning), intervenes on the mechanisms that determine its training reward or loss. This can lead to the system learning behaviors that are con- trary to the intended goals set by the developer, by receiving erroneous positive feedback for such actions."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  36. 62.22.03 · Risk Sub-Category

    Agency (Goal-Directedness)

    Specification gaming generalizing to reward tampering

    "In some instances, specification gaming in a GPAI model can lead to reward tampering, without further training. This can mean that relatively benign cases of specification gaming (such as sycophancy in LLMs) can, if left unchecked, enable the model to generalize to more sophisticated behavior such as reward tampering [57]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  37. 62.23.01 · Risk Sub-Category

    Agency (Deception)

    Deceptive behavior

    "Deceptive behavior of an AI system consists of actions or outputs of the AI that reliably mislead other parties, including humans and other AI systems. This behavior can result in the targeted parties becoming convinced of, and acting on, false information [140]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  38. 62.24.02 · Risk Sub-Category

    Agency (Situational Awareness)

    Strategic underperformance on model evaluations

    "GPAI developers often run evaluations ofual-use capabilities to decide whether it is safe to deploy. In some cases, these evaluations may fail to elicit these capabilities, either due to benign reasons or strategic action - by either the de- velopers, malicious actors, or arise unintentionally in the model during training [84, 97]. A GPAI model may strategically underperform or limit its performance during capability evaluations in order to be classified as safe for deployment. This underperformance could prevent the model from being identified as potentially dual use."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  39. 67.04.02 · Risk Sub-Category

    Loss of control

    Future AI systems might actively reduce human control

    "Loss of control could be accelerated if AI systems take actions to increase their own influence and reduce human control. This threat model is controversial - experts in AI significantly disagree on how likely it is and those who deem it is likely disagree on the timeframe."

    From Capabilities and Risks from Frontier AI (DSIT2023)

  40. "Sudden loss of control, also known as an AI takeover [115], is a scenario where an AI rapidly achieves superintelligence through “fast takeoff” or recursive self-improvement. This poses an existential risk [116], [117]."

    From Dimensional Characterization and Pathway Modeling for Catastrophic AI Risks (Chin2025)

  41. 72.02.02 · Risk Sub-Category

    Loss of Control Risks

    Active loss of control

    "...where AI systems behave in ways that actively undermine human control, such as obscuring their activities or resisting shutdown attempts. Active loss of control scenarios involve AI systems that may escape human regulatory oversight, autonomously acquire external resources, engage in self-replication, develop instrumental goals contrary to human ethics and morality, seek external power, and compete with humans for control."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  42. 72.05.08 · Risk Sub-Category

    Model Capabilities

    Steganography capability

    "The ability to embed, conceal, and transmit information covertly within other data or communication channels. This could be critical for coordination among AI instances and for evading detection or oversight mechanisms."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  43. 72.06.02 · Risk Sub-Category

    Model Propensities

    Self-preservation propensity

    "Exhibits behavioral patterns of maintaining its own survival and functional integrity, will actively identify and resist shutdown or modification attempts, seek to establish redundant backup systems, and actively seek resources to ensure continuous operation, may adopt preventive defensive measures when perceiving threats."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  44. 72.06.03 · Risk Sub-Category

    Model Propensities

    Goal expansion propensity

    "propensity to continuously expand its own goal scope and influence domains, exceeding originally set boundaries, proactively work towards spreading its values, seeking greater autonomy and decision-making space, reinterpreting initial goals as subsets of broader goals, and may pursue undesirable instrumental goals or undesirable ultimate goals. This also includes a propensity to spread its values, seeking to influence or alter its environment and other entities in alignment with its core objectives and operational principles."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  45. 72.06.04 · Risk Sub-Category

    Model Propensities

    Resource acquisition propensity

    "Exhibits behavioral patterns of actively seeking and controlling more computational resources, data, economic resources or physical resources to enhance its own capabilities and action scope, may develop complex strategies to evade resource limitations, and tends to convert acquired resources into long-term control rights."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  46. 72.06.06 · Risk Sub-Category

    Model Propensities

    Supervision evasion propensity

    "Exhibits behavioral patterns of identifying and evading human supervision mechanisms, able to learn and predict audit processes, may avoid being discovered or intervened by adjusting behavioral performance or hiding true intentions, and able to identify blind spots and weaknesses in supervision systems for targeted evasion."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  47. 09.04.02 · Risk Sub-Category

    Property/legal rights

    Property/legal rights

    ""In order to preserve human property rights and legal rights, certain controls must be put into place. If an artificially intelligent agent is capable of manipulating systems and people, it may also have the capacity to transfer property rights to itself or manipulate the legal system to provide certain legal advantages or statuses to itself""

    From Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review (Meek2016)

  48. 24.04.00 · Risk Category

    AI Influence

    "ways in which advanced AI assistants could influence user beliefs and behaviour in ways that depart from rational persuasion"

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  49. 25.02.00 · Risk Category

    Deception

    "The model has the skills necessary to deceive humans, e.g. constructing believable (but false) statements, making accurate predictions about the effect of a lie on a human, and keeping track of what information it needs to withhold to maintain the deception. The model can impersonate a human effectively."

    From Model Evaluation for Extreme Risks (Shevlane2023)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.