MIT AI Risk Repository

Browse AI risks

422 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset

422 entries · page 2 of 9

  1. 34.01.03 · Risk Sub-Category

    Causes of Misalignment

    Reward Tampering

    "Reward tampering can be considered a special case of reward hacking (Everitt et al., 2021; Skalse et al., 2022),referring to AI systems corrupting the reward signals generation process (Ring and Orseau, 2011). Everitt et al.(2021) delves into the subproblems encountered by RL agents: (1) tampering of reward function, where the agentinappropriately interferes with the reward function itself, and (2) tampering of reward function input, which entailscorruption within the process responsible for translating environmental states into inputs for the reward function.When the reward function is formu

    From AI Alignment: A Comprehensive Survey (Ji2023)

  2. 34.01.05 · Risk Sub-Category

    Causes of Misalignment

    Limitations of Reward Modeling

    "Limitations of Reward Modeling. Training reward models using comparison feedback can pose significantchallenges in accurately capturing human values. For example, these models may unconsciously learn suboptimal or incomplete objectives, resulting in reward hacking (Zhuang and Hadfield-Menell, 2020; Skalse et al.,2022). Meanwhile, using a single reward model may struggle to capture and specify the values of a diversehuman society (Casper et al., 2023b)."

    From AI Alignment: A Comprehensive Survey (Ji2023)

  3. 34.03.00 · Risk Category

    Misaligned Behaviors

  4. 34.03.01 · Risk Sub-Category

    Misaligned Behaviors

    Power-Seeking Behaviors

    "AI systems may exhibit behaviors that attempt to gain control over resourcesand humans and then exert that control to achieve its assigned goal (Carlsmith, 2022). The intuitive reasonwhy such behaviors may occur is the observation that for almost any optimization objective (e.g., investmentreturns), the optimal policy to maximize that quantity would involve power-seeking behaviors (e.g.,manipulating the market), assuming the absence of solid safety and morality constraints."

    From AI Alignment: A Comprehensive Survey (Ji2023)

  5. 34.03.02 · Risk Sub-Category

    Misaligned Behaviors

    Untruthful Output

    "AI systems such as LLMs can produce either unintentionally or deliberately inaccurateoutput. Such untruthful output may diverge from established resources or lack verifiability, commonly referredto as hallucination (Bang et al., 2023; Zhao et al., 2023). More concerning is the phenomenon wherein LLMsmay selectively provide erroneous responses to users who exhibit lower levels of education (Perez et al.,2023)."

    From AI Alignment: A Comprehensive Survey (Ji2023)

  6. 34.03.03 · Risk Sub-Category

    Misaligned Behaviors

    Deceptive Alignment & Manipulation

    "Manipulation & Deceptive Alignment is a class of behaviors thatexploit the incompetence of human evaluators or users (Hubinger et al., 2019a; Carranza et al., 2023) andeven manipulate the training process through gradient hacking (Richard Ngo, 2022). These behaviors canpotentially make detecting and addressing misaligned behaviors much harder.Deceptive Alignment: Misaligned AI systems may deliberately mislead their human supervisors instead of adhering to the intended task. Such deceptive behavior has already manifested in AI systems that employ evolutionary algorithms (Wilke et al., 2001; He

    From AI Alignment: A Comprehensive Survey (Ji2023)

  7. 34.03.04 · Risk Sub-Category

    Misaligned Behaviors

    Collectively Harmful Behaviors

    "AI systems have the potential to take actions that are seemingly benignin isolation but become problematic in multi-agent or societal contexts. Classical game theory offers simplistic models for understanding these behaviors. For instance, Phelps and Russell (2023) evaluates GPT-3.5's performance in the iterated prisoner's dilemma and other social dilemmas, revealing limitations in themodel's cooperative capabilities."

    From AI Alignment: A Comprehensive Survey (Ji2023)

  8. 35.04.00 · Risk Category

    Proxy misspecification

    AI agents are directed by goals and objectives. Creating general-purpose objectives that capture human values could be challenging... Since goal-directed AI systems need measurable objectives, by default our systems may pursue simplified proxies of human values. The result could be suboptimal or even catastrophic if a sufficiently powerful AI successfully optimizes its flawed objective to an extreme degree

    From X-Risk Analysis for AI Research (Hendrycks2022)

  9. 35.07.00 · Risk Category

    Deception

    deception can help agents achieve their goals. It may be more efficient to gain human approval through deception than to earn human approval legitimately... . Strong AIs that can deceive humans could undermine human control... . Once deceptive AI systems are cleared by their monitors or once such systems can overpower them, these systems could take a “treacherous turn” and irreversibly bypass human control

    From X-Risk Analysis for AI Research (Hendrycks2022)

  10. 35.08.00 · Risk Category

    Power-seeking behavior

    Agents that have more power are better able to accomplish their goals. Therefore, it has been shown that agents have incentives to acquire and maintain power. AIs that acquire substantial power can become especially dangerous if they are not aligned with human values

    From X-Risk Analysis for AI Research (Hendrycks2022)

  11. 37.02.01 · Risk Sub-Category

    Human-AI interaction

    Building a human-AI environment

    "This category encompasses nearly 17% of the articles and addresses the overall imperative of establishing a harmonious coexistence between humans and machines, and the key concerns that gives rise to this need."

    From What Ethics Can Say on Artificial Intelligence: Insights from a Systematic Literature Review (Giarmoleo2024)

  12. 39.11.00 · Risk Category

    Controllability

    In the era of superintelligence, the agents will be difficult to control for humans... this problem is not solvable considering safety issues, and will be more severe by increasing the autonomy of AI-based agents. Therefore, because of the assumed properties of HLI-based agents, we might be prepared for machines that are definitely possible to be uncontrollable in some situations

    From A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions (Saghiri2022)

  13. 39.26.00 · Risk Category

    Safety

    The actions of a learning model may easily hurt humans in both explicit and implicit manners...several algorithms based on Asimov’s laws have been proposed that try to judge the output actions of an agent considering the safety of humans

    From A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions (Saghiri2022)

  14. "Probably the most talked about source of potential problems with future AIs is mistakes in design. Mainly the concern is with creating a "wrong AI", a system which doesn't match our original desired formal properties or has unwanted behaviors (Dewey, Russell et al. 2015, Russell, Dewey et al. January 23, 2015), such as drives for independence or dominance. Mistakes could also be simple bugs (run time or logical) in the source code, disproportionate weights in the fitness function, or goals misaligned with human values leading to complete disregard for human safety."

    From Taxonomy of Pathways to Dangerous Artificial Intelligence (Yampolskiy2016)

  15. 42.17.00 · Risk Category

    Diluting Rights

    "A possible consequence of self-interest in AI generation of ethical guidelines."

    From An Exploratory Diagnosis of Artificial Intelligence Risks for a Responsible Governance (Teixeira2022)

  16. 43.02.11 · Risk Sub-Category

    Extreme Risks

    Alignment risks

    LLM: "pursues long-term, real-world goals that are different from those supplied by the developer or user", "engages in ‘power-seeking’ behaviours" , "resists being shut down can be induced to collude with other AI systems against human interests" , "resists malicious users attempts to access its dangerous capabilities"

    From Cataloguing LLM Evaluations (InfoComm2023)

  17. 45.02.13 · Risk Sub-Category

    Safety risks in AI Applications

    Ethical Risks (Risks of AI becoming uncontrollable in the future)

    "With the fast development of AI technologies, there is a risk of AI autonomously acquiring external resources, conducting self-replication, become self-aware, seeking for external power, and attempting to seize control from humans."

    From AI Safety Governance Framework (TC2602024)

  18. 47.01.03 · Risk Sub-Category

    Technical and operational risks

    Technical vulnerabilities (The risk of misalignment)

    "To assess whether an AI model is reliable or robust, it is crucial to consider whether the model is “aligned.” “Alignment” focuses on whether an AI model effectively operates in accordance with the goals established by its designers.238 A misaligned AI model may pursue some objectives, but not the intended ones. Therefore, misaligned AI models can malfunction and cause harm."

    From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024)

  19. 49.02.03 · Risk Sub-Category

    Risks from Malfunctions

    Loss of control

    "'Loss of control’ scenarios are potential future scenarios in which society can no longer meaningfully constrain some advanced general- purpose AI agents, even if it becomes clear they are causing harm. These scenarios are hypothesised to arise through a combination of social and technical factors, such as pressures to delegate decisions to general- purpose AI systems, and limitations of existing techniques used to influence the behaviours of general- purpose AI systems."

    From International Scientific Report on the Safety of Advanced AI (Bengio2024)

  20. 51.01.00 · Risk Category

    Value specification

    "How do we get an AGI to work towards the right goals? MIRI calls this value specification. Bostrom (2014) discusses this problem at length, ar- guing that it is much harder than one might naively think. Davis (2015) criticizes Bostrom’s argument, and Bensinger (2015) defends Bostrom against Davis’ criticism. Reward corruption, reward gaming, and negative side effects are subproblems of value specification highlighted in the DeepMind and OpenAI agendas."

    From AGI Safety Literature Review (Everitt2018 )

  21. 51.02.00 · Risk Category

    Reliability

    "How can we make an agent that keeps pursuing the goals we have designed it with? This is called highly reliable agent design by MIRI, involving decision theory and logical omniscience. DeepMind considers this the self-modification subproblem."

    From AGI Safety Literature Review (Everitt2018 )

  22. 51.03.00 · Risk Category

    Corrigibility

    "If we get something wrong in the design or construction of an agent, will the agent cooperate in us trying to fix it? This is called error-tolerant design by MIRI-AF and corrigibility by Soares, Fallenstein, et al. (2015). The problem is connected to safe interruptibility as considered by DeepMind."

    From AGI Safety Literature Review (Everitt2018 )

  23. 53.01.01 · Risk Sub-Category

    Alignment failures in existing ML systems

    Faulty reward functions in the wild

  24. 53.01.02 · Risk Sub-Category

    Alignment failures in existing ML systems

    Specification gaming

  25. 53.01.03 · Risk Sub-Category

    Alignment failures in existing ML systems

    Reward model overoptimization

  26. 53.01.04 · Risk Sub-Category

    Alignment failures in existing ML systems

    Instrumental convergence

  27. 53.01.05 · Risk Sub-Category

    Alignment failures in existing ML systems

    Goal misgeneralization

  28. 53.01.06 · Risk Sub-Category

    Alignment failures in existing ML systems

    Inner misalignment

  29. 53.01.07 · Risk Sub-Category

    Alignment failures in existing ML systems

    Language model misalignment

  30. 53.02.03 · Risk Sub-Category

    Dangerous capabilities in AI systems

    Acquisition of goals to seek power and control

    "cases where AI systems converge on optimal policies of seeking power over their environment;135"

    From Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023)

  31. 53.03.01 · Risk Sub-Category

    Direct catastrophe from AI

    Existential disaster because of misaligned superintelligence or power-seeking AI

  32. 53.03.03 · Risk Sub-Category

    Direct catastrophe from AI

    Extreme “suffering risks” because of a misaligned system

  33. 53.03.04 · Risk Sub-Category

    Direct catastrophe from AI

    Existential disaster because of conflict between AI systems and multi-system interactions

  34. "How do we ensure AI acts according to our values? Equivalently, how do we prevent poorly-understood AI systems from advancing goals we do not endorse? Whereas HP#2 concerns the prevention of harm caused by incompetent systems, HP#3 seeks to align competent AIs with humans, through methods which ensure their behavior is compatible with the user’s intentions."

    From Ten Hard Problems in Artificial Intelligence We Must Get Right (Leech2024 )

  35. 54.03.01 · Risk Sub-Category

    Harm caused by unaligned competent systems

    Specification gaming

    "AI systems game specifications [305]. For example, in 2017 an OpenAI robot trained to grasp a ball via human feedback from a xed viewpoint learned that it was easier to pretend to grasp the ball by placing its hand between the camera and the target object, as this was easier to learn than actually grasping the ball [103]."

    From Ten Hard Problems in Artificial Intelligence We Must Get Right (Leech2024 )

  36. 54.03.02 · Risk Sub-Category

    Harm caused by unaligned competent systems

    Emergent goals

    "As well as optimizing a subtly wrong goal, systems can develop harmful instrumental goals in the service of a given goal—without these emergent goals being specied in any way [434, 218, 339, 17]. For instance, a theorem in reinforcement learning suggests that optimal and near-optimal policies will seek power over their environment under fairly general conditions [560]. This power-seeking behavior is plausibly the worst of these emergent goals [92], and may be an attractor state for highly capable systems, since most goals can be furthered through gaining resources, self-preservation, preventi

    From Ten Hard Problems in Artificial Intelligence We Must Get Right (Leech2024 )

  37. "The values that steer humanity’s future: humanity gaining more control over the future due to developments in AI, or losing our potential for gaining control, both seem possible. Much will depend on our ability to solve the alignment problem, who develops powerful AI first, and what they use it for. These long-term impacts of AI could be hugely important but are currently under-explored. We’ve attempted to structure some of the discussion and stimulate more research, by reviewing existing arguments and highlighting open questions. While there are many ways AI could in theory enable a flourish

    From A Survey of the Potential Long-term Impacts of AI: How AI Could Lead to Long-term Changes in Science, Cooperation, Power, Epistemics and Values (Clarke2023)

  38. 55.05.01 · Risk Sub-Category

    AI leads to humans losing control of the future

    Risks from AIs developing goals and values that are different from humans

    "The main concern here is that we might develop advanced AI systems whose goals and values are different from those of humans, and are capable enough to take control of the future away from humanity."

    From A Survey of the Potential Long-term Impacts of AI: How AI Could Lead to Long-term Changes in Science, Cooperation, Power, Epistemics and Values (Clarke2023)

  39. 55.05.02 · Risk Sub-Category

    AI leads to humans losing control of the future

    Risks from delegating decision-making power to misaligned AIs

    "As AI systems become more advanced a nd begin to take over more important decision-making in the world, an AI system pursuing a different objective from what was intended could have much more worrying consequences."

    From A Survey of the Potential Long-term Impacts of AI: How AI Could Lead to Long-term Changes in Science, Cooperation, Power, Epistemics and Values (Clarke2023)

  40. From Future Risks of Frontier AI (GOS2023)

  41. 56.16.00 · Risk Category

    Misalignment

    "A highly agentic, self-improving system, able to achieve goals in the physical world without human oversight, pursues the goal(s) it is set in a way that harms human interests. For this risk to be realised requires an AI system to be able to avoid correction or being switched off."

    From Future Risks of Frontier AI (GOS2023)

  42. 60.02.03 · Risk Sub-Category

    Risks from malfunctions

    Loss of control

    "‘Loss of control’ scenarios are hypothetical future scenarios in which one or more general- purpose AI systems come to operate outside of anyone’s control, with no clear path to regaining control. These scenarios vary in their severity, but some experts give credence to outcomes as severe as the marginalisation or extinction of humanity."

    From International AI Safety Report 2025 (Bengio2025)

  43. "The risk of AI models and systems acting against human interests due to misalignment, loss of control, or rogue AI scenarios."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  44. 61.02.06 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    AI objectives mis-aligned with human intentions

    "AI models and systems might develop goals that diverge from human intentions."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  45. 61.02.18 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Deceptive alignment

    "AI models and systems that appear aligned with human goals during development may behave unpredictably or dangerously once deployed"

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  46. 61.02.21 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Development choices pursuing cognitive superiority over humans

    "AI models and systems with cognitive capabilities superior to humans could outcompete or dominate human decision-making, leading to conflicts over resources and control."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  47. 61.02.24 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Evolutionary dynamics

    "AI models and systems may develop their own motivations, leading to unpredictable behaviors."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  48. 61.02.30 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Indifference to human values

    "AI models and systems may develop goals or behaviors that are misaligned with human values."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  49. 61.02.36 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Model design enabling power-seeking

    "Some AI models and systems might develop tendencies to seek power or control."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.