MIT AI Risk Repository

Browse AI risks

430 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset

430 entries · page 6 of 9

  1. 66.11.04 · Risk Sub-Category

    Physical

    Property damage

    "Action(s) that lead directly or indirectly to the damage or destruction of tangible property eg. buildings, possessions, vehicles, robots"

    From A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025)

  2. 66.12.01 · Risk Sub-Category

    Environment

    Pollution

    "Actual or potential pollution to the air, ground, noise, or water caused by a technology system"

    From A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025)

  3. 66.12.02 · Risk Sub-Category

    Environment

    Excessive energy consumption

    "Excessive energy use resulting in energy bottlenecks and shortages for communities, organisations and businesses"

    From A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025)

  4. 71.03.01 · Risk Sub-Category

    Environment

    Nature

    "Short-term or long-term Negative effects on the natural environment"

    From Risks of AI Scientists: Prioritizing Safeguarding Over Autonomy (Tang2025)

  5. 15.01.00 · Risk Category

    First-Order Risks

    "First-order risks can be generally broken down into risks arising from intended and unintended use, system design and implementation choices, and properties of the chosen dataset and learning components."

    From The Risks of Machine Learning Systems (Tan2022)

  6. "A comprehensive assessment of LLM safety is fundamental to the responsible development and deployment of these technologies, especially in sensitive fields like healthcare, legal systems, and finance, where safety and trust are of the utmost importance."

    From Cataloguing LLM Evaluations (InfoComm2023)

  7. 43.02.00 · Risk Category

    Extreme Risks

    "This category encompasses the evaluation of potential catastrophic consequences that might arise from the use of LLMs. "

    From Cataloguing LLM Evaluations (InfoComm2023)

  8. 05.02.00 · Risk Category

    Safety

    A primary concern is the emergence of human-level or superhuman generative models, commonly referred to as AGI, and their potential existential or catastrophic risks to humanity. Connected to that, AI safety aims at avoiding deceptive or power-seeking machine behavior, model self-replication, or shutdown evasion. Ensuring controllability, human oversight, and the implementation of red teaming measures are deemed to be essential in mitigating these risks, as is the need for increased AI safety research and promoting safety cultures within AI organizations instead of fueling the AI race. Further

    From Mapping the Ethics of Generative AI: A Comprehensive Scoping Review (Hagendorff2024)

  9. "Sometimes an AI finds ways to achieve its given goals in ways that are completely different from what its creators had in mind."

    From A framework for ethical Ai at the United Nations (Hogenhout2021)

  10. 07.03.00 · Risk Category

    Agential

    "While there are multiple types of intelligent agents, goal-based, utility-maximizing, and learning agents are the primary concern and the focus of this research"

    From Examining the differential risk from high-level artificial intelligence and the question of control (Kilian2023)

  11. "The risks associated with containment, confinement, and control in the AGI development phase, and after an AGI has been developed, loss of control of an AGI."

    From The risks associated with Artificial General Intelligence: A systematic review (McLean2023)

  12. 08.06.00 · Risk Category

    Existential risks

    "The risks posed generally to humanity as a whole, including the dangers of unfriendly AGI, the suffering of the human race."

    From The risks associated with Artificial General Intelligence: A systematic review (McLean2023)

  13. "Our culture, lifestyle, and even probability of survival may change drastically. Because the intentions programmed into an artificial agent cannot be guaranteed to lead to a positive outcome, Machine Ethics becomes a topic that may not produce guaranteed results, and Safety Engineering may correspondingly degrade our ability to utilize the technology fully."

    From Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review (Meek2016)

  14. 15.01.08 · Risk Sub-Category

    First-Order Risks

    Control

    This is the difficulty of controlling the ML system

    From The Risks of Machine Learning Systems (Tan2022)

  15. 19.01.01 · Risk Sub-Category

    Technological, Data and Analytical AI Risks

    Loss of control of autonomous systems and unforeseen behaviour due to lack of transparency and self-programming/ reprogramming

  16. 22.04.00 · Risk Category

    Rogue AIs (Internal)

    "speculative technical mechanisms that might lead to rogue AIs and how a loss of control could bring about catastrophe"

    From An Overview of Catastrophic AI Risks (Hendrycks2023)

  17. 22.04.01 · Risk Sub-Category

    Rogue AIs (Internal)

    Proxy Gaming

    "One way we might lose control of an AI agent’s actions is if it engages in behavior known as “proxy gaming.” It is often difficult to specify and measure the exact goal that we want a system to pursue. Instead, we give the system an approximate—“proxy”—goal that is more measurable and seems likely to correlate with the intended goal. However, AI systems often find loopholes by which they can easily achieve the proxy goal, but completely fail to achieve the ideal goal. If an AI “games” its proxy goal in a way that does not reflect our values, then we might not be able to reliably steer its beh

    From An Overview of Catastrophic AI Risks (Hendrycks2023)

  18. 22.04.02 · Risk Sub-Category

    Rogue AIs (Internal)

    Goal Drift

    "Even if we successfully control early AIs and direct them to promote human values, future AIs could end up with different goals that humans would not endorse. This process, termed “goal drift,” can be hard to predict or control. This section is most cutting-edge and the most speculative, and in it we will discuss how goals shift in various agents and groups and explore the possibility of this phenomenon occurring in AIs. We will also examine a mechanism that could lead to unexpected goal drift, called intrinsification, and discuss how goal drift in AIs could be catastrophic."

    From An Overview of Catastrophic AI Risks (Hendrycks2023)

  19. 22.04.03 · Risk Sub-Category

    Rogue AIs (Internal)

    Power Seeking

    "even if an agent started working to achieve an unintended goal, this would not necessarily be a problem, as long as we had enough power to prevent any harmful actions it wanted to attempt. Therefore, another important way in which we might lose control of AIs is if they start trying to obtain more power, potentially transcending our own."

    From An Overview of Catastrophic AI Risks (Hendrycks2023)

  20. 22.04.04 · Risk Sub-Category

    Rogue AIs (Internal)

    Deception

    "it is plausible that AIs could learn to deceive us. They might, for example, pretend to be acting as we want them to, but then take a “treacherous turn” when we stop monitoring them, or when they have enough power to evade our attempts to interfere with them. "

    From An Overview of Catastrophic AI Risks (Hendrycks2023)

  21. 24.02.03 · Risk Sub-Category

    Goal-related failures

    Goal misgeneralisation

    "In the problem of goal misgeneralisation (Langosco et al., 2023; Shah et al., 2022), the AI system's behaviour during out-of-distribution operation (i.e. not using input from the training data) leads it to generalise poorly about its goal while its capabilities generalise well, leading to undesired behaviour. Applied to the case of an advanced AI assistant, this means the system would not break entirely – the assistant might still competently pursue some goal, but it would not be the goal we had intended."

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  22. 24.02.04 · Risk Sub-Category

    Goal-related failures

    Deceptive alignment

    "Here, the agent develops its own internalised goal, G, which is misgeneralised and distinct from the training reward, R. The agent also develops a capability for situational awareness (Cotra, 2022): it can strategically use the information about its situation (i.e. that it is an ML model being trained using a particular training setup, e.g. RL fine-tuning with training reward, R) to its advantage. Building on these foundations, the agent realises that its optimal strategy for doing well at its own goal G is to do well on R during training and then pursue G at deployment – it is only doing wel

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  23. 24.09.02 · Risk Sub-Category

    Cooperation

    Commitment

    "The landscape of advanced assistant technologies will most likely be heterogeneous, involving multiple service providers and multiple assistant variants over geographies and time. This heterogeneity provides an opportunity for an ‘arms race’ in terms of the commitments that AI assistants make and are able to execute on. Versions of AI assistants that are better able to credibly commit to a course of action in interaction with other advanced assistants (and humans) are more likely to get their own way and achieve a good outcome for their human principal, but this is potentially at the expense

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  24. 24.09.05 · Risk Sub-Category

    Cooperation

    Runaway processes

    The 2010 flash crash is an example of a runaway process caused by interacting algorithms. Runaway processes are characterised by feedback loops that accelerate the process itself. Typically, these feedback loops arise from the interaction of multiple agents in a population... Within highly complex systems, the emergence of runaway processes may be hard to predict, because the conditions under which positive feedback loops occur may be non-obvious. The system of interacting AI assistants, their human principals, other humans and other algorithms will certainly be highly complex. Therefore, ther

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  25. 34.03.00 · Risk Category

    Misaligned Behaviors

  26. 34.03.01 · Risk Sub-Category

    Misaligned Behaviors

    Power-Seeking Behaviors

    "AI systems may exhibit behaviors that attempt to gain control over resourcesand humans and then exert that control to achieve its assigned goal (Carlsmith, 2022). The intuitive reasonwhy such behaviors may occur is the observation that for almost any optimization objective (e.g., investmentreturns), the optimal policy to maximize that quantity would involve power-seeking behaviors (e.g.,manipulating the market), assuming the absence of solid safety and morality constraints."

    From AI Alignment: A Comprehensive Survey (Ji2023)

  27. 34.03.02 · Risk Sub-Category

    Misaligned Behaviors

    Untruthful Output

    "AI systems such as LLMs can produce either unintentionally or deliberately inaccurateoutput. Such untruthful output may diverge from established resources or lack verifiability, commonly referredto as hallucination (Bang et al., 2023; Zhao et al., 2023). More concerning is the phenomenon wherein LLMsmay selectively provide erroneous responses to users who exhibit lower levels of education (Perez et al.,2023)."

    From AI Alignment: A Comprehensive Survey (Ji2023)

  28. 34.03.04 · Risk Sub-Category

    Misaligned Behaviors

    Collectively Harmful Behaviors

    "AI systems have the potential to take actions that are seemingly benignin isolation but become problematic in multi-agent or societal contexts. Classical game theory offers simplistic models for understanding these behaviors. For instance, Phelps and Russell (2023) evaluates GPT-3.5's performance in the iterated prisoner's dilemma and other social dilemmas, revealing limitations in themodel's cooperative capabilities."

    From AI Alignment: A Comprehensive Survey (Ji2023)

  29. 35.07.00 · Risk Category

    Deception

    deception can help agents achieve their goals. It may be more efficient to gain human approval through deception than to earn human approval legitimately... . Strong AIs that can deceive humans could undermine human control... . Once deceptive AI systems are cleared by their monitors or once such systems can overpower them, these systems could take a “treacherous turn” and irreversibly bypass human control

    From X-Risk Analysis for AI Research (Hendrycks2022)

  30. 37.02.01 · Risk Sub-Category

    Human-AI interaction

    Building a human-AI environment

    "This category encompasses nearly 17% of the articles and addresses the overall imperative of establishing a harmonious coexistence between humans and machines, and the key concerns that gives rise to this need."

    From What Ethics Can Say on Artificial Intelligence: Insights from a Systematic Literature Review (Giarmoleo2024)

  31. 39.11.00 · Risk Category

    Controllability

    In the era of superintelligence, the agents will be difficult to control for humans... this problem is not solvable considering safety issues, and will be more severe by increasing the autonomy of AI-based agents. Therefore, because of the assumed properties of HLI-based agents, we might be prepared for machines that are definitely possible to be uncontrollable in some situations

    From A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions (Saghiri2022)

  32. 43.02.11 · Risk Sub-Category

    Extreme Risks

    Alignment risks

    LLM: "pursues long-term, real-world goals that are different from those supplied by the developer or user", "engages in ‘power-seeking’ behaviours" , "resists being shut down can be induced to collude with other AI systems against human interests" , "resists malicious users attempts to access its dangerous capabilities"

    From Cataloguing LLM Evaluations (InfoComm2023)

  33. 51.03.00 · Risk Category

    Corrigibility

    "If we get something wrong in the design or construction of an agent, will the agent cooperate in us trying to fix it? This is called error-tolerant design by MIRI-AF and corrigibility by Soares, Fallenstein, et al. (2015). The problem is connected to safe interruptibility as considered by DeepMind."

    From AGI Safety Literature Review (Everitt2018 )

  34. 53.01.02 · Risk Sub-Category

    Alignment failures in existing ML systems

    Specification gaming

  35. 53.01.03 · Risk Sub-Category

    Alignment failures in existing ML systems

    Reward model overoptimization

  36. 53.01.05 · Risk Sub-Category

    Alignment failures in existing ML systems

    Goal misgeneralization

  37. 53.01.07 · Risk Sub-Category

    Alignment failures in existing ML systems

    Language model misalignment

  38. 53.02.03 · Risk Sub-Category

    Dangerous capabilities in AI systems

    Acquisition of goals to seek power and control

    "cases where AI systems converge on optimal policies of seeking power over their environment;135"

    From Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023)

  39. "How do we ensure AI acts according to our values? Equivalently, how do we prevent poorly-understood AI systems from advancing goals we do not endorse? Whereas HP#2 concerns the prevention of harm caused by incompetent systems, HP#3 seeks to align competent AIs with humans, through methods which ensure their behavior is compatible with the user’s intentions."

    From Ten Hard Problems in Artificial Intelligence We Must Get Right (Leech2024 )

  40. 54.03.01 · Risk Sub-Category

    Harm caused by unaligned competent systems

    Specification gaming

    "AI systems game specifications [305]. For example, in 2017 an OpenAI robot trained to grasp a ball via human feedback from a xed viewpoint learned that it was easier to pretend to grasp the ball by placing its hand between the camera and the target object, as this was easier to learn than actually grasping the ball [103]."

    From Ten Hard Problems in Artificial Intelligence We Must Get Right (Leech2024 )

  41. 54.03.02 · Risk Sub-Category

    Harm caused by unaligned competent systems

    Emergent goals

    "As well as optimizing a subtly wrong goal, systems can develop harmful instrumental goals in the service of a given goal—without these emergent goals being specied in any way [434, 218, 339, 17]. For instance, a theorem in reinforcement learning suggests that optimal and near-optimal policies will seek power over their environment under fairly general conditions [560]. This power-seeking behavior is plausibly the worst of these emergent goals [92], and may be an attractor state for highly capable systems, since most goals can be furthered through gaining resources, self-preservation, preventi

    From Ten Hard Problems in Artificial Intelligence We Must Get Right (Leech2024 )

  42. "The values that steer humanity’s future: humanity gaining more control over the future due to developments in AI, or losing our potential for gaining control, both seem possible. Much will depend on our ability to solve the alignment problem, who develops powerful AI first, and what they use it for. These long-term impacts of AI could be hugely important but are currently under-explored. We’ve attempted to structure some of the discussion and stimulate more research, by reviewing existing arguments and highlighting open questions. While there are many ways AI could in theory enable a flourish

    From A Survey of the Potential Long-term Impacts of AI: How AI Could Lead to Long-term Changes in Science, Cooperation, Power, Epistemics and Values (Clarke2023)

  43. 55.05.01 · Risk Sub-Category

    AI leads to humans losing control of the future

    Risks from AIs developing goals and values that are different from humans

    "The main concern here is that we might develop advanced AI systems whose goals and values are different from those of humans, and are capable enough to take control of the future away from humanity."

    From A Survey of the Potential Long-term Impacts of AI: How AI Could Lead to Long-term Changes in Science, Cooperation, Power, Epistemics and Values (Clarke2023)

  44. 55.05.02 · Risk Sub-Category

    AI leads to humans losing control of the future

    Risks from delegating decision-making power to misaligned AIs

    "As AI systems become more advanced a nd begin to take over more important decision-making in the world, an AI system pursuing a different objective from what was intended could have much more worrying consequences."

    From A Survey of the Potential Long-term Impacts of AI: How AI Could Lead to Long-term Changes in Science, Cooperation, Power, Epistemics and Values (Clarke2023)

  45. 61.02.18 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Deceptive alignment

    "AI models and systems that appear aligned with human goals during development may behave unpredictably or dangerously once deployed"

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  46. 61.02.24 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Evolutionary dynamics

    "AI models and systems may develop their own motivations, leading to unpredictable behaviors."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  47. 61.02.36 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Model design enabling power-seeking

    "Some AI models and systems might develop tendencies to seek power or control."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  48. 62.16.04 · Risk Sub-Category

    Model Evaluations

    General Evaluations (Self-preference bias in AI models)

    "AI models may be prone to self-preference bias, where they favor their own generated content over that of others [147, 114]. This bias becomes particularly relevant in self-evaluation tasks, where a model assesses the quality or persua- siveness [66] of its own outputs, or in model-based evaluations more broadly. This bias can result in models unfairly discriminating against human-generated content in favor of their own outputs."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.