MIT AI Risk Repository

Browse AI risks

100 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset Also filtered by subdomain 7.1 ×

100 entries · page 1 of 2

  1. 05.02.00 · Risk Category

    Safety

    A primary concern is the emergence of human-level or superhuman generative models, commonly referred to as AGI, and their potential existential or catastrophic risks to humanity. Connected to that, AI safety aims at avoiding deceptive or power-seeking machine behavior, model self-replication, or shutdown evasion. Ensuring controllability, human oversight, and the implementation of red teaming measures are deemed to be essential in mitigating these risks, as is the need for increased AI safety research and promoting safety cultures within AI organizations instead of fueling the AI race. Further

    From Mapping the Ethics of Generative AI: A Comprehensive Scoping Review (Hagendorff2024)

  2. 05.09.00 · Risk Category

    Alignment

    The general tenet of AI alignment involves training generative AI systems to be harmless, helpful, and honest, ensuring their behavior aligns with and respects human values. However, a central debate in this area concerns the methodological challenges in selecting appropriate values. While AI systems can acquire human values through feedback, observation, or debate, there remains ambiguity over which individuals are qualified or legitimized to provide these guiding signals. Another prominent issue pertains to deceptive alignment, which might cause generative AI systems to tamper evaluations. A

    From Mapping the Ethics of Generative AI: A Comprehensive Scoping Review (Hagendorff2024)

  3. "Sometimes an AI finds ways to achieve its given goals in ways that are completely different from what its creators had in mind."

    From A framework for ethical Ai at the United Nations (Hogenhout2021)

  4. 07.03.00 · Risk Category

    Agential

    "While there are multiple types of intelligent agents, goal-based, utility-maximizing, and learning agents are the primary concern and the focus of this research"

    From Examining the differential risk from high-level artificial intelligence and the question of control (Kilian2023)

  5. "The risks associated with containment, confinement, and control in the AGI development phase, and after an AGI has been developed, loss of control of an AGI."

    From The risks associated with Artificial General Intelligence: A systematic review (McLean2023)

  6. "The risks associated with AGI goal safety, including human attempts at making goals safe, as well as the AGI making its own goals safe during self-improvement."

    From The risks associated with Artificial General Intelligence: A systematic review (McLean2023)

  7. 08.06.00 · Risk Category

    Existential risks

    "The risks posed generally to humanity as a whole, including the dangers of unfriendly AGI, the suffering of the human race."

    From The risks associated with Artificial General Intelligence: A systematic review (McLean2023)

  8. "A sufficiently intelligent AI could possess the ability to subtly influence societal behaviors through a sophisticated understanding of human nature"

    From Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review (Meek2016)

  9. "Our culture, lifestyle, and even probability of survival may change drastically. Because the intentions programmed into an artificial agent cannot be guaranteed to lead to a positive outcome, Machine Ethics becomes a topic that may not produce guaranteed results, and Safety Engineering may correspondingly degrade our ability to utilize the technology fully."

    From Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review (Meek2016)

  10. "The speculative potential for future advanced AI systems to harm human civilization, either through misuse or due to challenges in aligning AI objectives with human values."

    From AI Risk Profiles: A Standards Proposal for Pre-Deployment AI Risk Disclosures (Sherman2023)

  11. "The degree of automation and control describes the extent to which an AI system functions independently of human supervision and control."

    From Sources of Risk of AI Systems (Steimers2022)

  12. 15.01.08 · Risk Sub-Category

    First-Order Risks

    Control

    This is the difficulty of controlling the ML system

    From The Risks of Machine Learning Systems (Tan2022)

  13. 15.01.09 · Risk Sub-Category

    First-Order Risks

    Emergent behavior

    "This is the risk resulting from novel behavior acquired through continual learning or self-organization after deployment."

    From The Risks of Machine Learning Systems (Tan2022)

  14. "AI systems compromising human agency, or circumventing meaningful human control"

    From Sociotechnical Safety Evaluation of Generative AI Systems (Weidinger2023)

  15. 18.05.02 · Risk Sub-Category

    Human Autonomy and Intregrity Harms

    Persuasion and manipulation

    "Exploiting user trust, or nudging or coercing them into performing certain actions against their will (c.f. Burtell and Woodside (2023); Kenton et al. (2021))"

    From Sociotechnical Safety Evaluation of Generative AI Systems (Weidinger2023)

  16. 19.01.01 · Risk Sub-Category

    Technological, Data and Analytical AI Risks

    Loss of control of autonomous systems and unforeseen behaviour due to lack of transparency and self-programming/ reprogramming

  17. 22.04.00 · Risk Category

    Rogue AIs (Internal)

    "speculative technical mechanisms that might lead to rogue AIs and how a loss of control could bring about catastrophe"

    From An Overview of Catastrophic AI Risks (Hendrycks2023)

  18. 22.04.01 · Risk Sub-Category

    Rogue AIs (Internal)

    Proxy Gaming

    "One way we might lose control of an AI agent’s actions is if it engages in behavior known as “proxy gaming.” It is often difficult to specify and measure the exact goal that we want a system to pursue. Instead, we give the system an approximate—“proxy”—goal that is more measurable and seems likely to correlate with the intended goal. However, AI systems often find loopholes by which they can easily achieve the proxy goal, but completely fail to achieve the ideal goal. If an AI “games” its proxy goal in a way that does not reflect our values, then we might not be able to reliably steer its beh

    From An Overview of Catastrophic AI Risks (Hendrycks2023)

  19. 22.04.02 · Risk Sub-Category

    Rogue AIs (Internal)

    Goal Drift

    "Even if we successfully control early AIs and direct them to promote human values, future AIs could end up with different goals that humans would not endorse. This process, termed “goal drift,” can be hard to predict or control. This section is most cutting-edge and the most speculative, and in it we will discuss how goals shift in various agents and groups and explore the possibility of this phenomenon occurring in AIs. We will also examine a mechanism that could lead to unexpected goal drift, called intrinsification, and discuss how goal drift in AIs could be catastrophic."

    From An Overview of Catastrophic AI Risks (Hendrycks2023)

  20. 22.04.03 · Risk Sub-Category

    Rogue AIs (Internal)

    Power Seeking

    "even if an agent started working to achieve an unintended goal, this would not necessarily be a problem, as long as we had enough power to prevent any harmful actions it wanted to attempt. Therefore, another important way in which we might lose control of AIs is if they start trying to obtain more power, potentially transcending our own."

    From An Overview of Catastrophic AI Risks (Hendrycks2023)

  21. 22.04.04 · Risk Sub-Category

    Rogue AIs (Internal)

    Deception

    "it is plausible that AIs could learn to deceive us. They might, for example, pretend to be acting as we want them to, but then take a “treacherous turn” when we stop monitoring them, or when they have enough power to evade our attempts to interfere with them. "

    From An Overview of Catastrophic AI Risks (Hendrycks2023)

  22. 24.02.00 · Risk Category

    Goal-related failures

    "As we think about even more intelligent and advanced AI assistants, perhaps outperforming humans on many cognitive tasks, the question of how humans can successfully control such an assistant looms large. To achieve the goals we set for an assistant, it is possible (Shah, 2022) that the AI assistant will implement some form of consequentialist reasoning: considering many different plans, predicting their consequences and executing the plan that does best according to some metric, M. This kind of reasoning can arise because it is a broadly useful capability (e.g. planning ahead, considering mo

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  23. 24.02.02 · Risk Sub-Category

    Goal-related failures

    Specification gaming

    "Specification gaming (Krakovna et al., 2020) occurs when some faulty feedback is provided to the assistant in the training data (i.e. the training objective O does not fully capture what the user/designer wants the assistant to do). It is typified by the sort of behaviour that exploits loopholes in the task specification to satisfy the literal specification of a goal without achieving the intended outcome."

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  24. 24.02.03 · Risk Sub-Category

    Goal-related failures

    Goal misgeneralisation

    "In the problem of goal misgeneralisation (Langosco et al., 2023; Shah et al., 2022), the AI system's behaviour during out-of-distribution operation (i.e. not using input from the training data) leads it to generalise poorly about its goal while its capabilities generalise well, leading to undesired behaviour. Applied to the case of an advanced AI assistant, this means the system would not break entirely – the assistant might still competently pursue some goal, but it would not be the goal we had intended."

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  25. 24.02.04 · Risk Sub-Category

    Goal-related failures

    Deceptive alignment

    "Here, the agent develops its own internalised goal, G, which is misgeneralised and distinct from the training reward, R. The agent also develops a capability for situational awareness (Cotra, 2022): it can strategically use the information about its situation (i.e. that it is an ML model being trained using a particular training setup, e.g. RL fine-tuning with training reward, R) to its advantage. Building on these foundations, the agent realises that its optimal strategy for doing well at its own goal G is to do well on R during training and then pursue G at deployment – it is only doing wel

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  26. 24.09.00 · Risk Category

    Cooperation

    "" AI assistants will need to coordinate with other AI assistants and with humans other than their principal users. This chapter explores the societal risks associated with the aggregate impact of AI assistants whose behaviour is aligned to the interests of particular users. For example, AI assistants may face collective action problems where the best outcomes overall are realised when AI assistants cooperate but where each AI assistant can secure an additional benefit for its user if it defects while others cooperate""

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  27. 24.09.02 · Risk Sub-Category

    Cooperation

    Commitment

    "The landscape of advanced assistant technologies will most likely be heterogeneous, involving multiple service providers and multiple assistant variants over geographies and time. This heterogeneity provides an opportunity for an ‘arms race’ in terms of the commitments that AI assistants make and are able to execute on. Versions of AI assistants that are better able to credibly commit to a course of action in interaction with other advanced assistants (and humans) are more likely to get their own way and achieve a good outcome for their human principal, but this is potentially at the expense

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  28. 24.09.05 · Risk Sub-Category

    Cooperation

    Runaway processes

    The 2010 flash crash is an example of a runaway process caused by interacting algorithms. Runaway processes are characterised by feedback loops that accelerate the process itself. Typically, these feedback loops arise from the interaction of multiple agents in a population... Within highly complex systems, the emergence of runaway processes may be hard to predict, because the conditions under which positive feedback loops occur may be non-obvious. The system of interacting AI assistants, their human principals, other humans and other algorithms will certainly be highly complex. Therefore, ther

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  29. 34.01.00 · Risk Category

    Causes of Misalignment

    we aim to further analyze why and how the misalignment issues occur. We will first give an overview of common failure modes, and then focus on the mechanism of feedback-induced misalignment, and finally shift our emphasis towards an examination of misaligned behaviors and dangerous capabilities

    From AI Alignment: A Comprehensive Survey (Ji2023)

  30. 34.01.01 · Risk Sub-Category

    Causes of Misalignment

    Reward Hacking

    "Reward Hacking: In practice, proxy rewards are often easy to optimize and measure, yet they frequently fall shortof capturing the full spectrum of the actual rewards (Pan et al., 2021). This limitation is denoted as misspecifiedrewards. The pursuit of optimization based on such misspecified rewards may lead to a phenomenon knownas reward hacking, wherein agents may appear highly proficient according to specific metrics but fall short whenevaluated against human standards (Amodei et al., 2016; Everitt et al., 2017). The discrepancy between proxyrewards and true rewards often manifests as a sha

    From AI Alignment: A Comprehensive Survey (Ji2023)

  31. 34.01.02 · Risk Sub-Category

    Causes of Misalignment

    Goal Misgeneralization

    "Goal Misgeneralization: Goal misgeneralization is another failure mode, wherein the agent actively pursuesobjectives distinct from the training objectives in deployment while retaining the capabilities it acquired duringtraining (Di Langosco et al., 2022). For instance, in CoinRun games, the agent frequently prefers reachingthe end of a level, often neglecting relocated coins during testing scenarios. Di Langosco et al. (2022) drawattention to the fundamental disparity between capability generalization and goal generalization, emphasizing howthe inductive biases inherent in the model and its

    From AI Alignment: A Comprehensive Survey (Ji2023)

  32. 34.01.03 · Risk Sub-Category

    Causes of Misalignment

    Reward Tampering

    "Reward tampering can be considered a special case of reward hacking (Everitt et al., 2021; Skalse et al., 2022),referring to AI systems corrupting the reward signals generation process (Ring and Orseau, 2011). Everitt et al.(2021) delves into the subproblems encountered by RL agents: (1) tampering of reward function, where the agentinappropriately interferes with the reward function itself, and (2) tampering of reward function input, which entailscorruption within the process responsible for translating environmental states into inputs for the reward function.When the reward function is formu

    From AI Alignment: A Comprehensive Survey (Ji2023)

  33. 34.01.05 · Risk Sub-Category

    Causes of Misalignment

    Limitations of Reward Modeling

    "Limitations of Reward Modeling. Training reward models using comparison feedback can pose significantchallenges in accurately capturing human values. For example, these models may unconsciously learn suboptimal or incomplete objectives, resulting in reward hacking (Zhuang and Hadfield-Menell, 2020; Skalse et al.,2022). Meanwhile, using a single reward model may struggle to capture and specify the values of a diversehuman society (Casper et al., 2023b)."

    From AI Alignment: A Comprehensive Survey (Ji2023)

  34. 34.03.00 · Risk Category

    Misaligned Behaviors

  35. 34.03.01 · Risk Sub-Category

    Misaligned Behaviors

    Power-Seeking Behaviors

    "AI systems may exhibit behaviors that attempt to gain control over resourcesand humans and then exert that control to achieve its assigned goal (Carlsmith, 2022). The intuitive reasonwhy such behaviors may occur is the observation that for almost any optimization objective (e.g., investmentreturns), the optimal policy to maximize that quantity would involve power-seeking behaviors (e.g.,manipulating the market), assuming the absence of solid safety and morality constraints."

    From AI Alignment: A Comprehensive Survey (Ji2023)

  36. 34.03.02 · Risk Sub-Category

    Misaligned Behaviors

    Untruthful Output

    "AI systems such as LLMs can produce either unintentionally or deliberately inaccurateoutput. Such untruthful output may diverge from established resources or lack verifiability, commonly referredto as hallucination (Bang et al., 2023; Zhao et al., 2023). More concerning is the phenomenon wherein LLMsmay selectively provide erroneous responses to users who exhibit lower levels of education (Perez et al.,2023)."

    From AI Alignment: A Comprehensive Survey (Ji2023)

  37. 34.03.03 · Risk Sub-Category

    Misaligned Behaviors

    Deceptive Alignment & Manipulation

    "Manipulation & Deceptive Alignment is a class of behaviors thatexploit the incompetence of human evaluators or users (Hubinger et al., 2019a; Carranza et al., 2023) andeven manipulate the training process through gradient hacking (Richard Ngo, 2022). These behaviors canpotentially make detecting and addressing misaligned behaviors much harder.Deceptive Alignment: Misaligned AI systems may deliberately mislead their human supervisors instead of adhering to the intended task. Such deceptive behavior has already manifested in AI systems that employ evolutionary algorithms (Wilke et al., 2001; He

    From AI Alignment: A Comprehensive Survey (Ji2023)

  38. 34.03.04 · Risk Sub-Category

    Misaligned Behaviors

    Collectively Harmful Behaviors

    "AI systems have the potential to take actions that are seemingly benignin isolation but become problematic in multi-agent or societal contexts. Classical game theory offers simplistic models for understanding these behaviors. For instance, Phelps and Russell (2023) evaluates GPT-3.5's performance in the iterated prisoner's dilemma and other social dilemmas, revealing limitations in themodel's cooperative capabilities."

    From AI Alignment: A Comprehensive Survey (Ji2023)

  39. 35.04.00 · Risk Category

    Proxy misspecification

    AI agents are directed by goals and objectives. Creating general-purpose objectives that capture human values could be challenging... Since goal-directed AI systems need measurable objectives, by default our systems may pursue simplified proxies of human values. The result could be suboptimal or even catastrophic if a sufficiently powerful AI successfully optimizes its flawed objective to an extreme degree

    From X-Risk Analysis for AI Research (Hendrycks2022)

  40. 35.07.00 · Risk Category

    Deception

    deception can help agents achieve their goals. It may be more efficient to gain human approval through deception than to earn human approval legitimately... . Strong AIs that can deceive humans could undermine human control... . Once deceptive AI systems are cleared by their monitors or once such systems can overpower them, these systems could take a “treacherous turn” and irreversibly bypass human control

    From X-Risk Analysis for AI Research (Hendrycks2022)

  41. 35.08.00 · Risk Category

    Power-seeking behavior

    Agents that have more power are better able to accomplish their goals. Therefore, it has been shown that agents have incentives to acquire and maintain power. AIs that acquire substantial power can become especially dangerous if they are not aligned with human values

    From X-Risk Analysis for AI Research (Hendrycks2022)

  42. 37.02.01 · Risk Sub-Category

    Human-AI interaction

    Building a human-AI environment

    "This category encompasses nearly 17% of the articles and addresses the overall imperative of establishing a harmonious coexistence between humans and machines, and the key concerns that gives rise to this need."

    From What Ethics Can Say on Artificial Intelligence: Insights from a Systematic Literature Review (Giarmoleo2024)

  43. 39.11.00 · Risk Category

    Controllability

    In the era of superintelligence, the agents will be difficult to control for humans... this problem is not solvable considering safety issues, and will be more severe by increasing the autonomy of AI-based agents. Therefore, because of the assumed properties of HLI-based agents, we might be prepared for machines that are definitely possible to be uncontrollable in some situations

    From A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions (Saghiri2022)

  44. 39.26.00 · Risk Category

    Safety

    The actions of a learning model may easily hurt humans in both explicit and implicit manners...several algorithms based on Asimov’s laws have been proposed that try to judge the output actions of an agent considering the safety of humans

    From A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions (Saghiri2022)

  45. "Probably the most talked about source of potential problems with future AIs is mistakes in design. Mainly the concern is with creating a "wrong AI", a system which doesn't match our original desired formal properties or has unwanted behaviors (Dewey, Russell et al. 2015, Russell, Dewey et al. January 23, 2015), such as drives for independence or dominance. Mistakes could also be simple bugs (run time or logical) in the source code, disproportionate weights in the fitness function, or goals misaligned with human values leading to complete disregard for human safety."

    From Taxonomy of Pathways to Dangerous Artificial Intelligence (Yampolskiy2016)

  46. 42.17.00 · Risk Category

    Diluting Rights

    "A possible consequence of self-interest in AI generation of ethical guidelines."

    From An Exploratory Diagnosis of Artificial Intelligence Risks for a Responsible Governance (Teixeira2022)

  47. 43.02.11 · Risk Sub-Category

    Extreme Risks

    Alignment risks

    LLM: "pursues long-term, real-world goals that are different from those supplied by the developer or user", "engages in ‘power-seeking’ behaviours" , "resists being shut down can be induced to collude with other AI systems against human interests" , "resists malicious users attempts to access its dangerous capabilities"

    From Cataloguing LLM Evaluations (InfoComm2023)

  48. 45.02.13 · Risk Sub-Category

    Safety risks in AI Applications

    Ethical Risks (Risks of AI becoming uncontrollable in the future)

    "With the fast development of AI technologies, there is a risk of AI autonomously acquiring external resources, conducting self-replication, become self-aware, seeking for external power, and attempting to seize control from humans."

    From AI Safety Governance Framework (TC2602024)

  49. 47.01.03 · Risk Sub-Category

    Technical and operational risks

    Technical vulnerabilities (The risk of misalignment)

    "To assess whether an AI model is reliable or robust, it is crucial to consider whether the model is “aligned.” “Alignment” focuses on whether an AI model effectively operates in accordance with the goals established by its designers.238 A misaligned AI model may pursue some objectives, but not the intended ones. Therefore, misaligned AI models can malfunction and cause harm."

    From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024)

  50. 49.02.03 · Risk Sub-Category

    Risks from Malfunctions

    Loss of control

    "'Loss of control’ scenarios are potential future scenarios in which society can no longer meaningfully constrain some advanced general- purpose AI agents, even if it becomes clear they are causing harm. These scenarios are hypothesised to arise through a combination of social and technical factors, such as pressures to delegate decisions to general- purpose AI systems, and limitations of existing techniques used to influence the behaviours of general- purpose AI systems."

    From International Scientific Report on the Safety of Advanced AI (Bengio2024)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.