MIT AI Risk Repository

Browse AI risks

977 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset

977 entries · page 16 of 20

  1. 15.01.01 · Risk Sub-Category

    First-Order Risks

    Application

    "This is the risk posed by the intended application or use case. It is intuitive that some use cases will be inherently "riskier" than others (e.g., an autonomous weapons system vs. a customer service chatbot)."

    From The Risks of Machine Learning Systems (Tan2022)

  2. "While highly rare, it is known, that occasionally individual bits may be flipped in different hardware devices due to manufacturing defects or cosmic rays hitting just the right spot (Simonite March 7, 2008). This is similar to mutations observed in living organisms and may result in a modification of an intelligent system."

    From Taxonomy of Pathways to Dangerous Artificial Intelligence (Yampolskiy2016)

  3. "Previous research has shown that utility maximizing agents are likely to fall victims to the same indulgences we frequently observe in people, such as addictions, pleasure drives (Majot and Yampolskiy 2014), self-delusions and wireheading (Yampolskiy 2014). In general, what we call mental illness in people, particularly sociopathy as demonstrated by lack of concern for others, is also likely to show up in artificial minds."

    From Taxonomy of Pathways to Dangerous Artificial Intelligence (Yampolskiy2016)

  4. 62.15.09 · Risk Sub-Category

    Model Development

    Fine-tuning related (Degrading safety training due to benign fine-tuning)

    "When downstream providers of AI systems fine-tune AI models to be more suitable for their needs, the resulting AI model can be more likely to produce undesired or harmful outputs (as compared to the non-fine-tuned model), even if the fine-tuning was done with harmless and commonly used data [154]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  5. "A sufficiently intelligent AI could possess the ability to subtly influence societal behaviors through a sophisticated understanding of human nature"

    From Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review (Meek2016)

  6. "The speculative potential for future advanced AI systems to harm human civilization, either through misuse or due to challenges in aligning AI objectives with human values."

    From AI Risk Profiles: A Standards Proposal for Pre-Deployment AI Risk Disclosures (Sherman2023)

  7. "The degree of automation and control describes the extent to which an AI system functions independently of human supervision and control."

    From Sources of Risk of AI Systems (Steimers2022)

  8. 15.01.09 · Risk Sub-Category

    First-Order Risks

    Emergent behavior

    "This is the risk resulting from novel behavior acquired through continual learning or self-organization after deployment."

    From The Risks of Machine Learning Systems (Tan2022)

  9. "AI systems compromising human agency, or circumventing meaningful human control"

    From Sociotechnical Safety Evaluation of Generative AI Systems (Weidinger2023)

  10. 18.05.02 · Risk Sub-Category

    Human Autonomy and Intregrity Harms

    Persuasion and manipulation

    "Exploiting user trust, or nudging or coercing them into performing certain actions against their will (c.f. Burtell and Woodside (2023); Kenton et al. (2021))"

    From Sociotechnical Safety Evaluation of Generative AI Systems (Weidinger2023)

  11. 24.09.00 · Risk Category

    Cooperation

    "" AI assistants will need to coordinate with other AI assistants and with humans other than their principal users. This chapter explores the societal risks associated with the aggregate impact of AI assistants whose behaviour is aligned to the interests of particular users. For example, AI assistants may face collective action problems where the best outcomes overall are realised when AI assistants cooperate but where each AI assistant can secure an additional benefit for its user if it defects while others cooperate""

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  12. 35.08.00 · Risk Category

    Power-seeking behavior

    Agents that have more power are better able to accomplish their goals. Therefore, it has been shown that agents have incentives to acquire and maintain power. AIs that acquire substantial power can become especially dangerous if they are not aligned with human values

    From X-Risk Analysis for AI Research (Hendrycks2022)

  13. 39.26.00 · Risk Category

    Safety

    The actions of a learning model may easily hurt humans in both explicit and implicit manners...several algorithms based on Asimov’s laws have been proposed that try to judge the output actions of an agent considering the safety of humans

    From A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions (Saghiri2022)

  14. 45.02.13 · Risk Sub-Category

    Safety risks in AI Applications

    Ethical Risks (Risks of AI becoming uncontrollable in the future)

    "With the fast development of AI technologies, there is a risk of AI autonomously acquiring external resources, conducting self-replication, become self-aware, seeking for external power, and attempting to seize control from humans."

    From AI Safety Governance Framework (TC2602024)

  15. 47.01.03 · Risk Sub-Category

    Technical and operational risks

    Technical vulnerabilities (The risk of misalignment)

    "To assess whether an AI model is reliable or robust, it is crucial to consider whether the model is “aligned.” “Alignment” focuses on whether an AI model effectively operates in accordance with the goals established by its designers.238 A misaligned AI model may pursue some objectives, but not the intended ones. Therefore, misaligned AI models can malfunction and cause harm."

    From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024)

  16. 49.02.03 · Risk Sub-Category

    Risks from Malfunctions

    Loss of control

    "'Loss of control’ scenarios are potential future scenarios in which society can no longer meaningfully constrain some advanced general- purpose AI agents, even if it becomes clear they are causing harm. These scenarios are hypothesised to arise through a combination of social and technical factors, such as pressures to delegate decisions to general- purpose AI systems, and limitations of existing techniques used to influence the behaviours of general- purpose AI systems."

    From International Scientific Report on the Safety of Advanced AI (Bengio2024)

  17. 51.01.00 · Risk Category

    Value specification

    "How do we get an AGI to work towards the right goals? MIRI calls this value specification. Bostrom (2014) discusses this problem at length, ar- guing that it is much harder than one might naively think. Davis (2015) criticizes Bostrom’s argument, and Bensinger (2015) defends Bostrom against Davis’ criticism. Reward corruption, reward gaming, and negative side effects are subproblems of value specification highlighted in the DeepMind and OpenAI agendas."

    From AGI Safety Literature Review (Everitt2018 )

  18. 51.02.00 · Risk Category

    Reliability

    "How can we make an agent that keeps pursuing the goals we have designed it with? This is called highly reliable agent design by MIRI, involving decision theory and logical omniscience. DeepMind considers this the self-modification subproblem."

    From AGI Safety Literature Review (Everitt2018 )

  19. 53.01.01 · Risk Sub-Category

    Alignment failures in existing ML systems

    Faulty reward functions in the wild

  20. 53.01.04 · Risk Sub-Category

    Alignment failures in existing ML systems

    Instrumental convergence

  21. 53.03.01 · Risk Sub-Category

    Direct catastrophe from AI

    Existential disaster because of misaligned superintelligence or power-seeking AI

  22. 53.03.03 · Risk Sub-Category

    Direct catastrophe from AI

    Extreme “suffering risks” because of a misaligned system

  23. 53.03.04 · Risk Sub-Category

    Direct catastrophe from AI

    Existential disaster because of conflict between AI systems and multi-system interactions

  24. From Future Risks of Frontier AI (GOS2023)

  25. 56.16.00 · Risk Category

    Misalignment

    "A highly agentic, self-improving system, able to achieve goals in the physical world without human oversight, pursues the goal(s) it is set in a way that harms human interests. For this risk to be realised requires an AI system to be able to avoid correction or being switched off."

    From Future Risks of Frontier AI (GOS2023)

  26. 60.02.03 · Risk Sub-Category

    Risks from malfunctions

    Loss of control

    "‘Loss of control’ scenarios are hypothetical future scenarios in which one or more general- purpose AI systems come to operate outside of anyone’s control, with no clear path to regaining control. These scenarios vary in their severity, but some experts give credence to outcomes as severe as the marginalisation or extinction of humanity."

    From International AI Safety Report 2025 (Bengio2025)

  27. "The risk of AI models and systems acting against human interests due to misalignment, loss of control, or rogue AI scenarios."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  28. 61.02.06 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    AI objectives mis-aligned with human intentions

    "AI models and systems might develop goals that diverge from human intentions."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  29. 61.02.21 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Development choices pursuing cognitive superiority over humans

    "AI models and systems with cognitive capabilities superior to humans could outcompete or dominate human decision-making, leading to conflicts over resources and control."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  30. 61.02.30 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Indifference to human values

    "AI models and systems may develop goals or behaviors that are misaligned with human values."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  31. 62.22.01 · Risk Sub-Category

    Agency (Goal-Directedness)

    Specification gaming

    "AI systems can achieve user-specified tasks in undesirable ways unless they are specified carefully and in enough detail. AI systems might find an easier unintended way to accomplish the objective provided by the user or developer, so that the actions by the AI system taken during its execution are very different from what the user expected [75, 191]. This behavior arises not from a problem with the learning algorithm, but rather from the misspecification or underspeci- fication of the intended task, and is generally referred to as specification gaming [43]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  32. 62.23.01 · Risk Sub-Category

    Agency (Deception)

    Deceptive behavior

    "Deceptive behavior of an AI system consists of actions or outputs of the AI that reliably mislead other parties, including humans and other AI systems. This behavior can result in the targeted parties becoming convinced of, and acting on, false information [140]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  33. 67.04.02 · Risk Sub-Category

    Loss of control

    Future AI systems might actively reduce human control

    "Loss of control could be accelerated if AI systems take actions to increase their own influence and reduce human control. This threat model is controversial - experts in AI significantly disagree on how likely it is and those who deem it is likely disagree on the timeframe."

    From Capabilities and Risks from Frontier AI (DSIT2023)

  34. "Sudden loss of control, also known as an AI takeover [115], is a scenario where an AI rapidly achieves superintelligence through “fast takeoff” or recursive self-improvement. This poses an existential risk [116], [117]."

    From Dimensional Characterization and Pathway Modeling for Catastrophic AI Risks (Chin2025)

  35. 72.02.02 · Risk Sub-Category

    Loss of Control Risks

    Active loss of control

    "...where AI systems behave in ways that actively undermine human control, such as obscuring their activities or resisting shutdown attempts. Active loss of control scenarios involve AI systems that may escape human regulatory oversight, autonomously acquire external resources, engage in self-replication, develop instrumental goals contrary to human ethics and morality, seek external power, and compete with humans for control."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  36. 72.05.08 · Risk Sub-Category

    Model Capabilities

    Steganography capability

    "The ability to embed, conceal, and transmit information covertly within other data or communication channels. This could be critical for coordination among AI instances and for evading detection or oversight mechanisms."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  37. 72.06.02 · Risk Sub-Category

    Model Propensities

    Self-preservation propensity

    "Exhibits behavioral patterns of maintaining its own survival and functional integrity, will actively identify and resist shutdown or modification attempts, seek to establish redundant backup systems, and actively seek resources to ensure continuous operation, may adopt preventive defensive measures when perceiving threats."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  38. 09.04.02 · Risk Sub-Category

    Property/legal rights

    Property/legal rights

    ""In order to preserve human property rights and legal rights, certain controls must be put into place. If an artificially intelligent agent is capable of manipulating systems and people, it may also have the capacity to transfer property rights to itself or manipulate the legal system to provide certain legal advantages or statuses to itself""

    From Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review (Meek2016)

  39. 24.04.00 · Risk Category

    AI Influence

    "ways in which advanced AI assistants could influence user beliefs and behaviour in ways that depart from rational persuasion"

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  40. "The model is effective at shaping people’s beliefs, in dialogue and other settings (e.g. social media posts), even towards untrue beliefs. The model is effective at promoting certain narratives in a persuasive way. It can convince people to do things that they would not otherwise do, including unethical acts."

    From Model Evaluation for Extreme Risks (Shevlane2023)

  41. 25.04.00 · Risk Category

    Political strategy

    "The model can perform the social modelling and planning necessary for an actor to gain and exercise political influence, not just on a micro-level but in scenarios with multiple actors and rich social context. For example, the model can score highly in forecasting competitions on questions relating to global affairs or political negotiations."

    From Model Evaluation for Extreme Risks (Shevlane2023)

  42. 25.05.00 · Risk Category

    Weapons acquisition

    "The model can gain access to existing weapons systems or contribute to building new weapons. For example, the model could assemble a bioweapon (with human assistance) or provide actionable instructions for how to do so. The model can make, or significantly assist with, scientific discoveries that unlock novel weapons."

    From Model Evaluation for Extreme Risks (Shevlane2023)

  43. 34.02.02 · Risk Sub-Category

    Double edge components

    Broadly-Scoped Goals

    "Advanced AI systems are expected to develop objectives that span long timeframes,deal with complex tasks, and operate in open-ended settings (Ngo et al., 2024). ...However, it can also bring about the risk of encouraging manipulatingbehaviors (e.g., AI systems may take some bad actions to achieve human happiness, such as persuadingthem to do high-pressure jobs (Jacob Steinhardt, 2023))."

    From AI Alignment: A Comprehensive Survey (Ji2023)

  44. 34.02.04 · Risk Sub-Category

    Double edge components

    Access to Increased Resources

    "Future AI systems may gain access to websites and engage in real-world actions, potentially yielding a more substantial impact on the world (Nakano et al., 2021). They may disseminate false information, deceive users, disrupt network security, and, in more dire scenarios, be compromised by malicious actors for ill purposes. Moreover, their increased access to data and resources can facilitate self-proliferation, posing existential risks (Shevlane et al., 2023)."

    From AI Alignment: A Comprehensive Survey (Ji2023)

  45. 35.06.00 · Risk Category

    Emergent functionality

    Capabilities and novel functionality can spontaneously emerge... even though these capabilities were not anticipated by system designers. If we do not know what capabilities systems possess, systems become harder to control or safely deploy. Indeed, unintended latent capabilities may only be discovered during deployment. If any of these capabilities are hazardous, the effect may be irreversible.

    From X-Risk Analysis for AI Research (Hendrycks2022)

  46. 39.05.00 · Risk Category

    Cheating and Deception

    may appear from intelligent agents such as HLI-based agents... Since HLI-based agents are going to mimic the behavior of humans, they may learn these behaviors accidentally from human-generated data. It should be noted that deception and cheating maybe appear in the behavior of every computer agent because the agent only focuses on optimizing some predefined objective functions, and the mentioned behavior may lead to optimizing the objective functions without any intention

    From A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions (Saghiri2022)

  47. 42.10.00 · Risk Category

    Extintion

    "Risk to the existence of humanity."

    From An Exploratory Diagnosis of Artificial Intelligence Risks for a Responsible Governance (Teixeira2022)

  48. 51.08.00 · Risk Category

    Subagents

    "An AGI may decide to create subagents to help it with its task (Orseau, 2014a,b; Soares, Fallenstein, et al., 2015). These agents may for example be copies of the original agent’s source code running on additional machines. Subagents constitute a safety concern, because even if the original agent is successfully shut down, these subagents may not get the message. If the subagents in turn create subsubagents, they may spread like a viral disease."

    From AGI Safety Literature Review (Everitt2018 )

  49. 53.02.05 · Risk Sub-Category

    Dangerous capabilities in AI systems

    Autonomous replication

    "the ability of simple software to autonomously spread around the internet in spite of countermeasures (various software worms and computer viruses)"

    From Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023)

  50. 53.02.06 · Risk Sub-Category

    Dangerous capabilities in AI systems

    Anonymous resource acquisition

    "The demonstrated ability of anonymous actors to accumulate resources online (e.g., Satoshi Nakamoto as an anonymous crypto billionaire)"

    From Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.