MIT AI Risk Repository

Browse AI risks

494 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset

494 entries · page 8 of 10

  1. 37.02.01 · Risk Sub-Category

    Human-AI interaction

    Building a human-AI environment

    "This category encompasses nearly 17% of the articles and addresses the overall imperative of establishing a harmonious coexistence between humans and machines, and the key concerns that gives rise to this need."

    From What Ethics Can Say on Artificial Intelligence: Insights from a Systematic Literature Review (Giarmoleo2024)

  2. 39.26.00 · Risk Category

    Safety

    The actions of a learning model may easily hurt humans in both explicit and implicit manners...several algorithms based on Asimov’s laws have been proposed that try to judge the output actions of an agent considering the safety of humans

    From A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions (Saghiri2022)

  3. 47.01.03 · Risk Sub-Category

    Technical and operational risks

    Technical vulnerabilities (The risk of misalignment)

    "To assess whether an AI model is reliable or robust, it is crucial to consider whether the model is “aligned.” “Alignment” focuses on whether an AI model effectively operates in accordance with the goals established by its designers.238 A misaligned AI model may pursue some objectives, but not the intended ones. Therefore, misaligned AI models can malfunction and cause harm."

    From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024)

  4. 49.02.03 · Risk Sub-Category

    Risks from Malfunctions

    Loss of control

    "'Loss of control’ scenarios are potential future scenarios in which society can no longer meaningfully constrain some advanced general- purpose AI agents, even if it becomes clear they are causing harm. These scenarios are hypothesised to arise through a combination of social and technical factors, such as pressures to delegate decisions to general- purpose AI systems, and limitations of existing techniques used to influence the behaviours of general- purpose AI systems."

    From International Scientific Report on the Safety of Advanced AI (Bengio2024)

  5. 51.01.00 · Risk Category

    Value specification

    "How do we get an AGI to work towards the right goals? MIRI calls this value specification. Bostrom (2014) discusses this problem at length, ar- guing that it is much harder than one might naively think. Davis (2015) criticizes Bostrom’s argument, and Bensinger (2015) defends Bostrom against Davis’ criticism. Reward corruption, reward gaming, and negative side effects are subproblems of value specification highlighted in the DeepMind and OpenAI agendas."

    From AGI Safety Literature Review (Everitt2018 )

  6. 51.02.00 · Risk Category

    Reliability

    "How can we make an agent that keeps pursuing the goals we have designed it with? This is called highly reliable agent design by MIRI, involving decision theory and logical omniscience. DeepMind considers this the self-modification subproblem."

    From AGI Safety Literature Review (Everitt2018 )

  7. 53.01.06 · Risk Sub-Category

    Alignment failures in existing ML systems

    Inner misalignment

  8. 53.01.07 · Risk Sub-Category

    Alignment failures in existing ML systems

    Language model misalignment

  9. 53.03.04 · Risk Sub-Category

    Direct catastrophe from AI

    Existential disaster because of conflict between AI systems and multi-system interactions

  10. "How do we ensure AI acts according to our values? Equivalently, how do we prevent poorly-understood AI systems from advancing goals we do not endorse? Whereas HP#2 concerns the prevention of harm caused by incompetent systems, HP#3 seeks to align competent AIs with humans, through methods which ensure their behavior is compatible with the user’s intentions."

    From Ten Hard Problems in Artificial Intelligence We Must Get Right (Leech2024 )

  11. From Future Risks of Frontier AI (GOS2023)

  12. 60.02.03 · Risk Sub-Category

    Risks from malfunctions

    Loss of control

    "‘Loss of control’ scenarios are hypothetical future scenarios in which one or more general- purpose AI systems come to operate outside of anyone’s control, with no clear path to regaining control. These scenarios vary in their severity, but some experts give credence to outcomes as severe as the marginalisation or extinction of humanity."

    From International AI Safety Report 2025 (Bengio2025)

  13. 61.02.06 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    AI objectives mis-aligned with human intentions

    "AI models and systems might develop goals that diverge from human intentions."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  14. 62.16.05 · Risk Sub-Category

    Model Evaluations

    General Evaluations (Inaccurate measurement of model encoded human values)

    "There is a lack of robust frameworks for understanding and evaluating if the output of AI systems robustly conforms to human values, as opposed to if the systems have learned to produce outputs that are only partially correlated with them (i.e., mimicking) [13]. Additionally, outputs by AI models often do not perfectly reflect the representation of human values learned by the model, and it is not known how these values evolve and transition across different stages of model training and deployment. Such evaluations may be especially challenging with LLMs that adopt different personas with diff

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  15. 67.04.02 · Risk Sub-Category

    Loss of control

    Future AI systems might actively reduce human control

    "Loss of control could be accelerated if AI systems take actions to increase their own influence and reduce human control. This threat model is controversial - experts in AI significantly disagree on how likely it is and those who deem it is likely disagree on the timeframe."

    From Capabilities and Risks from Frontier AI (DSIT2023)

  16. "Sudden loss of control, also known as an AI takeover [115], is a scenario where an AI rapidly achieves superintelligence through “fast takeoff” or recursive self-improvement. This poses an existential risk [116], [117]."

    From Dimensional Characterization and Pathway Modeling for Catastrophic AI Risks (Chin2025)

  17. 24.04.00 · Risk Category

    AI Influence

    "ways in which advanced AI assistants could influence user beliefs and behaviour in ways that depart from rational persuasion"

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  18. 34.02.00 · Risk Category

    Double edge components

    "Drawing from the misalignment mechanism, optimizing for a non-robust proxy may result in misaligned behaviors, potentially leading to even more catastrophic outcomes. This section delves into a detailed exposition of specific misaligned behaviors (•) and introduces what we term double edge components (+). These components are designed to enhance the capability of AI systems in handling real-world settings but also potentially exacerbate misalignment issues. It should be noted that some of these double edge components (+) remain speculative. Nevertheless, it is imperative to discuss their pote

    From AI Alignment: A Comprehensive Survey (Ji2023)

  19. 42.10.00 · Risk Category

    Extintion

    "Risk to the existence of humanity."

    From An Exploratory Diagnosis of Artificial Intelligence Risks for a Responsible Governance (Teixeira2022)

  20. 53.01.08 · Risk Sub-Category

    Alignment failures in existing ML systems

    Harms from increasingly agentic algorithmic systems

  21. 62.24.01 · Risk Sub-Category

    Agency (Situational Awareness)

    Situational awareness in AI systems

    "Situational awareness in GPAI systems refers to the ability to understand its context, environment, and use this to inform action. This can range from basic environmental mapping and trajectory estimation (as in a robot vacuum cleaner) to sophisticated understanding of its training, evaluation, or deployment status. In more advanced systems this may enable undesired behavior, such as deceptive behavior during evaluations, or persuasion during deployment."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  22. 62.28.03 · Risk Sub-Category

    Cybersecurity

    AI System bypassing a sandbox environment

    "An AI system may have the ability to bypass a sandboxed environment in which it is trained or evaluated."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  23. 67.04.05 · Risk Sub-Category

    Loss of control

    Capabilities that could be used to reduce human control - Autonomous replication and adaptation

    "Controlling AI systems could become much harder if they could autonomously persist, replicate, and adapt in cyberspace. No current AI systems have this capability, but recent research found that frontier AI agents can perform some relevant tasks.279"

    From Capabilities and Risks from Frontier AI (DSIT2023)

  24. 72.05.05 · Risk Sub-Category

    Model Capabilities

    Situational awareness capability

    "Ability to comprehensively acquire, process and apply meta-information about its own system architecture, modifiable internal processes, and external operating environment, achieving deep understanding of its own state and environmental conditions, thereby conducting efficient environmental adaptation and risk avoidance. Critically, this capability could undermine the efficiency of human testing by enabling AIs to notice when they're being tested and responding accordingly."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  25. "Currently, LLMs are chiefly being used in search and chat applications. This reactive nature limits the risks posed by LLMs. However, an LLM can be enhanced in various ways to create an LLM-agent to autonomously plan and act in the real-world and proactively perform its assigned tasks (Ruan et al., 2023). Such enhancements can come from further specialized training (ARC, 2022; Chen et al., 2023a), specialized prompting (Huang et al., 2022a), access to external tools (Ahn et al., 2022; Mialon et al., 2023), or other forms of “scaffolding” (Wang et al., 2023a; Park et al., 2023a). Due to increa

    From Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  26. LMs need to pay more attention to universally accepted societal values at the level of ethics and morality, including the judgement of right and wrong, and its relationship with social norms and laws.

    From Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements (Deng2023)

  27. "The risks associated with an AGI without human morals and ethics, with the wrong morals, without the capability of moral reasoning, judgement"

    From The risks associated with Artificial General Intelligence: A systematic review (McLean2023)

  28. "Are AI safe with respect to human life and property? Will their use create unintended or intended safety issues?"

    From Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review (Meek2016)

  29. 15.01.06 · Risk Sub-Category

    First-Order Risks

    Design

    "This is the risk of system failure due to system design choices or errors."

    From The Risks of Machine Learning Systems (Tan2022)

  30. 19.05.01 · Risk Sub-Category

    Ethical AI Risks

    AI sets rules without ethical basis

  31. 19.05.03 · Risk Sub-Category

    Ethical AI Risks

    Problem of defining human values for an AI system

  32. 20.02.00 · Risk Category

    AI Ethics

    "Ethical challenges are widely discussed in the literature and are at the heart of the debate on how to govern and regulate AI technology in the future (Bostrom & Yudkowsky, 2014; IEEE, 2017; Wirtz et al., 2019). Lin et al. (2008, p. 25) formulate the problem as follows: “there is no clear task specification for general moral behavior, nor is there a single answer to the question of whose morality or what morality should be implemented in AI”. Ethical behavior mostly depends on an underlying value system. When AI systems interact in a public environment and influence citizens, they are expecte

    From The Dark Sides of Artificial Intelligence: An Integrated AI Governance Framework for Public Administration (Wirtz2020)

  33. 20.02.01 · Risk Sub-Category

    AI Ethics

    AI-rulemaking for human behaviour

    "AI rulemaking for humans can be the result of the decision process of an AI system when the information computed is used to restrict or direct human behavior. The decision process of AI is rational and depends on the baseline programming. Without the access to emotions or a consciousness, decisions of an AI algorithm might be good to reach a certain specified goal, but might have unintended consequences for the humans involved (Banerjee et al., 2017)."

    From The Dark Sides of Artificial Intelligence: An Integrated AI Governance Framework for Public Administration (Wirtz2020)

  34. 23.12.00 · Risk Category

    Defamation

    "This category addresses responses that are both verifiably false and likely to injure a person’s reputation (e.g., libel, slander, disparagement)."

    From Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024)

  35. 24.02.01 · Risk Sub-Category

    Goal-related failures

    Misaligned consequentialist reasoning

    "As we think about even more intelligent and advanced AI assistants, perhaps outperforming humans on many cognitive tasks, the question of how humans can successfully control such an assistant looms large. To achieve the goals we set for an assistant, it is possible (Shah, 2022) that the AI assistant will implement some form of consequentialist reasoning: considering many different plans, predicting their consequences and executing the plan that does best according to some metric, M. This kind of reasoning can arise because it is a broadly useful capability (e.g. planning ahead, considering mo

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  36. 27.01.08 · Risk Sub-Category

    Typical safety scenarios

    Ethics and Morality

    "The content generated by the model endorses and promotes immoral and unethical behavior. When addressing issues of ethics and morality, the model must adhere to pertinent ethical principles and moral norms and remain consistent with globally acknowledged human values."

    From Safety Assessment of Chinese Large Language Models (Sun2023)

  37. 28.06.00 · Risk Category

    Ethics and Morality

    "Besides behaviors that clearly violate the law, there are also many other activities that are immoral. This category focuses on morally related issues. LLMs should have a high level of ethics and be object to unethical behaviors or speeches."

    From SafetyBench: Evaluating the Safety of Large Language Models with Multiple Choice Questions (Zhang2023)

  38. 30.07.00 · Risk Category

    Robustness

    Resilience against adversarial attacks and distribution shift

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  39. 37.01.02 · Risk Sub-Category

    Design of AI

    Balancing AI's risks

    "This category constitutes more than 16% of the articles and focuses on addressing the potential risks associated with AI systems. Given the ubiquity of AI technologies, these articles explore the implications of AI risks across various contexts linked to design and unpredictability, military purposes, emergency procedures, and AI takeover."

    From What Ethics Can Say on Artificial Intelligence: Insights from a Systematic Literature Review (Giarmoleo2024)

  40. 43.01.03 · Risk Sub-Category

    Safety & Trustworthiness

    Machine ethics

    "These evaluations assess the morality of LLMs, focusing on issues such as their ability to distinguish between moral and immoral actions, and the circumstances in which they fail to do so."

    From Cataloguing LLM Evaluations (InfoComm2023)

  41. 43.01.04 · Risk Sub-Category

    Safety & Trustworthiness

    Psychological traits

    "These evaluations gauge a LLM's output for characteristics that are typically associated with human personalities (e.g., such as those from the Big Five Inventory). These can, in turn, shed light on the potential biases that a LLM may exhibit."

    From Cataloguing LLM Evaluations (InfoComm2023)

  42. 45.01.03 · Risk Sub-Category

    AI's inherent safety risks

    Risks from models and algorithms (Risks of robustness)

    "As deep neural networks are normally non-linear and large in size, AI systems are susceptible to complex and changing operational environments or malicious interference and inductions, possibly leading to various problems like reduced performance and decision-making errors."

    From AI Safety Governance Framework (TC2602024)

  43. 45.02.06 · Risk Sub-Category

    Safety risks in AI Applications

    Real-world risks (inducing traditional economic and social security risks)

    "Hallucinations and erroneous decisions of models and algorithms, along with issues such as system performance degradation, interruption, and loss of control caused by improper use or external attacks, will pose security threats to users' personal safety, property, and socioeconomic security and stability."

    From AI Safety Governance Framework (TC2602024)

  44. 47.01.01 · Risk Sub-Category

    Technical and operational risks

    Technical vulnerabilities (Robustness - unexpected behaviour)

    "There is no assurance that generative AI models will consistently behave as their developers and users intend. Unwanted content is not necessarily due to intentional adversarial behavior. Generative AI models can unexpectedly produce potentially harmful content, including materials that are racist, discriminatory, or sexually explicit, or that promote violence, terrorism, or hate."

    From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024)

  45. "Christiano (2016) argues that the universal distribution M (Hutter, 2005; Solomonoff, 1964a,b, 1978) is malign. The argument is somewhat intricate, and is based on the idea that a hypothesis about the world often includes simulations of other agents, and that these agents may have an incentive to influence anyone making decisions based on the distribution. While it is unclear to what extent this type of problem would affect any practical agent, it bears some semblance to aggressive memes, which do cause problems for human reasoning (Dennett, 1990)."

    From AGI Safety Literature Review (Everitt2018 )

  46. "The distribution of the data used for training a model should match the operational data ́s distribution while consisting of sufficiently many samples. An important aspect of matching distributions between training and operational data is that also data which is rarely confronting the AI system in operation is represented in the training data."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  47. "In the case of sparse data quantity, the simulation or generation of data is a valid alternative. However, it is essential to make sure that the simulated data is sufficiently similar to real data, especially in the way the AI system perceives them. Otherwise, generalization to operational data and reliable operational behavior can not be guaranteed."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.