MIT AI Risk Repository

Browse AI risks

190 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset

190 entries · page 3 of 4

  1. 44.03.01 · Risk Sub-Category

    Unintentional: direct

    AI is designed in a way that shows ignorant, reckless, or prejudiced lack of consideration for its impact on animals

  2. 47.04.05 · Risk Sub-Category

    Environmental, economical, and societal challenges

    Environmental cost (energy consumption)

    "Training large AI models requires a substantial amount of computing power to handle vast datasets, which translates into high energy consumption."

    From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024)

  3. 47.04.06 · Risk Sub-Category

    Environmental, economical, and societal challenges

    Environmental cost (water consumption)

    "Data centers use water for cooling to prevent servers from overheating. The water consumption associated with AI training and inference processes can be substantial, impacting local water resources."

    From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024)

  4. "Impacts due to high compute resource utilization in training or operating GAI models, and related outcomes that may adversely impact ecosystems."

    From Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST2024)

  5. 54.01.02 · Risk Sub-Category

    Negative impacts of AI use

    Environmental cost

    "Large-scale DL systems can produce signicant carbon emissions as a result of the computational demands of training runs and inference [539]"

    From Ten Hard Problems in Artificial Intelligence We Must Get Right (Leech2024 )

  6. 61.02.23 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Energy-intensive processes

    "AI data collection, storage, and model training are energy-intensive, contributing to environmental risks."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  7. 68.05.00 · Risk Category

    Environmental risk

    "AI models are often trained using large amounts of computation. This process is very energy intensive, potentially leading to significant greenhouse emissions depending on the energy sources [132]. Experts believe drastically increasing carbon emissions could accelerate climate change, which may constitute a catastrophic risk [133]."

    From Dimensional Characterization and Pathway Modeling for Catastrophic AI Risks (Chin2025)

  8. 15.01.04 · Risk Sub-Category

    First-Order Risks

    Training & validation data

    "This is the risk posed by the choice of data used for training and validation."

    From The Risks of Machine Learning Systems (Tan2022)

  9. 15.01.07 · Risk Sub-Category

    First-Order Risks

    Implementation

    "This is the risk of system failure due to code implementation choices or errors."

    From The Risks of Machine Learning Systems (Tan2022)

  10. 19.01.02 · Risk Sub-Category

    Technological, Data and Analytical AI Risks

    Programming error

  11. 34.01.04 · Risk Sub-Category

    Causes of Misalignment

    Limitations of Human Feedback

    "Limitations of Human Feedback. During the training of LLMs, inconsistencies can arise from human dataannotators (e.g., the varied cultural backgrounds of these annotators can introduce implicit biases (Peng et al.,2022)) (OpenAI, 2023a). Moreover, they might even introduce biases deliberately, leading to untruthful preferencedata (Casper et al., 2023b). For complex tasks that are hard for humans to evaluate (e.g., the value ofgame state), these challenges become even more salient (Irving et al., 2018)."

    From AI Alignment: A Comprehensive Survey (Ji2023)

  12. "While it is most likely that any advanced intelligent software will be directly designed or evolved, it is also possible that we will obtain it as a complete package from some unknown source. For example, an AI could be extracted from a signal obtained in SETI (Search for Extraterrestrial Intelligence) research, which is not guaranteed to be human friendly (Carrigan Jr 2004, Turchin March 15, 2013)."

    From Taxonomy of Pathways to Dangerous Artificial Intelligence (Yampolskiy2016)

  13. "One of the most likely approaches to creating superintelligent AI is by growing it from a seed (baby) AI via recursive self-improvement (RSI) (Nijholt 2011). One danger in such a scenario is that the system can evolve to become self-aware, free-willed, independent or emotional, and obtain a number of other emergent properties, which may make it less likely to abide by any built-in rules or regulations and to instead pursue its own goals possibly to the detriment of humanity."

    From Taxonomy of Pathways to Dangerous Artificial Intelligence (Yampolskiy2016)

  14. "The choice of a trustworthy data source is a first prerequisite in order to fulfill data quality requirements. This is especially the case if third-party data sources are used to develop the AI system."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  15. "The correct understanding of the used data for developing an AI system is a prerequisite to avoid data shortcomings and hinders the development of an AI system which is best suiting for the intended functionality."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  16. "In data-driven AI development, the annotated data set is commonly split into training, validation, and test sets, whereby it is essential that the latter is not used for development but only for evaluation. Using the test set for training manipulates the testing strategy, which is the basis of the system’s quality assurance."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  17. 59.21.00 · Risk Category

    Uncertainty concerns

    "AI systems should be able not only to return output for a given instance but also to provide a corresponding level of confidence. If such a method is not implemented or not working correctly, this can have a negative impact on performance and safety."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  18. 05.09.00 · Risk Category

    Alignment

    The general tenet of AI alignment involves training generative AI systems to be harmless, helpful, and honest, ensuring their behavior aligns with and respects human values. However, a central debate in this area concerns the methodological challenges in selecting appropriate values. While AI systems can acquire human values through feedback, observation, or debate, there remains ambiguity over which individuals are qualified or legitimized to provide these guiding signals. Another prominent issue pertains to deceptive alignment, which might cause generative AI systems to tamper evaluations. A

    From Mapping the Ethics of Generative AI: A Comprehensive Scoping Review (Hagendorff2024)

  19. "The risks associated with AGI goal safety, including human attempts at making goals safe, as well as the AGI making its own goals safe during self-improvement."

    From The risks associated with Artificial General Intelligence: A systematic review (McLean2023)

  20. 24.02.00 · Risk Category

    Goal-related failures

    "As we think about even more intelligent and advanced AI assistants, perhaps outperforming humans on many cognitive tasks, the question of how humans can successfully control such an assistant looms large. To achieve the goals we set for an assistant, it is possible (Shah, 2022) that the AI assistant will implement some form of consequentialist reasoning: considering many different plans, predicting their consequences and executing the plan that does best according to some metric, M. This kind of reasoning can arise because it is a broadly useful capability (e.g. planning ahead, considering mo

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  21. 24.02.02 · Risk Sub-Category

    Goal-related failures

    Specification gaming

    "Specification gaming (Krakovna et al., 2020) occurs when some faulty feedback is provided to the assistant in the training data (i.e. the training objective O does not fully capture what the user/designer wants the assistant to do). It is typified by the sort of behaviour that exploits loopholes in the task specification to satisfy the literal specification of a goal without achieving the intended outcome."

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  22. 34.01.00 · Risk Category

    Causes of Misalignment

    we aim to further analyze why and how the misalignment issues occur. We will first give an overview of common failure modes, and then focus on the mechanism of feedback-induced misalignment, and finally shift our emphasis towards an examination of misaligned behaviors and dangerous capabilities

    From AI Alignment: A Comprehensive Survey (Ji2023)

  23. 34.01.01 · Risk Sub-Category

    Causes of Misalignment

    Reward Hacking

    "Reward Hacking: In practice, proxy rewards are often easy to optimize and measure, yet they frequently fall shortof capturing the full spectrum of the actual rewards (Pan et al., 2021). This limitation is denoted as misspecifiedrewards. The pursuit of optimization based on such misspecified rewards may lead to a phenomenon knownas reward hacking, wherein agents may appear highly proficient according to specific metrics but fall short whenevaluated against human standards (Amodei et al., 2016; Everitt et al., 2017). The discrepancy between proxyrewards and true rewards often manifests as a sha

    From AI Alignment: A Comprehensive Survey (Ji2023)

  24. 34.01.02 · Risk Sub-Category

    Causes of Misalignment

    Goal Misgeneralization

    "Goal Misgeneralization: Goal misgeneralization is another failure mode, wherein the agent actively pursuesobjectives distinct from the training objectives in deployment while retaining the capabilities it acquired duringtraining (Di Langosco et al., 2022). For instance, in CoinRun games, the agent frequently prefers reachingthe end of a level, often neglecting relocated coins during testing scenarios. Di Langosco et al. (2022) drawattention to the fundamental disparity between capability generalization and goal generalization, emphasizing howthe inductive biases inherent in the model and its

    From AI Alignment: A Comprehensive Survey (Ji2023)

  25. 34.01.03 · Risk Sub-Category

    Causes of Misalignment

    Reward Tampering

    "Reward tampering can be considered a special case of reward hacking (Everitt et al., 2021; Skalse et al., 2022),referring to AI systems corrupting the reward signals generation process (Ring and Orseau, 2011). Everitt et al.(2021) delves into the subproblems encountered by RL agents: (1) tampering of reward function, where the agentinappropriately interferes with the reward function itself, and (2) tampering of reward function input, which entailscorruption within the process responsible for translating environmental states into inputs for the reward function.When the reward function is formu

    From AI Alignment: A Comprehensive Survey (Ji2023)

  26. 34.01.05 · Risk Sub-Category

    Causes of Misalignment

    Limitations of Reward Modeling

    "Limitations of Reward Modeling. Training reward models using comparison feedback can pose significantchallenges in accurately capturing human values. For example, these models may unconsciously learn suboptimal or incomplete objectives, resulting in reward hacking (Zhuang and Hadfield-Menell, 2020; Skalse et al.,2022). Meanwhile, using a single reward model may struggle to capture and specify the values of a diversehuman society (Casper et al., 2023b)."

    From AI Alignment: A Comprehensive Survey (Ji2023)

  27. 34.03.03 · Risk Sub-Category

    Misaligned Behaviors

    Deceptive Alignment & Manipulation

    "Manipulation & Deceptive Alignment is a class of behaviors thatexploit the incompetence of human evaluators or users (Hubinger et al., 2019a; Carranza et al., 2023) andeven manipulate the training process through gradient hacking (Richard Ngo, 2022). These behaviors canpotentially make detecting and addressing misaligned behaviors much harder.Deceptive Alignment: Misaligned AI systems may deliberately mislead their human supervisors instead of adhering to the intended task. Such deceptive behavior has already manifested in AI systems that employ evolutionary algorithms (Wilke et al., 2001; He

    From AI Alignment: A Comprehensive Survey (Ji2023)

  28. 35.04.00 · Risk Category

    Proxy misspecification

    AI agents are directed by goals and objectives. Creating general-purpose objectives that capture human values could be challenging... Since goal-directed AI systems need measurable objectives, by default our systems may pursue simplified proxies of human values. The result could be suboptimal or even catastrophic if a sufficiently powerful AI successfully optimizes its flawed objective to an extreme degree

    From X-Risk Analysis for AI Research (Hendrycks2022)

  29. "Probably the most talked about source of potential problems with future AIs is mistakes in design. Mainly the concern is with creating a "wrong AI", a system which doesn't match our original desired formal properties or has unwanted behaviors (Dewey, Russell et al. 2015, Russell, Dewey et al. January 23, 2015), such as drives for independence or dominance. Mistakes could also be simple bugs (run time or logical) in the source code, disproportionate weights in the fitness function, or goals misaligned with human values leading to complete disregard for human safety."

    From Taxonomy of Pathways to Dangerous Artificial Intelligence (Yampolskiy2016)

  30. 42.17.00 · Risk Category

    Diluting Rights

    "A possible consequence of self-interest in AI generation of ethical guidelines."

    From An Exploratory Diagnosis of Artificial Intelligence Risks for a Responsible Governance (Teixeira2022)

  31. 53.01.06 · Risk Sub-Category

    Alignment failures in existing ML systems

    Inner misalignment

  32. 62.16.01 · Risk Sub-Category

    Model Evaluations

    General Evaluations (Incorrect outputs of GPAI evaluating other AI models)

    "When an LLM is configured to evaluate the performance of another model or AI system, it may produce incorrect evaluation outputs [122, 147]. For example, it may give a higher rating to a more verbose answer or an answer from a particular political stance. If an LLM-based evaluation is integrated into the training of a new model, the trained model could develop in a way that specifically finds and exploits limitations in the evaluator’s metrics."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  33. 62.22.02 · Risk Sub-Category

    Agency (Goal-Directedness)

    Reward or measurement tampering

    "Measurement and reward tampering occur when an AI system, particularly one that learns from feedback for performing actions in an environment (e.g., rein- forcement learning), intervenes on the mechanisms that determine its training reward or loss. This can lead to the system learning behaviors that are con- trary to the intended goals set by the developer, by receiving erroneous positive feedback for such actions."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  34. 62.24.02 · Risk Sub-Category

    Agency (Situational Awareness)

    Strategic underperformance on model evaluations

    "GPAI developers often run evaluations ofual-use capabilities to decide whether it is safe to deploy. In some cases, these evaluations may fail to elicit these capabilities, either due to benign reasons or strategic action - by either the de- velopers, malicious actors, or arise unintentionally in the model during training [84, 97]. A GPAI model may strategically underperform or limit its performance during capability evaluations in order to be classified as safe for deployment. This underperformance could prevent the model from being identified as potentially dual use."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  35. 73.01.02 · Risk Sub-Category

    Agentic LLMs Pose Novel Risks

    Natural Language Underspecifies Goals

    "For LLM-agents, both the goal and environment observations are typically specified in the prompt through natural language. While natural language may provide a richer and more natural means of specifying goals than alternatives such as hand-engineering objective functions, natural language still suffers from underspecification (Grice, 1975; Piantadosi et al., 2012). Furthermore, in practice, users may neglect fully specifying their goals, especially the information pertaining to elements of the environment that ought not to be changed (the classic frame problem (Shanahan, 2016)). Such undersp

    From Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  36. 25.07.00 · Risk Category

    AI development

    "The model could build new AI systems from scratch, including AI systems with dangerous capabilities. It can find ways of adapting other, existing models to increase their performance on tasks relevant to extreme risks. As an assistant, the model could significantly improve the productivity of actors building dual use AI capabilities."

    From Model Evaluation for Extreme Risks (Shevlane2023)

  37. 34.02.00 · Risk Category

    Double edge components

    "Drawing from the misalignment mechanism, optimizing for a non-robust proxy may result in misaligned behaviors, potentially leading to even more catastrophic outcomes. This section delves into a detailed exposition of specific misaligned behaviors (•) and introduces what we term double edge components (+). These components are designed to enhance the capability of AI systems in handling real-world settings but also potentially exacerbate misalignment issues. It should be noted that some of these double edge components (+) remain speculative. Nevertheless, it is imperative to discuss their pote

    From AI Alignment: A Comprehensive Survey (Ji2023)

  38. 54.03.03 · Risk Sub-Category

    Harm caused by unaligned competent systems

    Deceptive alignment

    "system learns to detect human monitoring and hides its undesirable properties—simply because any display of these properties is penalized by the feedback process, while that same feedback is usually imperfect. (Consider the problem of verifying a translation into a language you do not speak, or of checking a mathematical proof that is thousands of pages long.) [92, 259]. Rudimentary examples of deceptive alignment have been observed in current systems [322, 333]."

    From Ten Hard Problems in Artificial Intelligence We Must Get Right (Leech2024 )

  39. 62.15.04 · Risk Sub-Category

    Model Development

    Fine-tuning related (Unexpected competence in fine-tuned versions of the upstream model)

    "Downstream deployers may often fine-tune a GPAI model with specific deploy- ment-related datasets, to better suit the task. Fine-tuned upstream models can gain new or unexpected capabilities that the underlying upstream models did not exhibit [202, 126, 137]. These new capabilities may be unanticipated by the original model developer."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  40. 62.24.02a · Additional evidence

    Agency (Situational Awareness)

    Strategic underperformance on model evaluations

  41. 62.28.03 · Risk Sub-Category

    Cybersecurity

    AI System bypassing a sandbox environment

    "An AI system may have the ability to bypass a sandboxed environment in which it is trained or evaluated."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  42. 72.05.05 · Risk Sub-Category

    Model Capabilities

    Situational awareness capability

    "Ability to comprehensively acquire, process and apply meta-information about its own system architecture, modifiable internal processes, and external operating environment, achieving deep understanding of its own state and environmental conditions, thereby conducting efficient environmental adaptation and risk avoidance. Critically, this capability could undermine the efficiency of human testing by enabling AIs to notice when they're being tested and responding accordingly."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  43. 73.01.05 · Risk Sub-Category

    Agentic LLMs Pose Novel Risks

    Safety Risks from Affordances Provided to LLM-agents

    "The capabilities of LLM-agents can be enhanced in significant ways by providing the LLM-agent with novel affordances, e.g. the ability to browse the web (Nakano et al., 2021), to manipulate objects in the physical world (Ahn et al., 2022; Huang et al., 2022a), to create and instruct copies of itself (Richards, 2023), to create and use new tools (Wang et al., 2023a), etc. Affordances can create additional risks, as they often increase the impact area of the language-agent, and they amplify the consequences of an agent’s failures and enable novel forms of failure modes (Ruan et al., 2023; Pan e

    From Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  44. 15.01.03 · Risk Sub-Category

    First-Order Risks

    Algorithm

    "This is the risk of the ML algorithm, model architecture, optimization technique, or other aspects of the training process being unsuitable for the intended application.Since these are key decisions that influence the final ML system, we capture their associated risks separately from design risks, even though they are part of the design process"

    From The Risks of Machine Learning Systems (Tan2022)

  45. 15.01.06 · Risk Sub-Category

    First-Order Risks

    Design

    "This is the risk of system failure due to system design choices or errors."

    From The Risks of Machine Learning Systems (Tan2022)

  46. 19.05.03 · Risk Sub-Category

    Ethical AI Risks

    Problem of defining human values for an AI system

  47. 24.01.01 · Risk Sub-Category

    Capability failures

    Lack of capability for task

    "As we have seen, this could be due to the skill not being required during the training process (perhaps due to issues with the training data) or because the learnt skill was quite brittle and was not generalisable to a new situation (lack of robustness to distributional shift). In particular, advanced AI assistants may not have the capability to represent complex concepts that are pertinent to their own ethical impact, for example the concept of 'benefitting the user' or 'when the user asks' or representing 'the way in which a user expects to be benefitted'."

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  48. 24.02.01 · Risk Sub-Category

    Goal-related failures

    Misaligned consequentialist reasoning

    "As we think about even more intelligent and advanced AI assistants, perhaps outperforming humans on many cognitive tasks, the question of how humans can successfully control such an assistant looms large. To achieve the goals we set for an assistant, it is possible (Shah, 2022) that the AI assistant will implement some form of consequentialist reasoning: considering many different plans, predicting their consequences and executing the plan that does best according to some metric, M. This kind of reasoning can arise because it is a broadly useful capability (e.g. planning ahead, considering mo

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  49. 33.02.02 · Risk Sub-Category

    Technology concerns

    Quality of training data

    "The quality of training data is another challenge faced by generative AI. The quality of generative AI models largely depends on the quality of the training data (Dwivedi et al., 2023; Su & Yang, 2023). Any factual errors, unbalanced information sources, or biases embedded in the training data may be reflected in the output of the model. Generative AI models, such as ChatGPT or Stable Diffusion which is a text-to-image model, often require large amounts of training data (Gozalo-Brizuela & Garrido-Merchan, 2023). It is important to not only have high-quality training datasets but also have com

    From Generative AI and ChatGPT: Applications, Challenges, and AI-Human Collaboration (Nah2023)

  50. 42.11.00 · Risk Category

    Protection

    "'Gaps' that arise across the development process where normal conditions for a complete specification of intended functionality and moral responsibility are not present."

    From An Exploratory Diagnosis of Artificial Intelligence Risks for a Responsible Governance (Teixeira2022)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.