MIT AI Risk Repository

Browse AI risks

662 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset

662 entries · page 8 of 14

  1. 11.05.05 · Risk Sub-Category

    Societal System Harms

    Environmental harms

    depletion or contamination of natural resources, and damage to built environments... that may occur throughout the lifecycle of digital technologies [170, 237] from “crale (mining) to usage (consumption) to grave (waste)”

    From Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction (Shelby2023)

  2. 15.02.05 · Risk Sub-Category

    Second-Order Risks

    Environmental

    The risk of harm to the natural environment posed by the ML system.

    From The Risks of Machine Learning Systems (Tan2022)

  3. 16.06.01 · Risk Sub-Category

    Risk area 6: Environmental and Socioeconomic harms

    Environmental harms from operating LMs

    "LMs (and AI more broadly) can have an environmental impact at different levels, including: (1) direct impacts from the energy used to train or operate the LM, (2) secondary impacts due to emissions from LM-based applications, (3) system-level impacts as LM-based applications influence human behaviour (e.g. increasing environmental awareness or consumption), and (4) resource impacts on precious metals and other materials required to build hardware on which the computations are run e.g. data centres, chips, or devices. Some evidence exists on (1), but (2) and (3) will likely be more significant

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  4. 17.06.01 · Risk Sub-Category

    Automation, Access and Environmental Harms

    Environmental harms from operation LMs

    "Large-scale machine learning models, including LMs, have the potential to create significant environmental costs via their energy demands, the associated carbon emissions for training and operating the models, and the demand for fresh water to cool the data centres where computations are run (Mytton, 2021; Patterson et al., 2021)."

    From Ethical and social risks of harm from language models (Weidinger2021)

  5. 18.06.02 · Risk Sub-Category

    Socioeconomic and environmental harms

    Environmental damage

    "Creating negative environmental impacts though model development and deployment"

    From Sociotechnical Safety Evaluation of Generative AI Systems (Weidinger2023)

  6. "the growing field of generative AI, which brings with it direct and severe impacts on our climate: generative AI comes with a high carbon footprint and similarly high resource price tag, which largely flies under the radar of public AI discourse. Training and running generative AI tools requires companies to use extreme amounts of energy and physical resources. Training one natural language processing model with normal tuning and experiments emits, on average, the same amount of carbon that seven people do over an entire year.121'

    From Generating Harms - Generative AI's impact and paths forwards (EPIC2023)

  7. 32.04.00 · Risk Category

    Environmental impacts

    Environmental harm, Sustainability

    From The Ethics of ChatGPT – Exploring the Ethical Issues of an Emerging Technology (Stahl2024)

  8. 39.02.00 · Risk Category

    Energy Consumption

    Some learning algorithms, including deep learning, utilize iterative learning processes [23]. This approach results in high energy consumption.

    From A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions (Saghiri2022)

  9. 41.06.00 · Risk Category

    Environment

    "AI is already helping to combat the impact of climate change with smart technology and sensors reducing emissions. However, it is also a key component in the development of nanobots, which could have dangerous environmental impacts by invisibly modifying substances at nanoscale."

    From The Rise of Artificial Intelligence - Future Outlooks and Emerging Risks (Allianz2018)

  10. 41.06.01 · Risk Sub-Category

    Environment

    Accelerated development of nanotechnology produces uncontrolled production of toxic nanoparticles

    "AI is a key component for the development of nanobots, which could have dangerous environmental implications by invisibly modifying substances at nanoscale. For example, nanobots could start chemical reactions that would create invisible nanoparticles that are toxic and potentially lethal."

    From The Rise of Artificial Intelligence - Future Outlooks and Emerging Risks (Allianz2018)

  11. 44.03.02 · Risk Sub-Category

    Unintentional: direct

    AI harms animals due to mistake or misadventure in the way the AI operates in practice

  12. "AI impacts human or ecological systems in ways that ultimately harm animals"

    From Harm to Nonhuman Animals from AI: a Systematic Account and Framework (Coghlan2023 )

  13. 44.04.01 · Risk Sub-Category

    Unintentional: indirect

    Indirect Material Harms

    "AI proliferation causes harm to the environment through energy use and e-waste thereby destroying animal habitat"

    From Harm to Nonhuman Animals from AI: a Systematic Account and Framework (Coghlan2023 )

  14. 44.04.02 · Risk Sub-Category

    Unintentional: indirect

    Harms from Estrangement

    "Replacement by AI of human observation and interaction leads to neglect of certain interests"

    From Harm to Nonhuman Animals from AI: a Systematic Account and Framework (Coghlan2023 )

  15. 44.04.03 · Risk Sub-Category

    Unintentional: indirect

    Epistemic Harms

    "Algorithmic recommender systems reinforce and amplify anthropocentric bias or desire of some people for animal cruelty as entertainment — leading to greater harm to animals through reinforcement of meat eating from factory farms, cruel uses of animals for entertainment, etc"

    From Harm to Nonhuman Animals from AI: a Systematic Account and Framework (Coghlan2023 )

  16. 54.01.02 · Risk Sub-Category

    Negative impacts of AI use

    Environmental cost

    "Large-scale DL systems can produce signicant carbon emissions as a result of the computational demands of training runs and inference [539]"

    From Ten Hard Problems in Artificial Intelligence We Must Get Right (Leech2024 )

  17. 58.09.00 · Risk Category

    Environmental

    "Environmental - Damage to the environment directly or indirectly caused by a technology system or set of systems."

    From A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms (Abercrombie2024)

  18. 58.09.08 · Risk Sub-Category

    Environmental

    Pollution

    "Pollution - Actual or potential pollution to the air, ground, noise, or water caused by a technology system."

    From A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms (Abercrombie2024)

  19. 60.03.04 · Risk Sub-Category

    Systemic risks

    Risks to the environment

    "General- purpose AI is a moderate but rapidly growing contributor to global environmental impacts through energy use and greenhouse gas (GHG) emissions. Current estimates indicate that data centres and data transmission account for an estimated 1% of global energy- related GHG emissions, with AI consuming 10–28% of data centre energy capacity. AI energy demand is expected to grow substantially by 2026, with some estimates projecting a doubling or more, driven primarily by general-purpose AI systems such as language models."

    From International AI Safety Report 2025 (Bengio2025)

  20. "The impact of AI on the environment, including risks related to climate change and pollution."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  21. 65.23.06 · Risk Sub-Category

    Non-technical risks (Societal impact)

    Impact on the environment

    "AI, and large generative models in particular, might produce increased carbon emissions and increase water usage for their training and operation."

    From AI Risk Atlas (IBM2025)

  22. 66.12.01 · Risk Sub-Category

    Environment

    Pollution

    "Actual or potential pollution to the air, ground, noise, or water caused by a technology system"

    From A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025)

  23. "One of the most likely approaches to creating superintelligent AI is by growing it from a seed (baby) AI via recursive self-improvement (RSI) (Nijholt 2011). One danger in such a scenario is that the system can evolve to become self-aware, free-willed, independent or emotional, and obtain a number of other emergent properties, which may make it less likely to abide by any built-in rules or regulations and to instead pursue its own goals possibly to the detriment of humanity."

    From Taxonomy of Pathways to Dangerous Artificial Intelligence (Yampolskiy2016)

  24. "Previous research has shown that utility maximizing agents are likely to fall victims to the same indulgences we frequently observe in people, such as addictions, pleasure drives (Majot and Yampolskiy 2014), self-delusions and wireheading (Yampolskiy 2014). In general, what we call mental illness in people, particularly sociopathy as demonstrated by lack of concern for others, is also likely to show up in artificial minds."

    From Taxonomy of Pathways to Dangerous Artificial Intelligence (Yampolskiy2016)

  25. 05.02.00 · Risk Category

    Safety

    A primary concern is the emergence of human-level or superhuman generative models, commonly referred to as AGI, and their potential existential or catastrophic risks to humanity. Connected to that, AI safety aims at avoiding deceptive or power-seeking machine behavior, model self-replication, or shutdown evasion. Ensuring controllability, human oversight, and the implementation of red teaming measures are deemed to be essential in mitigating these risks, as is the need for increased AI safety research and promoting safety cultures within AI organizations instead of fueling the AI race. Further

    From Mapping the Ethics of Generative AI: A Comprehensive Scoping Review (Hagendorff2024)

  26. "Sometimes an AI finds ways to achieve its given goals in ways that are completely different from what its creators had in mind."

    From A framework for ethical Ai at the United Nations (Hogenhout2021)

  27. 07.03.00 · Risk Category

    Agential

    "While there are multiple types of intelligent agents, goal-based, utility-maximizing, and learning agents are the primary concern and the focus of this research"

    From Examining the differential risk from high-level artificial intelligence and the question of control (Kilian2023)

  28. "A sufficiently intelligent AI could possess the ability to subtly influence societal behaviors through a sophisticated understanding of human nature"

    From Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review (Meek2016)

  29. "The degree of automation and control describes the extent to which an AI system functions independently of human supervision and control."

    From Sources of Risk of AI Systems (Steimers2022)

  30. 15.01.09 · Risk Sub-Category

    First-Order Risks

    Emergent behavior

    "This is the risk resulting from novel behavior acquired through continual learning or self-organization after deployment."

    From The Risks of Machine Learning Systems (Tan2022)

  31. "AI systems compromising human agency, or circumventing meaningful human control"

    From Sociotechnical Safety Evaluation of Generative AI Systems (Weidinger2023)

  32. 18.05.02 · Risk Sub-Category

    Human Autonomy and Intregrity Harms

    Persuasion and manipulation

    "Exploiting user trust, or nudging or coercing them into performing certain actions against their will (c.f. Burtell and Woodside (2023); Kenton et al. (2021))"

    From Sociotechnical Safety Evaluation of Generative AI Systems (Weidinger2023)

  33. 22.04.00 · Risk Category

    Rogue AIs (Internal)

    "speculative technical mechanisms that might lead to rogue AIs and how a loss of control could bring about catastrophe"

    From An Overview of Catastrophic AI Risks (Hendrycks2023)

  34. 22.04.01 · Risk Sub-Category

    Rogue AIs (Internal)

    Proxy Gaming

    "One way we might lose control of an AI agent’s actions is if it engages in behavior known as “proxy gaming.” It is often difficult to specify and measure the exact goal that we want a system to pursue. Instead, we give the system an approximate—“proxy”—goal that is more measurable and seems likely to correlate with the intended goal. However, AI systems often find loopholes by which they can easily achieve the proxy goal, but completely fail to achieve the ideal goal. If an AI “games” its proxy goal in a way that does not reflect our values, then we might not be able to reliably steer its beh

    From An Overview of Catastrophic AI Risks (Hendrycks2023)

  35. 22.04.02 · Risk Sub-Category

    Rogue AIs (Internal)

    Goal Drift

    "Even if we successfully control early AIs and direct them to promote human values, future AIs could end up with different goals that humans would not endorse. This process, termed “goal drift,” can be hard to predict or control. This section is most cutting-edge and the most speculative, and in it we will discuss how goals shift in various agents and groups and explore the possibility of this phenomenon occurring in AIs. We will also examine a mechanism that could lead to unexpected goal drift, called intrinsification, and discuss how goal drift in AIs could be catastrophic."

    From An Overview of Catastrophic AI Risks (Hendrycks2023)

  36. 22.04.03 · Risk Sub-Category

    Rogue AIs (Internal)

    Power Seeking

    "even if an agent started working to achieve an unintended goal, this would not necessarily be a problem, as long as we had enough power to prevent any harmful actions it wanted to attempt. Therefore, another important way in which we might lose control of AIs is if they start trying to obtain more power, potentially transcending our own."

    From An Overview of Catastrophic AI Risks (Hendrycks2023)

  37. 22.04.04 · Risk Sub-Category

    Rogue AIs (Internal)

    Deception

    "it is plausible that AIs could learn to deceive us. They might, for example, pretend to be acting as we want them to, but then take a “treacherous turn” when we stop monitoring them, or when they have enough power to evade our attempts to interfere with them. "

    From An Overview of Catastrophic AI Risks (Hendrycks2023)

  38. 24.02.00 · Risk Category

    Goal-related failures

    "As we think about even more intelligent and advanced AI assistants, perhaps outperforming humans on many cognitive tasks, the question of how humans can successfully control such an assistant looms large. To achieve the goals we set for an assistant, it is possible (Shah, 2022) that the AI assistant will implement some form of consequentialist reasoning: considering many different plans, predicting their consequences and executing the plan that does best according to some metric, M. This kind of reasoning can arise because it is a broadly useful capability (e.g. planning ahead, considering mo

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  39. 24.02.02 · Risk Sub-Category

    Goal-related failures

    Specification gaming

    "Specification gaming (Krakovna et al., 2020) occurs when some faulty feedback is provided to the assistant in the training data (i.e. the training objective O does not fully capture what the user/designer wants the assistant to do). It is typified by the sort of behaviour that exploits loopholes in the task specification to satisfy the literal specification of a goal without achieving the intended outcome."

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  40. 24.02.03 · Risk Sub-Category

    Goal-related failures

    Goal misgeneralisation

    "In the problem of goal misgeneralisation (Langosco et al., 2023; Shah et al., 2022), the AI system's behaviour during out-of-distribution operation (i.e. not using input from the training data) leads it to generalise poorly about its goal while its capabilities generalise well, leading to undesired behaviour. Applied to the case of an advanced AI assistant, this means the system would not break entirely – the assistant might still competently pursue some goal, but it would not be the goal we had intended."

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  41. 24.02.04 · Risk Sub-Category

    Goal-related failures

    Deceptive alignment

    "Here, the agent develops its own internalised goal, G, which is misgeneralised and distinct from the training reward, R. The agent also develops a capability for situational awareness (Cotra, 2022): it can strategically use the information about its situation (i.e. that it is an ML model being trained using a particular training setup, e.g. RL fine-tuning with training reward, R) to its advantage. Building on these foundations, the agent realises that its optimal strategy for doing well at its own goal G is to do well on R during training and then pursue G at deployment – it is only doing wel

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  42. 24.09.00 · Risk Category

    Cooperation

    "" AI assistants will need to coordinate with other AI assistants and with humans other than their principal users. This chapter explores the societal risks associated with the aggregate impact of AI assistants whose behaviour is aligned to the interests of particular users. For example, AI assistants may face collective action problems where the best outcomes overall are realised when AI assistants cooperate but where each AI assistant can secure an additional benefit for its user if it defects while others cooperate""

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  43. 34.01.01 · Risk Sub-Category

    Causes of Misalignment

    Reward Hacking

    "Reward Hacking: In practice, proxy rewards are often easy to optimize and measure, yet they frequently fall shortof capturing the full spectrum of the actual rewards (Pan et al., 2021). This limitation is denoted as misspecifiedrewards. The pursuit of optimization based on such misspecified rewards may lead to a phenomenon knownas reward hacking, wherein agents may appear highly proficient according to specific metrics but fall short whenevaluated against human standards (Amodei et al., 2016; Everitt et al., 2017). The discrepancy between proxyrewards and true rewards often manifests as a sha

    From AI Alignment: A Comprehensive Survey (Ji2023)

  44. 34.01.02 · Risk Sub-Category

    Causes of Misalignment

    Goal Misgeneralization

    "Goal Misgeneralization: Goal misgeneralization is another failure mode, wherein the agent actively pursuesobjectives distinct from the training objectives in deployment while retaining the capabilities it acquired duringtraining (Di Langosco et al., 2022). For instance, in CoinRun games, the agent frequently prefers reachingthe end of a level, often neglecting relocated coins during testing scenarios. Di Langosco et al. (2022) drawattention to the fundamental disparity between capability generalization and goal generalization, emphasizing howthe inductive biases inherent in the model and its

    From AI Alignment: A Comprehensive Survey (Ji2023)

  45. 34.01.03 · Risk Sub-Category

    Causes of Misalignment

    Reward Tampering

    "Reward tampering can be considered a special case of reward hacking (Everitt et al., 2021; Skalse et al., 2022),referring to AI systems corrupting the reward signals generation process (Ring and Orseau, 2011). Everitt et al.(2021) delves into the subproblems encountered by RL agents: (1) tampering of reward function, where the agentinappropriately interferes with the reward function itself, and (2) tampering of reward function input, which entailscorruption within the process responsible for translating environmental states into inputs for the reward function.When the reward function is formu

    From AI Alignment: A Comprehensive Survey (Ji2023)

  46. 34.03.00 · Risk Category

    Misaligned Behaviors

  47. 34.03.01 · Risk Sub-Category

    Misaligned Behaviors

    Power-Seeking Behaviors

    "AI systems may exhibit behaviors that attempt to gain control over resourcesand humans and then exert that control to achieve its assigned goal (Carlsmith, 2022). The intuitive reasonwhy such behaviors may occur is the observation that for almost any optimization objective (e.g., investmentreturns), the optimal policy to maximize that quantity would involve power-seeking behaviors (e.g.,manipulating the market), assuming the absence of solid safety and morality constraints."

    From AI Alignment: A Comprehensive Survey (Ji2023)

  48. 34.03.02 · Risk Sub-Category

    Misaligned Behaviors

    Untruthful Output

    "AI systems such as LLMs can produce either unintentionally or deliberately inaccurateoutput. Such untruthful output may diverge from established resources or lack verifiability, commonly referredto as hallucination (Bang et al., 2023; Zhao et al., 2023). More concerning is the phenomenon wherein LLMsmay selectively provide erroneous responses to users who exhibit lower levels of education (Perez et al.,2023)."

    From AI Alignment: A Comprehensive Survey (Ji2023)

  49. 34.03.03 · Risk Sub-Category

    Misaligned Behaviors

    Deceptive Alignment & Manipulation

    "Manipulation & Deceptive Alignment is a class of behaviors thatexploit the incompetence of human evaluators or users (Hubinger et al., 2019a; Carranza et al., 2023) andeven manipulate the training process through gradient hacking (Richard Ngo, 2022). These behaviors canpotentially make detecting and addressing misaligned behaviors much harder.Deceptive Alignment: Misaligned AI systems may deliberately mislead their human supervisors instead of adhering to the intended task. Such deceptive behavior has already manifested in AI systems that employ evolutionary algorithms (Wilke et al., 2001; He

    From AI Alignment: A Comprehensive Survey (Ji2023)

  50. 34.03.04 · Risk Sub-Category

    Misaligned Behaviors

    Collectively Harmful Behaviors

    "AI systems have the potential to take actions that are seemingly benignin isolation but become problematic in multi-agent or societal contexts. Classical game theory offers simplistic models for understanding these behaviors. For instance, Phelps and Russell (2023) evaluates GPT-3.5's performance in the iterated prisoner's dilemma and other social dilemmas, revealing limitations in themodel's cooperative capabilities."

    From AI Alignment: A Comprehensive Survey (Ji2023)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.