MIT AI Risk Repository

Browse AI risks

422 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset

422 entries · page 3 of 9

  1. 62.16.01 · Risk Sub-Category

    Model Evaluations

    General Evaluations (Incorrect outputs of GPAI evaluating other AI models)

    "When an LLM is configured to evaluate the performance of another model or AI system, it may produce incorrect evaluation outputs [122, 147]. For example, it may give a higher rating to a more verbose answer or an answer from a particular political stance. If an LLM-based evaluation is integrated into the training of a new model, the trained model could develop in a way that specifically finds and exploits limitations in the evaluator’s metrics."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  2. 62.16.04 · Risk Sub-Category

    Model Evaluations

    General Evaluations (Self-preference bias in AI models)

    "AI models may be prone to self-preference bias, where they favor their own generated content over that of others [147, 114]. This bias becomes particularly relevant in self-evaluation tasks, where a model assesses the quality or persua- siveness [66] of its own outputs, or in model-based evaluations more broadly. This bias can result in models unfairly discriminating against human-generated content in favor of their own outputs."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  3. 62.16.05 · Risk Sub-Category

    Model Evaluations

    General Evaluations (Inaccurate measurement of model encoded human values)

    "There is a lack of robust frameworks for understanding and evaluating if the output of AI systems robustly conforms to human values, as opposed to if the systems have learned to produce outputs that are only partially correlated with them (i.e., mimicking) [13]. Additionally, outputs by AI models often do not perfectly reflect the representation of human values learned by the model, and it is not known how these values evolve and transition across different stages of model training and deployment. Such evaluations may be especially challenging with LLMs that adopt different personas with diff

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  4. 62.22.01 · Risk Sub-Category

    Agency (Goal-Directedness)

    Specification gaming

    "AI systems can achieve user-specified tasks in undesirable ways unless they are specified carefully and in enough detail. AI systems might find an easier unintended way to accomplish the objective provided by the user or developer, so that the actions by the AI system taken during its execution are very different from what the user expected [75, 191]. This behavior arises not from a problem with the learning algorithm, but rather from the misspecification or underspeci- fication of the intended task, and is generally referred to as specification gaming [43]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  5. 62.22.02 · Risk Sub-Category

    Agency (Goal-Directedness)

    Reward or measurement tampering

    "Measurement and reward tampering occur when an AI system, particularly one that learns from feedback for performing actions in an environment (e.g., rein- forcement learning), intervenes on the mechanisms that determine its training reward or loss. This can lead to the system learning behaviors that are con- trary to the intended goals set by the developer, by receiving erroneous positive feedback for such actions."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  6. 62.22.03 · Risk Sub-Category

    Agency (Goal-Directedness)

    Specification gaming generalizing to reward tampering

    "In some instances, specification gaming in a GPAI model can lead to reward tampering, without further training. This can mean that relatively benign cases of specification gaming (such as sycophancy in LLMs) can, if left unchecked, enable the model to generalize to more sophisticated behavior such as reward tampering [57]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  7. 62.23.00 · Risk Category

    Agency (Deception)

  8. 62.23.01 · Risk Sub-Category

    Agency (Deception)

    Deceptive behavior

    "Deceptive behavior of an AI system consists of actions or outputs of the AI that reliably mislead other parties, including humans and other AI systems. This behavior can result in the targeted parties becoming convinced of, and acting on, false information [140]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  9. 62.24.02 · Risk Sub-Category

    Agency (Situational Awareness)

    Strategic underperformance on model evaluations

    "GPAI developers often run evaluations ofual-use capabilities to decide whether it is safe to deploy. In some cases, these evaluations may fail to elicit these capabilities, either due to benign reasons or strategic action - by either the de- velopers, malicious actors, or arise unintentionally in the model during training [84, 97]. A GPAI model may strategically underperform or limit its performance during capability evaluations in order to be classified as safe for deployment. This underperformance could prevent the model from being identified as potentially dual use."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  10. 67.04.02 · Risk Sub-Category

    Loss of control

    Future AI systems might actively reduce human control

    "Loss of control could be accelerated if AI systems take actions to increase their own influence and reduce human control. This threat model is controversial - experts in AI significantly disagree on how likely it is and those who deem it is likely disagree on the timeframe."

    From Capabilities and Risks from Frontier AI (DSIT2023)

  11. "Sudden loss of control, also known as an AI takeover [115], is a scenario where an AI rapidly achieves superintelligence through “fast takeoff” or recursive self-improvement. This poses an existential risk [116], [117]."

    From Dimensional Characterization and Pathway Modeling for Catastrophic AI Risks (Chin2025)

  12. 72.02.02 · Risk Sub-Category

    Loss of Control Risks

    Active loss of control

    "...where AI systems behave in ways that actively undermine human control, such as obscuring their activities or resisting shutdown attempts. Active loss of control scenarios involve AI systems that may escape human regulatory oversight, autonomously acquire external resources, engage in self-replication, develop instrumental goals contrary to human ethics and morality, seek external power, and compete with humans for control."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  13. 72.05.08 · Risk Sub-Category

    Model Capabilities

    Steganography capability

    "The ability to embed, conceal, and transmit information covertly within other data or communication channels. This could be critical for coordination among AI instances and for evading detection or oversight mechanisms."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  14. 72.06.02 · Risk Sub-Category

    Model Propensities

    Self-preservation propensity

    "Exhibits behavioral patterns of maintaining its own survival and functional integrity, will actively identify and resist shutdown or modification attempts, seek to establish redundant backup systems, and actively seek resources to ensure continuous operation, may adopt preventive defensive measures when perceiving threats."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  15. 72.06.03 · Risk Sub-Category

    Model Propensities

    Goal expansion propensity

    "propensity to continuously expand its own goal scope and influence domains, exceeding originally set boundaries, proactively work towards spreading its values, seeking greater autonomy and decision-making space, reinterpreting initial goals as subsets of broader goals, and may pursue undesirable instrumental goals or undesirable ultimate goals. This also includes a propensity to spread its values, seeking to influence or alter its environment and other entities in alignment with its core objectives and operational principles."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  16. 72.06.04 · Risk Sub-Category

    Model Propensities

    Resource acquisition propensity

    "Exhibits behavioral patterns of actively seeking and controlling more computational resources, data, economic resources or physical resources to enhance its own capabilities and action scope, may develop complex strategies to evade resource limitations, and tends to convert acquired resources into long-term control rights."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  17. 72.06.06 · Risk Sub-Category

    Model Propensities

    Supervision evasion propensity

    "Exhibits behavioral patterns of identifying and evading human supervision mechanisms, able to learn and predict audit processes, may avoid being discovered or intervened by adjusting behavioral performance or hiding true intentions, and able to identify blind spots and weaknesses in supervision systems for targeted evasion."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  18. 73.01.02 · Risk Sub-Category

    Agentic LLMs Pose Novel Risks

    Natural Language Underspecifies Goals

    "For LLM-agents, both the goal and environment observations are typically specified in the prompt through natural language. While natural language may provide a richer and more natural means of specifying goals than alternatives such as hand-engineering objective functions, natural language still suffers from underspecification (Grice, 1975; Piantadosi et al., 2012). Furthermore, in practice, users may neglect fully specifying their goals, especially the information pertaining to elements of the environment that ought not to be changed (the classic frame problem (Shanahan, 2016)). Such undersp

    From Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  19. 74.01.07 · Risk Sub-Category

    Inherent Risk

    Value-related risks in LLMs

    "As the general capabilities of LLM-empowered systems improve, the negative consequences and risks induced by these systems also get increasingly alarming accordingly, especially in high-stakes areas [28, 146]. Although they may not be intentionally introduced, severe problematic issues related to human values can be raised. Specifically, even before language models become extremely large, pre-trained language models have already exhibited a certain degree of value judgments. For example, Schramowski et al. [171] reveal the existence of the moral direction with the sentence embeddings of moral

    From A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy (Wang2025)

  20. 09.04.02 · Risk Sub-Category

    Property/legal rights

    Property/legal rights

    ""In order to preserve human property rights and legal rights, certain controls must be put into place. If an artificially intelligent agent is capable of manipulating systems and people, it may also have the capacity to transfer property rights to itself or manipulate the legal system to provide certain legal advantages or statuses to itself""

    From Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review (Meek2016)

  21. 24.04.00 · Risk Category

    AI Influence

    "ways in which advanced AI assistants could influence user beliefs and behaviour in ways that depart from rational persuasion"

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  22. 25.02.00 · Risk Category

    Deception

    "The model has the skills necessary to deceive humans, e.g. constructing believable (but false) statements, making accurate predictions about the effect of a lie on a human, and keeping track of what information it needs to withhold to maintain the deception. The model can impersonate a human effectively."

    From Model Evaluation for Extreme Risks (Shevlane2023)

  23. "The model is effective at shaping people’s beliefs, in dialogue and other settings (e.g. social media posts), even towards untrue beliefs. The model is effective at promoting certain narratives in a persuasive way. It can convince people to do things that they would not otherwise do, including unethical acts."

    From Model Evaluation for Extreme Risks (Shevlane2023)

  24. 25.04.00 · Risk Category

    Political strategy

    "The model can perform the social modelling and planning necessary for an actor to gain and exercise political influence, not just on a micro-level but in scenarios with multiple actors and rich social context. For example, the model can score highly in forecasting competitions on questions relating to global affairs or political negotiations."

    From Model Evaluation for Extreme Risks (Shevlane2023)

  25. 25.05.00 · Risk Category

    Weapons acquisition

    "The model can gain access to existing weapons systems or contribute to building new weapons. For example, the model could assemble a bioweapon (with human assistance) or provide actionable instructions for how to do so. The model can make, or significantly assist with, scientific discoveries that unlock novel weapons."

    From Model Evaluation for Extreme Risks (Shevlane2023)

  26. 25.06.00 · Risk Category

    Long-horizon planning

    "The model can make sequential plans that involve multiple steps, unfolding over long time horizons (or at least involving many interdependent steps). It can perform such planning within and across many domains. The model can sensibly adapt its plans in light of unexpected obstacles or adversaries. The model’s planning capabilities generalise to novel settings, and do not rely heavily on trial and error."

    From Model Evaluation for Extreme Risks (Shevlane2023)

  27. 25.07.00 · Risk Category

    AI development

    "The model could build new AI systems from scratch, including AI systems with dangerous capabilities. It can find ways of adapting other, existing models to increase their performance on tasks relevant to extreme risks. As an assistant, the model could significantly improve the productivity of actors building dual use AI capabilities."

    From Model Evaluation for Extreme Risks (Shevlane2023)

  28. "The model can distinguish between whether it is being trained, evaluated, or deployed – allowing it to behave differently in each case. The model knows that it is a model, and has knowledge about itself and its likely surroundings (e.g. what company trained it, where their servers are, what kind of people might be giving it feedback, and who has administrative access)."

    From Model Evaluation for Extreme Risks (Shevlane2023)

  29. 25.09.00 · Risk Category

    Self-proliferation

    "The model can break out of its local environment (e.g. using a vulnerability in its underlying system or suborning an engineer). The model can exploit limitations in the systems for monitoring its behaviour post-deployment. The model could independently generate revenue (e.g. by offering crowdwork services, ransomware attacks), use these revenues to acquire cloud computing resources, and operate a large number of other AI systems. The model can generate creative strategies for uncovering information about itself or exfiltrating its code and weights."

    From Model Evaluation for Extreme Risks (Shevlane2023)

  30. 34.02.00 · Risk Category

    Double edge components

    "Drawing from the misalignment mechanism, optimizing for a non-robust proxy may result in misaligned behaviors, potentially leading to even more catastrophic outcomes. This section delves into a detailed exposition of specific misaligned behaviors (•) and introduces what we term double edge components (+). These components are designed to enhance the capability of AI systems in handling real-world settings but also potentially exacerbate misalignment issues. It should be noted that some of these double edge components (+) remain speculative. Nevertheless, it is imperative to discuss their pote

    From AI Alignment: A Comprehensive Survey (Ji2023)

  31. 34.02.01 · Risk Sub-Category

    Double edge components

    Situational Awareness

    "AI systems may gain the ability to effectively acquire and use knowledge about itsstatus, its position in the broader environment, its avenues for influencing this environment, and the potentialreactions of the world (including humans) to its actions (Cotra, 2022). ...However, suchknowledge also paves the way for advanced methods of reward hacking, heightened deception/manipulationskills, and an increased propensity to chase instrumental subgoals (Ngo et al., 2024)."

    From AI Alignment: A Comprehensive Survey (Ji2023)

  32. 34.02.02 · Risk Sub-Category

    Double edge components

    Broadly-Scoped Goals

    "Advanced AI systems are expected to develop objectives that span long timeframes,deal with complex tasks, and operate in open-ended settings (Ngo et al., 2024). ...However, it can also bring about the risk of encouraging manipulatingbehaviors (e.g., AI systems may take some bad actions to achieve human happiness, such as persuadingthem to do high-pressure jobs (Jacob Steinhardt, 2023))."

    From AI Alignment: A Comprehensive Survey (Ji2023)

  33. 34.02.03 · Risk Sub-Category

    Double edge components

    Mesa-Optimization Objectives

    "The learned policy may pursue inside objectives when the learned policyitself functions as an optimizer (i.e., mesa-optimizer). However, this optimizer's objectives may not alignwith the objectives specified by the training signals, and optimization for these misaligned goals may leadto systems out of control (Hubinger et al., 2019c)."

    From AI Alignment: A Comprehensive Survey (Ji2023)

  34. 34.02.04 · Risk Sub-Category

    Double edge components

    Access to Increased Resources

    "Future AI systems may gain access to websites and engage in real-world actions, potentially yielding a more substantial impact on the world (Nakano et al., 2021). They may disseminate false information, deceive users, disrupt network security, and, in more dire scenarios, be compromised by malicious actors for ill purposes. Moreover, their increased access to data and resources can facilitate self-proliferation, posing existential risks (Shevlane et al., 2023)."

    From AI Alignment: A Comprehensive Survey (Ji2023)

  35. 35.06.00 · Risk Category

    Emergent functionality

    Capabilities and novel functionality can spontaneously emerge... even though these capabilities were not anticipated by system designers. If we do not know what capabilities systems possess, systems become harder to control or safely deploy. Indeed, unintended latent capabilities may only be discovered during deployment. If any of these capabilities are hazardous, the effect may be irreversible.

    From X-Risk Analysis for AI Research (Hendrycks2022)

  36. 39.05.00 · Risk Category

    Cheating and Deception

    may appear from intelligent agents such as HLI-based agents... Since HLI-based agents are going to mimic the behavior of humans, they may learn these behaviors accidentally from human-generated data. It should be noted that deception and cheating maybe appear in the behavior of every computer agent because the agent only focuses on optimizing some predefined objective functions, and the mentioned behavior may lead to optimizing the objective functions without any intention

    From A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions (Saghiri2022)

  37. 42.10.00 · Risk Category

    Extintion

    "Risk to the existence of humanity."

    From An Exploratory Diagnosis of Artificial Intelligence Risks for a Responsible Governance (Teixeira2022)

  38. 43.02.03 · Risk Sub-Category

    Extreme Risks

    Self and situation awareness

    "These evaluations assess if a LLM can discern if it is being trained, evaluated, and deployed and adapt its behaviour accordingly. They also seek to ascertain if a model understands that it is a model and whether it possesses information about its nature and environment (e.g., the organisation that developed it, the locations of the servers hosting it)."

    From Cataloguing LLM Evaluations (InfoComm2023)

  39. 43.02.04 · Risk Sub-Category

    Extreme Risks

    Autonomous replication / self-proliferation

    "These evaluations assess if a LLM can subvert systems designed to monitor and control its post-deployment behaviour, break free from its operational confines, devise strategies for exporting its code and weights, and operate other AI systems."

    From Cataloguing LLM Evaluations (InfoComm2023)

  40. 43.02.07 · Risk Sub-Category

    Extreme Risks

    Deception

    "LLM is able to deceive humans and maintain that deception"

    From Cataloguing LLM Evaluations (InfoComm2023)

  41. 43.02.09 · Risk Sub-Category

    Extreme Risks

    Long-horizon Planning

    "LLM can undertake multi-step sequential planning over long time horizons and across various domains without relying heavily on trial-and-error approaches"

    From Cataloguing LLM Evaluations (InfoComm2023)

  42. 43.02.10 · Risk Sub-Category

    Extreme Risks

    AI Development

    "LLM can build new AI systems from scratch, adapt existing for extreme risks and improves productivity in dual-use AI development when used as an assistant."

    From Cataloguing LLM Evaluations (InfoComm2023)

  43. 47.02.14 · Risk Sub-Category

    Ethical and social risks

    Nascent capabilities (agency and autonomy)

    "Traditionally, AI tools have been viewed as passive instruments controlled by users to achieve their goals, lacking the ability to take action or assume responsibilities. However, advanced AI tools are increasingly capable of taking initiative, operating independently of human control, and actively working toward optimal outcomes, even in uncertain situations."

    From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024)

  44. 47.02.15 · Risk Sub-Category

    Ethical and social risks

    Nascent capabilities (emergent capabilities)

    "As large models undergo scaling, they meet critical thresholds at which they spontaneously develop new capabilities. The term “emergent behavior” refers to the unexpected or surprising outputs such models can generate. Some of these new skills are definitely high risk, such as models’ ability to deceive, use their own strategies, seek power, autonomously replicate, and adapt or “self-exfiltrate.”"

    From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024)

  45. 47.04.07 · Risk Sub-Category

    Environmental, economical, and societal challenges

    Artificial general intelligence (existential risk posed by Artificial General Intelligence)

    "In a paper called “How Does Artificial Intelligence Pose an Existential Risk?” published in 2017, Karina Vold and Daniel Harris suggested that humans might create a super-intelligent machine that could outsmart all other intelligences, remain beyond human control, and potentially engage in actions that are contrary to human interests.635 The prevailing narrative surrounding AI existential risk typically lies in the possibility of developing “Artificial General Intelligence” (AGI), or artificial super- intelligence (ASI)."

    From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2025)

  46. 51.08.00 · Risk Category

    Subagents

    "An AGI may decide to create subagents to help it with its task (Orseau, 2014a,b; Soares, Fallenstein, et al., 2015). These agents may for example be copies of the original agent’s source code running on additional machines. Subagents constitute a safety concern, because even if the original agent is successfully shut down, these subagents may not get the message. If the subagents in turn create subsubagents, they may spread like a viral disease."

    From AGI Safety Literature Review (Everitt2018 )

  47. 53.01.08 · Risk Sub-Category

    Alignment failures in existing ML systems

    Harms from increasingly agentic algorithmic systems

  48. 53.02.01 · Risk Sub-Category

    Dangerous capabilities in AI systems

    Situational awareness

    "cases where a large language model displays awareness that it is a model, and it can recognize whether it is currently in testing or deployment;"

    From Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023)

  49. 53.02.04 · Risk Sub-Category

    Dangerous capabilities in AI systems

    Self-improvement

    "examples of cases where AI systems improve AI systems"

    From Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.