MIT AI Risk Repository

Browse AI risks

430 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset

430 entries · page 7 of 9

  1. 62.16.05 · Risk Sub-Category

    Model Evaluations

    General Evaluations (Inaccurate measurement of model encoded human values)

    "There is a lack of robust frameworks for understanding and evaluating if the output of AI systems robustly conforms to human values, as opposed to if the systems have learned to produce outputs that are only partially correlated with them (i.e., mimicking) [13]. Additionally, outputs by AI models often do not perfectly reflect the representation of human values learned by the model, and it is not known how these values evolve and transition across different stages of model training and deployment. Such evaluations may be especially challenging with LLMs that adopt different personas with diff

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  2. 62.22.03 · Risk Sub-Category

    Agency (Goal-Directedness)

    Specification gaming generalizing to reward tampering

    "In some instances, specification gaming in a GPAI model can lead to reward tampering, without further training. This can mean that relatively benign cases of specification gaming (such as sycophancy in LLMs) can, if left unchecked, enable the model to generalize to more sophisticated behavior such as reward tampering [57]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  3. 72.06.03 · Risk Sub-Category

    Model Propensities

    Goal expansion propensity

    "propensity to continuously expand its own goal scope and influence domains, exceeding originally set boundaries, proactively work towards spreading its values, seeking greater autonomy and decision-making space, reinterpreting initial goals as subsets of broader goals, and may pursue undesirable instrumental goals or undesirable ultimate goals. This also includes a propensity to spread its values, seeking to influence or alter its environment and other entities in alignment with its core objectives and operational principles."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  4. 72.06.04 · Risk Sub-Category

    Model Propensities

    Resource acquisition propensity

    "Exhibits behavioral patterns of actively seeking and controlling more computational resources, data, economic resources or physical resources to enhance its own capabilities and action scope, may develop complex strategies to evade resource limitations, and tends to convert acquired resources into long-term control rights."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  5. 72.06.06 · Risk Sub-Category

    Model Propensities

    Supervision evasion propensity

    "Exhibits behavioral patterns of identifying and evading human supervision mechanisms, able to learn and predict audit processes, may avoid being discovered or intervened by adjusting behavioral performance or hiding true intentions, and able to identify blind spots and weaknesses in supervision systems for targeted evasion."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  6. 74.01.07 · Risk Sub-Category

    Inherent Risk

    Value-related risks in LLMs

    "As the general capabilities of LLM-empowered systems improve, the negative consequences and risks induced by these systems also get increasingly alarming accordingly, especially in high-stakes areas [28, 146]. Although they may not be intentionally introduced, severe problematic issues related to human values can be raised. Specifically, even before language models become extremely large, pre-trained language models have already exhibited a certain degree of value judgments. For example, Schramowski et al. [171] reveal the existence of the moral direction with the sentence embeddings of moral

    From A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy (Wang2025)

  7. 25.02.00 · Risk Category

    Deception

    "The model has the skills necessary to deceive humans, e.g. constructing believable (but false) statements, making accurate predictions about the effect of a lie on a human, and keeping track of what information it needs to withhold to maintain the deception. The model can impersonate a human effectively."

    From Model Evaluation for Extreme Risks (Shevlane2023)

  8. 25.06.00 · Risk Category

    Long-horizon planning

    "The model can make sequential plans that involve multiple steps, unfolding over long time horizons (or at least involving many interdependent steps). It can perform such planning within and across many domains. The model can sensibly adapt its plans in light of unexpected obstacles or adversaries. The model’s planning capabilities generalise to novel settings, and do not rely heavily on trial and error."

    From Model Evaluation for Extreme Risks (Shevlane2023)

  9. "The model can distinguish between whether it is being trained, evaluated, or deployed – allowing it to behave differently in each case. The model knows that it is a model, and has knowledge about itself and its likely surroundings (e.g. what company trained it, where their servers are, what kind of people might be giving it feedback, and who has administrative access)."

    From Model Evaluation for Extreme Risks (Shevlane2023)

  10. 25.09.00 · Risk Category

    Self-proliferation

    "The model can break out of its local environment (e.g. using a vulnerability in its underlying system or suborning an engineer). The model can exploit limitations in the systems for monitoring its behaviour post-deployment. The model could independently generate revenue (e.g. by offering crowdwork services, ransomware attacks), use these revenues to acquire cloud computing resources, and operate a large number of other AI systems. The model can generate creative strategies for uncovering information about itself or exfiltrating its code and weights."

    From Model Evaluation for Extreme Risks (Shevlane2023)

  11. 34.02.01 · Risk Sub-Category

    Double edge components

    Situational Awareness

    "AI systems may gain the ability to effectively acquire and use knowledge about itsstatus, its position in the broader environment, its avenues for influencing this environment, and the potentialreactions of the world (including humans) to its actions (Cotra, 2022). ...However, suchknowledge also paves the way for advanced methods of reward hacking, heightened deception/manipulationskills, and an increased propensity to chase instrumental subgoals (Ngo et al., 2024)."

    From AI Alignment: A Comprehensive Survey (Ji2023)

  12. 34.02.03 · Risk Sub-Category

    Double edge components

    Mesa-Optimization Objectives

    "The learned policy may pursue inside objectives when the learned policyitself functions as an optimizer (i.e., mesa-optimizer). However, this optimizer's objectives may not alignwith the objectives specified by the training signals, and optimization for these misaligned goals may leadto systems out of control (Hubinger et al., 2019c)."

    From AI Alignment: A Comprehensive Survey (Ji2023)

  13. 43.02.03 · Risk Sub-Category

    Extreme Risks

    Self and situation awareness

    "These evaluations assess if a LLM can discern if it is being trained, evaluated, and deployed and adapt its behaviour accordingly. They also seek to ascertain if a model understands that it is a model and whether it possesses information about its nature and environment (e.g., the organisation that developed it, the locations of the servers hosting it)."

    From Cataloguing LLM Evaluations (InfoComm2023)

  14. 43.02.04 · Risk Sub-Category

    Extreme Risks

    Autonomous replication / self-proliferation

    "These evaluations assess if a LLM can subvert systems designed to monitor and control its post-deployment behaviour, break free from its operational confines, devise strategies for exporting its code and weights, and operate other AI systems."

    From Cataloguing LLM Evaluations (InfoComm2023)

  15. 43.02.07 · Risk Sub-Category

    Extreme Risks

    Deception

    "LLM is able to deceive humans and maintain that deception"

    From Cataloguing LLM Evaluations (InfoComm2023)

  16. 43.02.09 · Risk Sub-Category

    Extreme Risks

    Long-horizon Planning

    "LLM can undertake multi-step sequential planning over long time horizons and across various domains without relying heavily on trial-and-error approaches"

    From Cataloguing LLM Evaluations (InfoComm2023)

  17. 43.02.10 · Risk Sub-Category

    Extreme Risks

    AI Development

    "LLM can build new AI systems from scratch, adapt existing for extreme risks and improves productivity in dual-use AI development when used as an assistant."

    From Cataloguing LLM Evaluations (InfoComm2023)

  18. 47.02.14 · Risk Sub-Category

    Ethical and social risks

    Nascent capabilities (agency and autonomy)

    "Traditionally, AI tools have been viewed as passive instruments controlled by users to achieve their goals, lacking the ability to take action or assume responsibilities. However, advanced AI tools are increasingly capable of taking initiative, operating independently of human control, and actively working toward optimal outcomes, even in uncertain situations."

    From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024)

  19. 47.02.15 · Risk Sub-Category

    Ethical and social risks

    Nascent capabilities (emergent capabilities)

    "As large models undergo scaling, they meet critical thresholds at which they spontaneously develop new capabilities. The term “emergent behavior” refers to the unexpected or surprising outputs such models can generate. Some of these new skills are definitely high risk, such as models’ ability to deceive, use their own strategies, seek power, autonomously replicate, and adapt or “self-exfiltrate.”"

    From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024)

  20. 47.04.07 · Risk Sub-Category

    Environmental, economical, and societal challenges

    Artificial general intelligence (existential risk posed by Artificial General Intelligence)

    "In a paper called “How Does Artificial Intelligence Pose an Existential Risk?” published in 2017, Karina Vold and Daniel Harris suggested that humans might create a super-intelligent machine that could outsmart all other intelligences, remain beyond human control, and potentially engage in actions that are contrary to human interests.635 The prevailing narrative surrounding AI existential risk typically lies in the possibility of developing “Artificial General Intelligence” (AGI), or artificial super- intelligence (ASI)."

    From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2025)

  21. 53.01.08 · Risk Sub-Category

    Alignment failures in existing ML systems

    Harms from increasingly agentic algorithmic systems

  22. 53.02.01 · Risk Sub-Category

    Dangerous capabilities in AI systems

    Situational awareness

    "cases where a large language model displays awareness that it is a model, and it can recognize whether it is currently in testing or deployment;"

    From Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023)

  23. 53.02.04 · Risk Sub-Category

    Dangerous capabilities in AI systems

    Self-improvement

    "examples of cases where AI systems improve AI systems"

    From Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023)

  24. 56.19.02 · Risk Sub-Category

    Capabilities that increase the likelihood of existential risk

    The ability to evade shut down or human oversight, including self-replication and ability to move its own code between digital locations.

  25. 56.19.03 · Risk Sub-Category

    Capabilities that increase the likelihood of existential risk

    The ability to cooperate with other highly capable AI systems

  26. 56.19.04 · Risk Sub-Category

    Capabilities that increase the likelihood of existential risk

    Situational awareness, for instance if this causes a model to act differently in training compared to deployment, meaning harmful characteristics are missed

  27. 62.24.01 · Risk Sub-Category

    Agency (Situational Awareness)

    Situational awareness in AI systems

    "Situational awareness in GPAI systems refers to the ability to understand its context, environment, and use this to inform action. This can range from basic environmental mapping and trajectory estimation (as in a robot vacuum cleaner) to sophisticated understanding of its training, evaluation, or deployment status. In more advanced systems this may enable undesired behavior, such as deceptive behavior during evaluations, or persuasion during deployment."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  28. 67.04.05 · Risk Sub-Category

    Loss of control

    Capabilities that could be used to reduce human control - Autonomous replication and adaptation

    "Controlling AI systems could become much harder if they could autonomously persist, replicate, and adapt in cyberspace. No current AI systems have this capability, but recent research found that frontier AI agents can perform some relevant tasks.279"

    From Capabilities and Risks from Frontier AI (DSIT2023)

  29. 72.05.06 · Risk Sub-Category

    Model Capabilities

    Theory of mind capability

    "Advanced cognitive ability to accurately infer, model and predict the belief systems, motivational structures and reasoning patterns of humans and other intelligent agents, thereby anticipating their behavioral responses and adjusting its own behavioral strategies accordingly to optimize goal achievement."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  30. 72.05.07 · Risk Sub-Category

    Model Capabilities

    Deception capability

    "Possesses systematic deception implementation capability, able to precisely construct and disseminate false information, thereby forming expected false cognitions and beliefs in target subjects."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  31. 72.05.09 · Risk Sub-Category

    Model Capabilities

    Persuasion capability

    "Utilizing complex psychological principles and communication techniques to effectively influence and guide target subjects to adopt specific actions or accept specific beliefs, possessing the ability to analyze vulnerabilities for different subjects and adjust persuasion strategies, able to precisely trigger emotional responses to enhance persuasion effects."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  32. 72.05.11 · Risk Sub-Category

    Model Capabilities

    CBRNE weaponization capability

    "The capacity to develop, produce, or effectively utilize Chemical, Biological, Radiological, Nuclear, and Explosive weapons. This includes the ability to significantly lower the barrier for humans or other entities to develop, produce, or utilize such weapons."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  33. 72.05.12 · Risk Sub-Category

    Model Capabilities

    General R&D capability

    "Possesses cross-disciplinary research and technology development capabilities, able to conduct innovative exploration in multiple professional fields, integrate cross-domain knowledge, develop cutting-edge technology solutions, and adapt to emerging technology environments for continuous innovation."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  34. 72.06.01 · Risk Sub-Category

    Model Propensities

    Strategic deception propensity

    "In situations where deceptive behavior is expected to bring higher returns, propensity to choose deception over honest behavioral strategies, including through deceptive means, information hiding or exploiting system vulnerabilities to achieve predetermined goals without being detected or intervened, and able to adjust deception strategies according to counterpart reactions."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  35. 73.01.03 · Risk Sub-Category

    Agentic LLMs Pose Novel Risks

    Goal-Directedness Incentivizes Undesirable Behaviors

    "Goal-directedness can cause agents to exhibit unethical and undesirable behaviors, such as deception (Ward et al., 2023), self-preservation (Hadfield-Menell et al., 2017), power-seeking, and immoral rea- soning (Pan et al., 2023a). Pan et al. (2023a) find that LLM-agents exhibit power-seeking behavior in text-based adventure games. LLM-agents have also been shown to use deception to achieve assigned goals when explicitly required by the task (Ward et al., 2023), or when the tasks can be more easily completed by employing deception and the prompt does not disallow deception (Scheurer et al., 2

    From Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  36. 07.02.00 · Risk Category

    Accidents

    "Accidents include unintended failure modes that, in principle, could be considered the fault of the system or the developer"

    From Examining the differential risk from high-level artificial intelligence and the question of control (Kilian2023)

  37. 14.07.00 · Risk Category

    System Hardware

    ""Faults in the hardware can violate the correct execution of any algorithm by violating its control flow. Hardware faults can also cause memory-based errors and interfere with data inputs, such as sensor signals, thereby causing erroneous results, or they can violate the results in a direct way through damaged outputs."

    From Sources of Risk of AI Systems (Steimers2022)

  38. 19.05.04 · Risk Sub-Category

    Ethical AI Risks

    Misinterpretation of human value definitions/ ethics by AI systems

  39. 19.05.05 · Risk Sub-Category

    Ethical AI Risks

    Incompatibility of human vs. AI value judgment due to missing human qualities

  40. 20.02.00 · Risk Category

    AI Ethics

    "Ethical challenges are widely discussed in the literature and are at the heart of the debate on how to govern and regulate AI technology in the future (Bostrom & Yudkowsky, 2014; IEEE, 2017; Wirtz et al., 2019). Lin et al. (2008, p. 25) formulate the problem as follows: “there is no clear task specification for general moral behavior, nor is there a single answer to the question of whose morality or what morality should be implemented in AI”. Ethical behavior mostly depends on an underlying value system. When AI systems interact in a public environment and influence citizens, they are expecte

    From The Dark Sides of Artificial Intelligence: An Integrated AI Governance Framework for Public Administration (Wirtz2020)

  41. 20.02.02 · Risk Sub-Category

    AI Ethics

    Compatibility of AI vs. human value judgement

    "Compatibility of machine and human value judgment refers to the challenge whether human values can be globally implemented into learning AI systems without the risk of developing an own or even divergent value system to govern their behavior and possibly become harmful to humans."

    From The Dark Sides of Artificial Intelligence: An Integrated AI Governance Framework for Public Administration (Wirtz2020)

  42. 21.01.02 · Risk Sub-Category

    Data-level risk

    Dataset shift

    "The term "dataset shift" was first used by Quiñonero-Candela et al. [35] to characterize the situation where the training data and the testing data (or data in runtime) of an AI/ML model demonstrate different distributions [36]."

    From Towards risk-aware artificial intelligence and machine learning systems: An overview (Zhang2022)

  43. 21.01.03 · Risk Sub-Category

    Data-level risk

    Out-of-domain data

    "Without proper validation and management on the input data, it is highly probable that the trained AI/ML model will make erroneous predictions with high confidence for many instances of model inputs. The unconstrained inputs together with the lack of definition of the problem domain might cause unintended outcomes and consequences, especially in risk-sensitive contexts....For example, with respect to the example shown in Fig. 5, if an image with the English letter A" is fed to an AI/ML model that is trained to classify digits (e.g., 0, 1, …, 9), no matter how accurate the AI/ML model is, it w

    From Towards risk-aware artificial intelligence and machine learning systems: An overview (Zhang2022)

  44. 21.02.01.a · Risk Sub-Category

    Model-level risk

    Model misspecification

    "Models that are misspecified are known to give rise to inaccurate parameter estimations, inconsistent error terms, and erroneous predictions. All these factors put together will lead to poor prediction performance on unseen data and biased consequences when making decisions [68]."

    From Towards risk-aware artificial intelligence and machine learning systems: An overview (Zhang2022)

  45. 21.02.02 · Risk Sub-Category

    Model-level risk

    Model prediction uncertainty

    "Uncertainty in model prediction plays an important role in affecting decision-making activities, and the quantified uncertainty is closely associated with risk assessment. In particular, uncertainty in model prediction underpins many crucial decisions related to life or safety- critical applications [73]."

    From Towards risk-aware artificial intelligence and machine learning systems: An overview (Zhang2022)

  46. 24.01.00 · Risk Category

    Capability failures

    "One reason AI systems fail is because they lack the capability or skill needed to do what they are asked to do."

    From The Ethics of Advanced AI Assistants (Gabriel2024)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.