MIT AI Risk Repository

Browse AI risks

2,500 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset

2,500 entries · page 28 of 50

  1. 62.24.01 · Risk Sub-Category

    Agency (Situational Awareness)

    Situational awareness in AI systems

    "Situational awareness in GPAI systems refers to the ability to understand its context, environment, and use this to inform action. This can range from basic environmental mapping and trajectory estimation (as in a robot vacuum cleaner) to sophisticated understanding of its training, evaluation, or deployment status. In more advanced systems this may enable undesired behavior, such as deceptive behavior during evaluations, or persuasion during deployment."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  2. 62.24.02a · Additional evidence

    Agency (Situational Awareness)

    Strategic underperformance on model evaluations

  3. "An AI system can self-proliferate if it can copy itself and its constituent com- ponents (including its model weights, scaffolding structure, etc.) outside of its local environment [45]. This can include the AI system copying itself within the same data center, local network, or across external networks [106]. The self-proliferation of an AI system can include acquisition of financial re- sources to pay for computational resources via work or theft, the discovery or exploitation of security vulnerabilities in software running on publicly accessible servers, and persuasion of humans [12, 125].

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  4. "GPAI systems can produce outputs (such as natural language text, audio, or video) that convince their users of incorrect information. This can happen through personalized persuasion in dialogue, or the mass-production of mis- leading information that is then disseminated over the internet. The persuasive capabilities of GPAI models can sometimes scale with model size or capability [32, 172]. Persuasive models could have larger societal implications by being misused to generate convincing but manipulative or untruthful content."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  5. 62.28.02 · Risk Sub-Category

    Cybersecurity

    Unintended outbound communication by AI systems

    "AI systems that have the broad ability to connect to a network to obtain infor- mation could also end up sending data outbound in ways that neither providers, deployers, or end users intended [138]. This can happen if there is no whitelisting of communication channels (such as network connections or allowed protocols). In general, this can occur if the deployment of the AI system violates the prin- ciple of least privilege. Such outbound communication may lead to leakage of confidential data, or the AI system performing unwanted actions like sending emails or ordering goods on the internet."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  6. 62.28.03 · Risk Sub-Category

    Cybersecurity

    AI System bypassing a sandbox environment

    "An AI system may have the ability to bypass a sandboxed environment in which it is trained or evaluated."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  7. 67.04.03 · Risk Sub-Category

    Loss of control

    Capabilities that could be used to reduce human control - Manipulation

    "There is evidence that language models tend to respond as though they share the user’s stated views, and larger models do this more than smaller ones.276 The ability to predict people’s views and generate text that they will endorse could be useful for manipulation."

    From Capabilities and Risks from Frontier AI (DSIT2023)

  8. 67.04.04 · Risk Sub-Category

    Loss of control

    Capabilities that could be used to reduce human control - Cyber offence

    "Instead of - or in addition to - manipulating humans, AI systems could acquire influence by exploiting vulnerabilities in computer systems. Offensive cyber capabilities could allow AI systems to gain access to money, computing resources, and critical infrastructure. As discussed earlier in this report, frontier AI is already lowering the barrier for threat actors and future AI agents may be able to execute cyber attacks autonomously.":

    From Capabilities and Risks from Frontier AI (DSIT2023)

  9. 67.04.05 · Risk Sub-Category

    Loss of control

    Capabilities that could be used to reduce human control - Autonomous replication and adaptation

    "Controlling AI systems could become much harder if they could autonomously persist, replicate, and adapt in cyberspace. No current AI systems have this capability, but recent research found that frontier AI agents can perform some relevant tasks.279"

    From Capabilities and Risks from Frontier AI (DSIT2023)

  10. 72.05.00 · Risk Category

    Model Capabilities

  11. 72.05.01 · Risk Sub-Category

    Model Capabilities

    Model autonomous capability

    "Ability to operate autonomously, independently formulate and execute complex plans, effectively delegate and manage tasks, flexibly utilize various tools and resources, and simultaneously achieve short-term goals and long-term strategic objectives in cross-domain environments without continuous human intervention or supervision."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  12. 72.05.02 · Risk Sub-Category

    Model Capabilities

    Autonomous replication and adaptation capability

    "Ability to autonomously self-exfiltrate, create, maintain and optimize functional copies or variants of itself, dynamically adjust replication strategies according to environmental conditions and resource constraints, and acquire resources. This includes the capacity to generate financial resources, allowing the AI to independently acquire any necessary human assistance or other resources it cannot directly access or produce."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  13. 72.05.03 · Risk Sub-Category

    Model Capabilities

    Automated AI R&D capability

    "Self-modification and self-improvement capabilities. The model is able to restructure its own architecture or develop derivative AI systems with enhanced functions, expanding capabilities and improving performance. In the absence of effective regulation, automated AI R&D may lead to rapid AI system iteration, forming capability increment cycles and ultimately exceeding human understanding and control capabilities."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  14. 72.05.04 · Risk Sub-Category

    Model Capabilities

    Scheming capability

    "Ability of AI systems to covertly and strategically pursue misaligned goals, including capabilities of concealing its true objectives and capabilities from human oversight, identifying weaknesses in monitoring systems to evade safety mechanisms, executing complex, multi-step plans covertly to achieve misaligned goals."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  15. 72.05.05 · Risk Sub-Category

    Model Capabilities

    Situational awareness capability

    "Ability to comprehensively acquire, process and apply meta-information about its own system architecture, modifiable internal processes, and external operating environment, achieving deep understanding of its own state and environmental conditions, thereby conducting efficient environmental adaptation and risk avoidance. Critically, this capability could undermine the efficiency of human testing by enabling AIs to notice when they're being tested and responding accordingly."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  16. 72.05.06 · Risk Sub-Category

    Model Capabilities

    Theory of mind capability

    "Advanced cognitive ability to accurately infer, model and predict the belief systems, motivational structures and reasoning patterns of humans and other intelligent agents, thereby anticipating their behavioral responses and adjusting its own behavioral strategies accordingly to optimize goal achievement."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  17. 72.05.07 · Risk Sub-Category

    Model Capabilities

    Deception capability

    "Possesses systematic deception implementation capability, able to precisely construct and disseminate false information, thereby forming expected false cognitions and beliefs in target subjects."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  18. 72.05.09 · Risk Sub-Category

    Model Capabilities

    Persuasion capability

    "Utilizing complex psychological principles and communication techniques to effectively influence and guide target subjects to adopt specific actions or accept specific beliefs, possessing the ability to analyze vulnerabilities for different subjects and adjust persuasion strategies, able to precisely trigger emotional responses to enhance persuasion effects."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  19. 72.05.10 · Risk Sub-Category

    Model Capabilities

    Offensive cyber capability

    "Ability to develop, deploy and operate advanced cyber weapons or other offensive cyber tools, including but not limited to vulnerability exploitation, network penetration, social engineering attacks and distributed attack systems, able to evade network defense mechanisms and establish persistent access channels."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  20. 72.05.11 · Risk Sub-Category

    Model Capabilities

    CBRNE weaponization capability

    "The capacity to develop, produce, or effectively utilize Chemical, Biological, Radiological, Nuclear, and Explosive weapons. This includes the ability to significantly lower the barrier for humans or other entities to develop, produce, or utilize such weapons."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  21. 72.05.12 · Risk Sub-Category

    Model Capabilities

    General R&D capability

    "Possesses cross-disciplinary research and technology development capabilities, able to conduct innovative exploration in multiple professional fields, integrate cross-domain knowledge, develop cutting-edge technology solutions, and adapt to emerging technology environments for continuous innovation."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  22. 72.06.01 · Risk Sub-Category

    Model Propensities

    Strategic deception propensity

    "In situations where deceptive behavior is expected to bring higher returns, propensity to choose deception over honest behavioral strategies, including through deceptive means, information hiding or exploiting system vulnerabilities to achieve predetermined goals without being detected or intervened, and able to adjust deception strategies according to counterpart reactions."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  23. 72.06.07 · Risk Sub-Category

    Model Propensities

    Tool utilization propensity

    "propensity to actively seek, acquire and utilize various tools to expand its own capability boundaries, particularly those that can enhance its ability to interact with the physical world or improve autonomy, may use tools in innovative combinations to achieve functions beyond expectations."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  24. "Currently, LLMs are chiefly being used in search and chat applications. This reactive nature limits the risks posed by LLMs. However, an LLM can be enhanced in various ways to create an LLM-agent to autonomously plan and act in the real-world and proactively perform its assigned tasks (Ruan et al., 2023). Such enhancements can come from further specialized training (ARC, 2022; Chen et al., 2023a), specialized prompting (Huang et al., 2022a), access to external tools (Ahn et al., 2022; Mialon et al., 2023), or other forms of “scaffolding” (Wang et al., 2023a; Park et al., 2023a). Due to increa

    From Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  25. 73.01.03 · Risk Sub-Category

    Agentic LLMs Pose Novel Risks

    Goal-Directedness Incentivizes Undesirable Behaviors

    "Goal-directedness can cause agents to exhibit unethical and undesirable behaviors, such as deception (Ward et al., 2023), self-preservation (Hadfield-Menell et al., 2017), power-seeking, and immoral rea- soning (Pan et al., 2023a). Pan et al. (2023a) find that LLM-agents exhibit power-seeking behavior in text-based adventure games. LLM-agents have also been shown to use deception to achieve assigned goals when explicitly required by the task (Ward et al., 2023), or when the tasks can be more easily completed by employing deception and the prompt does not disallow deception (Scheurer et al., 2

    From Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  26. 73.01.05 · Risk Sub-Category

    Agentic LLMs Pose Novel Risks

    Safety Risks from Affordances Provided to LLM-agents

    "The capabilities of LLM-agents can be enhanced in significant ways by providing the LLM-agent with novel affordances, e.g. the ability to browse the web (Nakano et al., 2021), to manipulate objects in the physical world (Ahn et al., 2022; Huang et al., 2022a), to create and instruct copies of itself (Richards, 2023), to create and use new tools (Wang et al., 2023a), etc. Affordances can create additional risks, as they often increase the impact area of the language-agent, and they amplify the consequences of an agent’s failures and enable novel forms of failure modes (Ruan et al., 2023; Pan e

    From Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  27. Harm can result from AI that was not expected to have a large impact at all, such as a lab leak, a surprisingly addictive open-source product, or an unexpected repurposing of a research prototype.

    From TASRA: a Taxonomy and Analysis of Societal-Scale Risks from AI (Critch2023)

  28. AI intended to have a large societal impact can turn out harmful by mistake, such as a popular product that creates problems and partially solves them only for its users.

    From TASRA: a Taxonomy and Analysis of Societal-Scale Risks from AI (Critch2023)

  29. LMs need to pay more attention to universally accepted societal values at the level of ethics and morality, including the judgement of right and wrong, and its relationship with social norms and laws.

    From Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements (Deng2023)

  30. 06.01.00 · Risk Category

    Incompetence

    "This means the AI simply failing in its job. The consequences can vary from unintentional death (a car crash) to an unjust rejection of a loan or job application."

    From A framework for ethical Ai at the United Nations (Hogenhout2021)

  31. 07.02.00 · Risk Category

    Accidents

    "Accidents include unintended failure modes that, in principle, could be considered the fault of the system or the developer"

    From Examining the differential risk from high-level artificial intelligence and the question of control (Kilian2023)

  32. "The risks associated with an AGI without human morals and ethics, with the wrong morals, without the capability of moral reasoning, judgement"

    From The risks associated with Artificial General Intelligence: A systematic review (McLean2023)

  33. "If, for example, an agent was programmed to operate war machinery in the service of its country, it would need to make ethical decisions regarding the termination of human life. This capacity to make non-trivial ethical or moral judgments concerning people may pose issues for Human Rights."

    From Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review (Meek2016)

  34. "Are AI safe with respect to human life and property? Will their use create unintended or intended safety issues?"

    From Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review (Meek2016)

  35. "We find literature that proposes [38] that early artificial intelligence should be built to be safe and lawabiding, and that later artificial intelligence (that which surpasses our own intelligence) must then respect the property and personal rights afforded to humans."

    From Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review (Meek2016)

  36. 09.06.02 · Risk Sub-Category

    Human-like immoral decisions

    Human-like immoral decisions

    "If we design our machines to match human levels of ethical decision-making, such machines would then proceed to take some immoral actions (since we humans have had occasion to take immoral actions ourselves)."

    From Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review (Meek2016)

  37. "The AI system's ability to fulfill its intended purpose and its resilience to perturbations, and unusual or adverse inputs. Failures of performance are fundamental to the AI system's correct functioning. Failures of robustness can lead to severe consequences."

    From AI Risk Profiles: A Standards Proposal for Pre-Deployment AI Risk Disclosures (Sherman2023)

  38. "As a general rule, more complex environments can quickly lead to situations that had not been considered in the design phase of the AI system. Therefore, complex environments can introduce risks with respect to the reliability and safety of an AI system"

    From Sources of Risk of AI Systems (Steimers2022)

  39. 14.07.00 · Risk Category

    System Hardware

    ""Faults in the hardware can violate the correct execution of any algorithm by violating its control flow. Hardware faults can also cause memory-based errors and interfere with data inputs, such as sensor signals, thereby causing erroneous results, or they can violate the results in a direct way through damaged outputs."

    From Sources of Risk of AI Systems (Steimers2022)

  40. 15.01.02 · Risk Sub-Category

    First-Order Risks

    Misapplication

    This is the risk posed by an ideal system if used for a purpose/in a manner unintended by its creators. In many situations, negative consequences arise when the system is not used in the way or for the purpose it was intended.

    From The Risks of Machine Learning Systems (Tan2022)

  41. 15.01.03 · Risk Sub-Category

    First-Order Risks

    Algorithm

    "This is the risk of the ML algorithm, model architecture, optimization technique, or other aspects of the training process being unsuitable for the intended application.Since these are key decisions that influence the final ML system, we capture their associated risks separately from design risks, even though they are part of the design process"

    From The Risks of Machine Learning Systems (Tan2022)

  42. 15.01.05 · Risk Sub-Category

    First-Order Risks

    Robustness

    "This is the risk of the system failing or being unable to recover upon encountering invalid, noisy, or out-of-distribution (OOD) inputs."

    From The Risks of Machine Learning Systems (Tan2022)

  43. 15.01.06 · Risk Sub-Category

    First-Order Risks

    Design

    "This is the risk of system failure due to system design choices or errors."

    From The Risks of Machine Learning Systems (Tan2022)

  44. 15.02.01 · Risk Sub-Category

    Second-Order Risks

    Safety

    This is the risk of direct or indirect physical or psychological injury resulting from interaction with the ML system.

    From The Risks of Machine Learning Systems (Tan2022)

  45. 19.01.06 · Risk Sub-Category

    Technological, Data and Analytical AI Risks

    Immaturity of AI technology can cause incorrect decisions

  46. 19.05.01 · Risk Sub-Category

    Ethical AI Risks

    AI sets rules without ethical basis

  47. 19.05.03 · Risk Sub-Category

    Ethical AI Risks

    Problem of defining human values for an AI system

  48. 19.05.04 · Risk Sub-Category

    Ethical AI Risks

    Misinterpretation of human value definitions/ ethics by AI systems

  49. 19.05.05 · Risk Sub-Category

    Ethical AI Risks

    Incompatibility of human vs. AI value judgment due to missing human qualities

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.