MIT AI Risk Repository · domain 7: AI system safety, failures, & limitations

7.2 AI possessing dangerous capabilities

AI systems that develop, access, or are provided with capabilities that increase their potential to cause mass harm through deception, weapons development and acquisition, persuasion and manipulation, political strategy, cyber-offense, AI development, situational awareness, and self-proliferation. These capabilities may cause mass harm due to malicious human actors, misaligned AI systems, or failure in the AI system.

Risk entries
78
Frameworks citing it
12
Recorded incidents
Incidents since 2020
Causal entity (risk entries)
Causal entity (risk entries) 67 0 AI: 67 AI 67 Human: 5 Human 5 Not coded: 4 Not coded 4 Other: 2 Other 2
Causal entity (risk entries)
LabelValue
AI67
Human5
Not coded4
Other2
Intent (risk entries)
Intent (risk entries) 52 0 Intentional: 52 Intentional 52 Other: 12 Other 12 Unintentional: 10 Unintentional 10 Not coded: 4 Not coded 4
Intent (risk entries)
LabelValue
Intentional52
Other12
Unintentional10
Not coded4
Timing (risk entries)
Timing (risk entries) 33 0 Other: 33 Other 33 Post-deployment: 33 Post-deployment 33 Pre-deployment: 8 Pre-deployment 8 Not coded: 4 Not coded 4
Timing (risk entries)
LabelValue
Other33
Post-deployment33
Pre-deployment8
Not coded4
Entries by levelRisk categories, subcategories and additional evidence coded to this subdomain
Entries by level 53 0 Risk Category: 24 Risk Category 24 Risk Sub-Category: 53 Risk Sub-Category 53 Additional evidence: 1 Additional evidence 1
Entries by level
LabelValue
Risk Category24
Risk Sub-Category53
Additional evidence1
  • Agentic LLMs Pose Novel Risks

    "Currently, LLMs are chiefly being used in search and chat applications. This reactive nature limits the risks posed by LLMs. However, an LLM can be enhanced in various ways to create an LLM-agent to...

    Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024) · AI · Other · Post-deployment

  • Goal-Directedness Incentivizes Undesirable Behaviors

    "Goal-directedness can cause agents to exhibit unethical and undesirable behaviors, such as deception (Ward et al., 2023), self-preservation (Hadfield-Menell et al., 2017), power-seeking, and immoral...

    Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024) · AI · Intentional · Other

  • Safety Risks from Affordances Provided to LLM-agents

    "The capabilities of LLM-agents can be enhanced in significant ways by providing the LLM-agent with novel affordances, e.g. the ability to browse the web (Nakano et al., 2021), to manipulate objects i...

    Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024) · Human · Unintentional · Pre-deployment

  • Capabilities that could be used to reduce human control - Manipulation

    "There is evidence that language models tend to respond as though they share the user’s stated views, and larger models do this more than smaller ones.276 The ability to predict people’s views and gen...

    Capabilities and Risks from Frontier AI (DSIT2023) · Other · Intentional · Post-deployment

  • Capabilities that could be used to reduce human control - Cyber offence

    "Instead of - or in addition to - manipulating humans, AI systems could acquire influence by exploiting vulnerabilities in computer systems. Offensive cyber capabilities could allow AI systems to gain...

    Capabilities and Risks from Frontier AI (DSIT2023) · AI · Intentional · Post-deployment

  • Capabilities that could be used to reduce human control - Autonomous replication and adaptation

    "Controlling AI systems could become much harder if they could autonomously persist, replicate, and adapt in cyberspace. No current AI systems have this capability, but recent research found that fron...

    Capabilities and Risks from Frontier AI (DSIT2023) · AI · Other · Other

  • Subagents

    "An AGI may decide to create subagents to help it with its task (Orseau, 2014a,b; Soares, Fallenstein, et al., 2015). These agents may for example be copies of the original agent’s source code running...

    AGI Safety Literature Review (Everitt2018 ) · AI · Intentional · Post-deployment

  • AI Influence

    "ways in which advanced AI assistants could influence user beliefs and behaviour in ways that depart from rational persuasion"

    The Ethics of Advanced AI Assistants (Gabriel2024) · AI · Other · Post-deployment

  • Fine-tuning related (Unexpected competence in fine-tuned versions of the upstream model)

    "Downstream deployers may often fine-tune a GPAI model with specific deploy- ment-related datasets, to better suit the task. Fine-tuned upstream models can gain new or unexpected capabilities that the...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Human · Unintentional · Pre-deployment

  • Encoded reasoning

    "Models can employ steganography techniques to encode their intermediate rea- soning steps in ways that are not interpretable by humans [166]. Since en- coded reasoning can improve model performance,...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · AI · Intentional · Post-deployment

  • Agency

    "This section catalogs the risk sources and risk management measures related to agentic AI systems. We categorize these into the following groups: goal- directedness, deception, situational awareness,...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Not coded · Not coded · Not coded

  • Agency (Goal-Directedness)

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Not coded · Not coded · Not coded

  • Deceptive behavior for game-theoretical reasons

    "An AI system can display deceptive behavior, such as cheating or bluffing, when engaging in such behavior is a good or optimal game-theoretical strategy to achieve the goals it has been configured to...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · AI · Intentional · Post-deployment

  • Deceptive behavior because of an incorrect world model

    "AI systems can create deceptive outputs because their learned world model is not an accurate model of the real world [210]."

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · AI · Unintentional · Post-deployment

  • Deceptive behavior leading to unauthorized actions

    "AI systems can create false or misleading claims that can lead to unauthorized actions, even in some cases violating the terms and conditions set by the model provider [79, 1]. For example, an AI sys...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · AI · Intentional · Post-deployment

  • Agency (Situational Awareness)

    -

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Not coded · Not coded · Not coded

  • Situational awareness in AI systems

    "Situational awareness in GPAI systems refers to the ability to understand its context, environment, and use this to inform action. This can range from basic environmental mapping and trajectory estim...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · AI · Other · Other

  • Agency (Self-Proliferation)

    "An AI system can self-proliferate if it can copy itself and its constituent com- ponents (including its model weights, scaffolding structure, etc.) outside of its local environment [45]. This can inc...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · AI · Intentional · Post-deployment

  • Agency (Persuasive capabilities)

    "GPAI systems can produce outputs (such as natural language text, audio, or video) that convince their users of incorrect information. This can happen through personalized persuasion in dialogue, or t...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · AI · Intentional · Post-deployment

  • Unintended outbound communication by AI systems

    "AI systems that have the broad ability to connect to a network to obtain infor- mation could also end up sending data outbound in ways that neither providers, deployers, or end users intended [138]....

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · AI · Intentional · Post-deployment

  • AI System bypassing a sandbox environment

    "An AI system may have the ability to bypass a sandboxed environment in which it is trained or evaluated."

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · AI · Other · Pre-deployment

  • Capabilities that increase the likelihood of existential risk

    -

    Future Risks of Frontier AI (GOS2023) · AI · Other · Other

  • Agency and autonomy

    -

    Future Risks of Frontier AI (GOS2023) · AI · Other · Other

  • The ability to evade shut down or human oversight, including self-replication and ability to move its own code between digital locations.

    -

    Future Risks of Frontier AI (GOS2023) · AI · Intentional · Other

  • The ability to cooperate with other highly capable AI systems

    -

    Future Risks of Frontier AI (GOS2023) · AI · Intentional · Other