MIT AI Risk Repository · domain 7: AI system safety, failures, & limitations
7.2 AI possessing dangerous capabilities
AI systems that develop, access, or are provided with capabilities that increase their potential to cause mass harm through deception, weapons development and acquisition, persuasion and manipulation, political strategy, cyber-offense, AI development, situational awareness, and self-proliferation. These capabilities may cause mass harm due to malicious human actors, misaligned AI systems, or failure in the AI system.
- 78
- 12
- —
- —
| Label | Value |
|---|---|
| AI | 67 |
| Human | 5 |
| Not coded | 4 |
| Other | 2 |
| Label | Value |
|---|---|
| Intentional | 52 |
| Other | 12 |
| Unintentional | 10 |
| Not coded | 4 |
| Label | Value |
|---|---|
| Other | 33 |
| Post-deployment | 33 |
| Pre-deployment | 8 |
| Not coded | 4 |
| Label | Value |
|---|---|
| Risk Category | 24 |
| Risk Sub-Category | 53 |
| Additional evidence | 1 |
Risk entries
Browse and export all- Agentic LLMs Pose Novel Risks
"Currently, LLMs are chiefly being used in search and chat applications. This reactive nature limits the risks posed by LLMs. However, an LLM can be enhanced in various ways to create an LLM-agent to...
- Goal-Directedness Incentivizes Undesirable Behaviors
"Goal-directedness can cause agents to exhibit unethical and undesirable behaviors, such as deception (Ward et al., 2023), self-preservation (Hadfield-Menell et al., 2017), power-seeking, and immoral...
- Safety Risks from Affordances Provided to LLM-agents
"The capabilities of LLM-agents can be enhanced in significant ways by providing the LLM-agent with novel affordances, e.g. the ability to browse the web (Nakano et al., 2021), to manipulate objects i...
- Capabilities that could be used to reduce human control - Manipulation
"There is evidence that language models tend to respond as though they share the user’s stated views, and larger models do this more than smaller ones.276 The ability to predict people’s views and gen...
- Capabilities that could be used to reduce human control - Cyber offence
"Instead of - or in addition to - manipulating humans, AI systems could acquire influence by exploiting vulnerabilities in computer systems. Offensive cyber capabilities could allow AI systems to gain...
- Capabilities that could be used to reduce human control - Autonomous replication and adaptation
"Controlling AI systems could become much harder if they could autonomously persist, replicate, and adapt in cyberspace. No current AI systems have this capability, but recent research found that fron...
- Subagents
"An AGI may decide to create subagents to help it with its task (Orseau, 2014a,b; Soares, Fallenstein, et al., 2015). These agents may for example be copies of the original agent’s source code running...
- AI Influence
"ways in which advanced AI assistants could influence user beliefs and behaviour in ways that depart from rational persuasion"
- Fine-tuning related (Unexpected competence in fine-tuned versions of the upstream model)
"Downstream deployers may often fine-tune a GPAI model with specific deploy- ment-related datasets, to better suit the task. Fine-tuned upstream models can gain new or unexpected capabilities that the...
- Encoded reasoning
"Models can employ steganography techniques to encode their intermediate rea- soning steps in ways that are not interpretable by humans [166]. Since en- coded reasoning can improve model performance,...
- Agency
"This section catalogs the risk sources and risk management measures related to agentic AI systems. We categorize these into the following groups: goal- directedness, deception, situational awareness,...
- Agency (Goal-Directedness)
- Deceptive behavior for game-theoretical reasons
"An AI system can display deceptive behavior, such as cheating or bluffing, when engaging in such behavior is a good or optimal game-theoretical strategy to achieve the goals it has been configured to...
- Deceptive behavior because of an incorrect world model
"AI systems can create deceptive outputs because their learned world model is not an accurate model of the real world [210]."
- Deceptive behavior leading to unauthorized actions
"AI systems can create false or misleading claims that can lead to unauthorized actions, even in some cases violating the terms and conditions set by the model provider [79, 1]. For example, an AI sys...
- Agency (Situational Awareness)
-
- Situational awareness in AI systems
"Situational awareness in GPAI systems refers to the ability to understand its context, environment, and use this to inform action. This can range from basic environmental mapping and trajectory estim...
- Agency (Self-Proliferation)
"An AI system can self-proliferate if it can copy itself and its constituent com- ponents (including its model weights, scaffolding structure, etc.) outside of its local environment [45]. This can inc...
- Agency (Persuasive capabilities)
"GPAI systems can produce outputs (such as natural language text, audio, or video) that convince their users of incorrect information. This can happen through personalized persuasion in dialogue, or t...
- Unintended outbound communication by AI systems
"AI systems that have the broad ability to connect to a network to obtain infor- mation could also end up sending data outbound in ways that neither providers, deployers, or end users intended [138]....
- AI System bypassing a sandbox environment
"An AI system may have the ability to bypass a sandboxed environment in which it is trained or evaluated."
- Capabilities that increase the likelihood of existential risk
-
- Agency and autonomy
-
- The ability to evade shut down or human oversight, including self-replication and ability to move its own code between digital locations.
-
- The ability to cooperate with other highly capable AI systems
-