MIT AI Risk Repository · domain 7: AI system safety, failures, & limitations
7.2 AI possessing dangerous capabilities
AI systems that develop, access, or are provided with capabilities that increase their potential to cause mass harm through deception, weapons development and acquisition, persuasion and manipulation, political strategy, cyber-offense, AI development, situational awareness, and self-proliferation. These capabilities may cause mass harm due to malicious human actors, misaligned AI systems, or failure in the AI system.
- 78
- 12
- —
- —
| Label | Value |
|---|---|
| AI | 67 |
| Human | 5 |
| Not coded | 4 |
| Other | 2 |
| Label | Value |
|---|---|
| Intentional | 52 |
| Other | 12 |
| Unintentional | 10 |
| Not coded | 4 |
| Label | Value |
|---|---|
| Other | 33 |
| Post-deployment | 33 |
| Pre-deployment | 8 |
| Not coded | 4 |
| Label | Value |
|---|---|
| Risk Category | 24 |
| Risk Sub-Category | 53 |
| Additional evidence | 1 |
Risk entries
Browse and export all- Agentic LLMs Pose Novel Risks
"Currently, LLMs are chiefly being used in search and chat applications. This reactive nature limits the risks posed by LLMs. However, an LLM can be enhanced in various ways to create an LLM-agent to...
- Goal-Directedness Incentivizes Undesirable Behaviors
"Goal-directedness can cause agents to exhibit unethical and undesirable behaviors, such as deception (Ward et al., 2023), self-preservation (Hadfield-Menell et al., 2017), power-seeking, and immoral...
- Safety Risks from Affordances Provided to LLM-agents
"The capabilities of LLM-agents can be enhanced in significant ways by providing the LLM-agent with novel affordances, e.g. the ability to browse the web (Nakano et al., 2021), to manipulate objects i...
- Capabilities that could be used to reduce human control - Manipulation
"There is evidence that language models tend to respond as though they share the user’s stated views, and larger models do this more than smaller ones.276 The ability to predict people’s views and gen...
- Capabilities that could be used to reduce human control - Cyber offence
"Instead of - or in addition to - manipulating humans, AI systems could acquire influence by exploiting vulnerabilities in computer systems. Offensive cyber capabilities could allow AI systems to gain...
- Capabilities that could be used to reduce human control - Autonomous replication and adaptation
"Controlling AI systems could become much harder if they could autonomously persist, replicate, and adapt in cyberspace. No current AI systems have this capability, but recent research found that fron...
- Subagents
"An AGI may decide to create subagents to help it with its task (Orseau, 2014a,b; Soares, Fallenstein, et al., 2015). These agents may for example be copies of the original agent’s source code running...
- Nascent capabilities (agency and autonomy)
"Traditionally, AI tools have been viewed as passive instruments controlled by users to achieve their goals, lacking the ability to take action or assume responsibilities. However, advanced AI tools a...
- Nascent capabilities (emergent capabilities)
"As large models undergo scaling, they meet critical thresholds at which they spontaneously develop new capabilities. The term “emergent behavior” refers to the unexpected or surprising outputs such m...
- Artificial general intelligence (existential risk posed by Artificial General Intelligence)
"In a paper called “How Does Artificial Intelligence Pose an Existential Risk?” published in 2017, Karina Vold and Daniel Harris suggested that humans might create a super-intelligent machine that cou...
- Capabilities that increase the likelihood of existential risk
-
- Agency and autonomy
-
- The ability to evade shut down or human oversight, including self-replication and ability to move its own code between digital locations.
-
- The ability to cooperate with other highly capable AI systems
-
- Situational awareness, for instance if this causes a model to act differently in training compared to deployment, meaning harmful characteristics are missed
-
- Self-improvement
-
- AI Influence
"ways in which advanced AI assistants could influence user beliefs and behaviour in ways that depart from rational persuasion"
- Fine-tuning related (Unexpected competence in fine-tuned versions of the upstream model)
"Downstream deployers may often fine-tune a GPAI model with specific deploy- ment-related datasets, to better suit the task. Fine-tuned upstream models can gain new or unexpected capabilities that the...
- Encoded reasoning
"Models can employ steganography techniques to encode their intermediate rea- soning steps in ways that are not interpretable by humans [166]. Since en- coded reasoning can improve model performance,...
- Agency
"This section catalogs the risk sources and risk management measures related to agentic AI systems. We categorize these into the following groups: goal- directedness, deception, situational awareness,...
- Agency (Goal-Directedness)
- Deceptive behavior for game-theoretical reasons
"An AI system can display deceptive behavior, such as cheating or bluffing, when engaging in such behavior is a good or optimal game-theoretical strategy to achieve the goals it has been configured to...
- Deceptive behavior because of an incorrect world model
"AI systems can create deceptive outputs because their learned world model is not an accurate model of the real world [210]."
- Deceptive behavior leading to unauthorized actions
"AI systems can create false or misleading claims that can lead to unauthorized actions, even in some cases violating the terms and conditions set by the model provider [79, 1]. For example, an AI sys...
- Agency (Situational Awareness)
-