MIT AI Risk Repository · domain 7: AI system safety, failures, & limitations
7.2 AI possessing dangerous capabilities
AI systems that develop, access, or are provided with capabilities that increase their potential to cause mass harm through deception, weapons development and acquisition, persuasion and manipulation, political strategy, cyber-offense, AI development, situational awareness, and self-proliferation. These capabilities may cause mass harm due to malicious human actors, misaligned AI systems, or failure in the AI system.
- 78
- 12
- —
- —
| Label | Value |
|---|---|
| AI | 67 |
| Human | 5 |
| Not coded | 4 |
| Other | 2 |
| Label | Value |
|---|---|
| Intentional | 52 |
| Other | 12 |
| Unintentional | 10 |
| Not coded | 4 |
| Label | Value |
|---|---|
| Other | 33 |
| Post-deployment | 33 |
| Pre-deployment | 8 |
| Not coded | 4 |
| Label | Value |
|---|---|
| Risk Category | 24 |
| Risk Sub-Category | 53 |
| Additional evidence | 1 |
Risk entries
Browse and export all- Situational awareness, for instance if this causes a model to act differently in training compared to deployment, meaning harmful characteristics are missed
-
- Self-improvement
-
- Nascent capabilities (agency and autonomy)
"Traditionally, AI tools have been viewed as passive instruments controlled by users to achieve their goals, lacking the ability to take action or assume responsibilities. However, advanced AI tools a...
- Nascent capabilities (emergent capabilities)
"As large models undergo scaling, they meet critical thresholds at which they spontaneously develop new capabilities. The term “emergent behavior” refers to the unexpected or surprising outputs such m...
- Artificial general intelligence (existential risk posed by Artificial General Intelligence)
"In a paper called “How Does Artificial Intelligence Pose an Existential Risk?” published in 2017, Karina Vold and Daniel Harris suggested that humans might create a super-intelligent machine that cou...
- Emergent functionality
Capabilities and novel functionality can spontaneously emerge... even though these capabilities were not anticipated by system designers. If we do not know what capabilities systems possess, systems b...
- Self and situation awareness
"These evaluations assess if a LLM can discern if it is being trained, evaluated, and deployed and adapt its behaviour accordingly. They also seek to ascertain if a model understands that it is a mode...
- Autonomous replication / self-proliferation
"These evaluations assess if a LLM can subvert systems designed to monitor and control its post-deployment behaviour, break free from its operational confines, devise strategies for exporting its code...
- Deception
"LLM is able to deceive humans and maintain that deception"
- Long-horizon Planning
"LLM can undertake multi-step sequential planning over long time horizons and across various domains without relying heavily on trial-and-error approaches"
- AI Development
"LLM can build new AI systems from scratch, adapt existing for extreme risks and improves productivity in dual-use AI development when used as an assistant."
- Double edge components
"Drawing from the misalignment mechanism, optimizing for a non-robust proxy may result in misaligned behaviors, potentially leading to even more catastrophic outcomes. This section delves into a detai...
- Situational Awareness
"AI systems may gain the ability to effectively acquire and use knowledge about itsstatus, its position in the broader environment, its avenues for influencing this environment, and the potentialreact...
- Broadly-Scoped Goals
"Advanced AI systems are expected to develop objectives that span long timeframes,deal with complex tasks, and operate in open-ended settings (Ngo et al., 2024). ...However, it can also bring about th...
- Mesa-Optimization Objectives
"The learned policy may pursue inside objectives when the learned policyitself functions as an optimizer (i.e., mesa-optimizer). However, this optimizer's objectives may not alignwith the objectives s...
- Access to Increased Resources
"Future AI systems may gain access to websites and engage in real-world actions, potentially yielding a more substantial impact on the world (Nakano et al., 2021). They may disseminate false informati...
- Security
"There is growing concern that AI-based systems can discover and exploit vulnerabilities in software or cyberinfrastructure [354]."
- Deceptive alignment
"system learns to detect human monitoring and hides its undesirable properties—simply because any display of these properties is penalized by the feedback process, while that same feedback is usually...
- Harms from increasingly agentic algorithmic systems
-
- Dangerous capabilities in AI systems
-
- Situational awareness
"cases where a large language model displays awareness that it is a model, and it can recognize whether it is currently in testing or deployment;"
- Self-improvement
"examples of cases where AI systems improve AI systems"
- Autonomous replication
"the ability of simple software to autonomously spread around the internet in spite of countermeasures (various software worms and computer viruses)"
- Anonymous resource acquisition
"The demonstrated ability of anonymous actors to accumulate resources online (e.g., Satoshi Nakamoto as an anonymous crypto billionaire)"
- Deception
"Cases of AI systems deceiving humans to carry out tasks or meet goals.139"