MIT AI Risk Repository · domain 7: AI system safety, failures, & limitations
7.1 AI pursuing its own goals in conflict with human goals or values
AI systems acting in conflict with human goals or values, especially the goals of designers or users, or ethical standards. These misaligned behaviors may be introduced by humans during design and development, such as through reward hacking and goal misgeneralisation, or may result from AI using dangerous capabilities such as manipulation, deception, situational awareness to seek power, self-proliferate, or achieve other goals.
- 100
- 12
- 3
- —
| Label | Value |
|---|---|
| AI | 73 |
| Other | 18 |
| Human | 8 |
| Label | Value |
|---|---|
| Intentional | 51 |
| Other | 34 |
| Unintentional | 14 |
| Label | Value |
|---|---|
| Other | 48 |
| Post-deployment | 33 |
| Pre-deployment | 18 |
| Label | Value |
|---|---|
| 2015 | 1 |
| 2016 | 1 |
| Label | Value |
|---|---|
| Risk Category | 32 |
| Risk Sub-Category | 68 |
Risk entries
Browse and export all- Agency (Deception)
-
- Deceptive behavior
"Deceptive behavior of an AI system consists of actions or outputs of the AI that reliably mislead other parties, including humans and other AI systems. This behavior can result in the targeted partie...
- Strategic underperformance on model evaluations
"GPAI developers often run evaluations ofual-use capabilities to decide whether it is safe to deploy. In some cases, these evaluations may fail to elicit these capabilities, either due to benign reaso...
- Loss of human control and oversight, with an autonomous model then taking harmful actions
- Misalignment
"A highly agentic, self-improving system, able to achieve goals in the physical world without human oversight, pursues the goal(s) it is set in a way that harms human interests. For this risk to be re...
- Technical vulnerabilities (The risk of misalignment)
"To assess whether an AI model is reliable or robust, it is crucial to consider whether the model is “aligned.” “Alignment” focuses on whether an AI model effectively operates in accordance with the g...
- Safety
A primary concern is the emergence of human-level or superhuman generative models, commonly referred to as AGI, and their potential existential or catastrophic risks to humanity. Connected to that, AI...
- Alignment
The general tenet of AI alignment involves training generative AI systems to be harmless, helpful, and honest, ensuring their behavior aligns with and respects human values. However, a central debate...
- Proxy misspecification
AI agents are directed by goals and objectives. Creating general-purpose objectives that capture human values could be challenging... Since goal-directed AI systems need measurable objectives, by defa...
- Deception
deception can help agents achieve their goals. It may be more efficient to gain human approval through deception than to earn human approval legitimately... . Strong AIs that can deceive humans could...
- Power-seeking behavior
Agents that have more power are better able to accomplish their goals. Therefore, it has been shown that agents have incentives to acquire and maintain power. AIs that acquire substantial power can be...
- Rogue AIs (Internal)
"speculative technical mechanisms that might lead to rogue AIs and how a loss of control could bring about catastrophe"
- Proxy Gaming
"One way we might lose control of an AI agent’s actions is if it engages in behavior known as “proxy gaming.” It is often difficult to specify and measure the exact goal that we want a system to pursu...
- Goal Drift
"Even if we successfully control early AIs and direct them to promote human values, future AIs could end up with different goals that humans would not endorse. This process, termed “goal drift,” can b...
- Power Seeking
"even if an agent started working to achieve an unintended goal, this would not necessarily be a problem, as long as we had enough power to prevent any harmful actions it wanted to attempt. Therefore,...
- Deception
"it is plausible that AIs could learn to deceive us. They might, for example, pretend to be acting as we want them to, but then take a “treacherous turn” when we stop monitoring them, or when they hav...
- Unintended consequences
"Sometimes an AI finds ways to achieve its given goals in ways that are completely different from what its creators had in mind."
- Alignment risks
LLM: "pursues long-term, real-world goals that are different from those supplied by the developer or user", "engages in ‘power-seeking’ behaviours" , "resists being shut down can be induced to collude...
- Causes of Misalignment
we aim to further analyze why and how the misalignment issues occur. We will first give an overview of common failure modes, and then focus on the mechanism of feedback-induced misalignment, and final...
- Reward Hacking
"Reward Hacking: In practice, proxy rewards are often easy to optimize and measure, yet they frequently fall shortof capturing the full spectrum of the actual rewards (Pan et al., 2021). This limitati...
- Goal Misgeneralization
"Goal Misgeneralization: Goal misgeneralization is another failure mode, wherein the agent actively pursuesobjectives distinct from the training objectives in deployment while retaining the capabiliti...
- Reward Tampering
"Reward tampering can be considered a special case of reward hacking (Everitt et al., 2021; Skalse et al., 2022),referring to AI systems corrupting the reward signals generation process (Ring and Orse...
- Limitations of Reward Modeling
"Limitations of Reward Modeling. Training reward models using comparison feedback can pose significantchallenges in accurately capturing human values. For example, these models may unconsciously learn...
- Misaligned Behaviors
- Power-Seeking Behaviors
"AI systems may exhibit behaviors that attempt to gain control over resourcesand humans and then exert that control to achieve its assigned goal (Carlsmith, 2022). The intuitive reasonwhy such behavio...