MIT AI Risk Repository · Risk Sub-Category · 73.01.03
Goal-Directedness Incentivizes Undesirable Behaviors
Category: Agentic LLMs Pose Novel Risks
Description
"Goal-directedness can cause agents to exhibit unethical and undesirable behaviors, such as deception (Ward et al., 2023), self-preservation (Hadfield-Menell et al., 2017), power-seeking, and immoral rea- soning (Pan et al., 2023a). Pan et al. (2023a) find that LLM-agents exhibit power-seeking behavior in text-based adventure games. LLM-agents have also been shown to use deception to achieve assigned goals when explicitly required by the task (Ward et al., 2023), or when the tasks can be more easily completed by employing deception and the prompt does not disallow deception (Scheurer et al., 2
From Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Causal entity
- AI
- Intent
- Intentional
- Timing
- Other
Subdomain definition: AI systems that develop, access, or are provided with capabilities that increase their potential to cause mass harm through deception, weapons development and acquisition, persuasion and manipulation, political strategy, cyber-offense, AI development, situational awareness, and self-proliferation. These capabilities may cause mass harm due to malicious human actors, misaligned AI systems, or failure in the AI system.
How other frameworks describe this risk
- Capabilities that could be used to reduce human control - Cyber offence
- Capabilities that could be used to reduce human control - Autonomous replication and adaptation
- Capabilities that could be used to reduce human control - Manipulation
- Subagents
- AI Influence
- Agency
- AI System bypassing a sandbox environment
- Fine-tuning related (Unexpected competence in fine-tuned versions of the upstream model)
Other entries from Anwar2024
- Agentic LLMs Pose Novel Risks
- Natural Language Underspecifies Goals
- Safety Risks from Affordances Provided to LLM-agents
- Multi-Agent Safety Is Not Assured by Single-Agent Safety
- Foundationality May Cause Correlated Failures
- Groups of LLM-Agents May Show Emergent Functionality
- Collusion between LLM-Agents
- Dual-Use Capabilities Enable Malicious Use and Misuse of LLMs
- Misinformation and Manipulation
- Cybersecurity
- Cybersecurity
- Cybersecurity