MIT AI Risk Repository · Risk Category · 73.01.00
Agentic LLMs Pose Novel Risks
Description
"Currently, LLMs are chiefly being used in search and chat applications. This reactive nature limits the risks posed by LLMs. However, an LLM can be enhanced in various ways to create an LLM-agent to autonomously plan and act in the real-world and proactively perform its assigned tasks (Ruan et al., 2023). Such enhancements can come from further specialized training (ARC, 2022; Chen et al., 2023a), specialized prompting (Huang et al., 2022a), access to external tools (Ahn et al., 2022; Mialon et al., 2023), or other forms of “scaffolding” (Wang et al., 2023a; Park et al., 2023a). Due to increa
From Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Causal entity
- AI
- Intent
- Other
- Timing
- Post-deployment
Subdomain definition: AI systems that develop, access, or are provided with capabilities that increase their potential to cause mass harm through deception, weapons development and acquisition, persuasion and manipulation, political strategy, cyber-offense, AI development, situational awareness, and self-proliferation. These capabilities may cause mass harm due to malicious human actors, misaligned AI systems, or failure in the AI system.
How other frameworks describe this risk
- Capabilities that could be used to reduce human control - Cyber offence
- Capabilities that could be used to reduce human control - Autonomous replication and adaptation
- Capabilities that could be used to reduce human control - Manipulation
- Subagents
- AI Influence
- Agency
- AI System bypassing a sandbox environment
- Fine-tuning related (Unexpected competence in fine-tuned versions of the upstream model)
Other entries from Anwar2024
- Natural Language Underspecifies Goals
- Goal-Directedness Incentivizes Undesirable Behaviors
- Safety Risks from Affordances Provided to LLM-agents
- Multi-Agent Safety Is Not Assured by Single-Agent Safety
- Foundationality May Cause Correlated Failures
- Groups of LLM-Agents May Show Emergent Functionality
- Collusion between LLM-Agents
- Dual-Use Capabilities Enable Malicious Use and Misuse of LLMs
- Misinformation and Manipulation
- Cybersecurity
- Cybersecurity
- Cybersecurity