MIT AI Risk Repository · Risk Category · 62.21.00
Agency
Description
"This section catalogs the risk sources and risk management measures related to agentic AI systems. We categorize these into the following groups: goal- directedness, deception, situational awareness, self-proliferation, and persuasion"
From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
Subdomain definition: AI systems that develop, access, or are provided with capabilities that increase their potential to cause mass harm through deception, weapons development and acquisition, persuasion and manipulation, political strategy, cyber-offense, AI development, situational awareness, and self-proliferation. These capabilities may cause mass harm due to malicious human actors, misaligned AI systems, or failure in the AI system.
How other frameworks describe this risk
- Safety Risks from Affordances Provided to LLM-agents
- Agentic LLMs Pose Novel Risks
- Goal-Directedness Incentivizes Undesirable Behaviors
- Capabilities that could be used to reduce human control - Cyber offence
- Capabilities that could be used to reduce human control - Autonomous replication and adaptation
- Capabilities that could be used to reduce human control - Manipulation
- Subagents
- AI Influence