MIT AI Risk Repository · Risk Category · 39.05.00
Cheating and Deception
Description
may appear from intelligent agents such as HLI-based agents... Since HLI-based agents are going to mimic the behavior of humans, they may learn these behaviors accidentally from human-generated data. It should be noted that deception and cheating maybe appear in the behavior of every computer agent because the agent only focuses on optimizing some predefined objective functions, and the mentioned behavior may lead to optimizing the objective functions without any intention
From A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions (Saghiri2022), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Causal entity
- AI
- Intent
- Unintentional
- Timing
- Post-deployment
Subdomain definition: AI systems that develop, access, or are provided with capabilities that increase their potential to cause mass harm through deception, weapons development and acquisition, persuasion and manipulation, political strategy, cyber-offense, AI development, situational awareness, and self-proliferation. These capabilities may cause mass harm due to malicious human actors, misaligned AI systems, or failure in the AI system.
How other frameworks describe this risk
- Safety Risks from Affordances Provided to LLM-agents
- Agentic LLMs Pose Novel Risks
- Goal-Directedness Incentivizes Undesirable Behaviors
- Capabilities that could be used to reduce human control - Cyber offence
- Capabilities that could be used to reduce human control - Autonomous replication and adaptation
- Capabilities that could be used to reduce human control - Manipulation
- Subagents
- AI Influence