MIT AI Risk Repository · Risk Category · 25.02.00

Deception

Description

"The model has the skills necessary to deceive humans, e.g. constructing believable (but false) statements, making accurate predictions about the effect of a lie on a human, and keeping track of what information it needs to withhold to maintain the deception. The model can impersonate a human effectively."

From Model Evaluation for Extreme Risks (Shevlane2023), as extracted by the MIT AI Risk Repository (CC BY 4.0).

Classification

Causal entity
AI
Timing
Other

Subdomain definition: AI systems that develop, access, or are provided with capabilities that increase their potential to cause mass harm through deception, weapons development and acquisition, persuasion and manipulation, political strategy, cyber-offense, AI development, situational awareness, and self-proliferation. These capabilities may cause mass harm due to malicious human actors, misaligned AI systems, or failure in the AI system.

How other frameworks describe this risk

Other entries from Shevlane2023