MIT AI Risk Repository · Risk Sub-Category · 62.23.03

Deceptive behavior because of an incorrect world model

Category: Agency (Deception)

Description

"AI systems can create deceptive outputs because their learned world model is not an accurate model of the real world [210]."

From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).

Classification

Causal entity
AI

Subdomain definition: AI systems that develop, access, or are provided with capabilities that increase their potential to cause mass harm through deception, weapons development and acquisition, persuasion and manipulation, political strategy, cyber-offense, AI development, situational awareness, and self-proliferation. These capabilities may cause mass harm due to malicious human actors, misaligned AI systems, or failure in the AI system.

How other frameworks describe this risk

Other entries from Gipiškis2024