MIT AI Risk Repository · Risk Sub-Category · 24.02.01

Misaligned consequentialist reasoning

Category: Goal-related failures

Description

"As we think about even more intelligent and advanced AI assistants, perhaps outperforming humans on many cognitive tasks, the question of how humans can successfully control such an assistant looms large. To achieve the goals we set for an assistant, it is possible (Shah, 2022) that the AI assistant will implement some form of consequentialist reasoning: considering many different plans, predicting their consequences and executing the plan that does best according to some metric, M. This kind of reasoning can arise because it is a broadly useful capability (e.g. planning ahead, considering mo

From The Ethics of Advanced AI Assistants (Gabriel2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).

Classification

Causal entity
AI
Intent
Other

Subdomain definition: AI systems that fail to perform reliably or effectively under varying conditions, exposing them to errors and failures that can have significant consequences, especially in critical applications or areas that require moral reasoning.

Real-world incidents in this subdomain

Browse all incidents in this subdomain

How other frameworks describe this risk

Other entries from Gabriel2024