MIT AI Risk Repository · Risk Sub-Category · 53.03.02
Gradual, irretrievable ceding of human power over the future to AI systems
Category: Direct catastrophe from AI
Description
-
From Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Causal entity
- Human
- Intent
- Unintentional
- Timing
- Post-deployment
Subdomain definition: Humans delegating key decisions to AI systems, or AI systems making decisions that diminish human control and autonomy, potentially leading to humans feeling disempowered, losing the ability to shape a fulfilling life trajectory or becoming cognitively enfeebled.
Real-world incidents in this subdomain
- DisMech AI Curation Agent Reportedly Completed GitHub Issue Intended as New Contributor's Learning Task
- Waymo Driverless Taxi Allegedly Stalled During Pedestrian Harassment Incident in San Francisco
- California Police Turned on Music to Allegedly Trigger Instagram’s DCMA to Avoid Being Live-Streamed
- Hawaii Police Deployed Robot Dog to Patrol a Homeless Encampment
How other frameworks describe this risk
Other entries from Maas2023
- Alignment failures in existing ML systems
- Faulty reward functions in the wild
- Specification gaming
- Reward model overoptimization
- Instrumental convergence
- Goal misgeneralization
- Inner misalignment
- Language model misalignment
- Harms from increasingly agentic algorithmic systems
- Dangerous capabilities in AI systems
- Situational awareness
- Acquisition of a goal to harm society