MIT AI Risk Repository · Risk Sub-Category · 24.11.01
Entrenched viewpoints and reduced political efficacy
Category: Misinformation risks
Description
"Design choices such as greater personalisation of AI assistants and efforts to align them with human preferences could also reinforce people’s pre-existing biases and entrench specific ideologies. Increasingly agentic AI assistants trained using techniques such as reinforcement learning from human feedback (RLHF) and with the ability to access and analyse users’ behavioural data, for example, may learn to tailor their responses to users’ preferences and feedback. In doing so, these systems could end up producing partial or ideologically biased statements in an attempt to conform to user expec
From The Ethics of Advanced AI Assistants (Gabriel2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Domain
- 3. Misinformation
- Causal entity
- AI
- Intent
- Unintentional
- Timing
- Post-deployment
Subdomain definition: Highly personalized AI-generated misinformation creating “filter bubbles” where individuals only see what matches their existing beliefs, undermining shared reality, weakening social cohesion and political processes.
Real-world incidents in this subdomain
- Google Books Appears to Be Indexing Works Written by AI
- Uptick in Low-Quality AI-Produced Content Degraded Publishers' Submission Management
- Korean Politician Employed Deepfake as Campaign Representative
- Facebook Political Ad Delivery Algorithms Inferred Users' Political Alignment, Inhibiting Political Campaigns' Reach
How other frameworks describe this risk
- Radicalisation
- Information degradation
- Institutional trust loss
- Worsened epistemic processes for society
- AI contributes to increased online polarisation
- Widespread use of persuasive tools contributes to splintered epistemic communities
- Reduced decision-making capacity as a result of decreased trust in information
- Degradation of the information environment
Other entries from Gabriel2024
- Capability failures
- Lack of capability for task
- Difficult to develop metrics for evaluating benefits or harms caused by AI assistants
- Safe exploration problem with widely deployed AI assistants
- Goal-related failures
- Misaligned consequentialist reasoning
- Specification gaming
- Goal misgeneralisation
- Deceptive alignment
- Malicious Uses
- Offensive Cyber Operations (General)
- AI-Powered Spear-Phishing at Scale