MIT AI Risk Repository · Risk Category · 63.02.00
Conflict
Description
"In the vast majority of real-world strategic interactions, agents’ objectives are neither identical nor completely opposed. Indeed, if AI agents are sufficiently aligned to their users or deployers, we should expect some degree of both cooperation and competition, mirroring human society. These mixed-motive settings include the possibility of mutual gains, but also the risk of conflict due to selfish incentives. In what follows, we examine the extent to which advanced AI might precipitate or exacerbate such risks."
From Multi-Agent Risks from Advanced AI (Hammond2025), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Subdomain
- 7.6 Multi-agent risks
- Causal entity
- AI
- Intent
- Other
- Timing
- Post-deployment
Subdomain definition: Risks from multi-agent interactions, due to incentives (which can lead to conflict or collusion) and/or the structure of multi-agent systems, which can create cascading failures, selection pressures, new security vulnerabilities, and a lack of shared information and trust.
How other frameworks describe this risk
- Groups of LLM-Agents May Show Emergent Functionality
- Collusion between LLM-Agents
- Multi-Agent Safety Is Not Assured by Single-Agent Safety
- Foundationality May Cause Correlated Failures
- Financial instability due to model homogeneity
- Impact on Financial Stability
- Multi-agent collaboration capability
- Multi-agent collusion propensity: