MIT AI Risk Repository · domain 7: AI system safety, failures, & limitations
7.6 Multi-agent risks
Risks from multi-agent interactions, due to incentives (which can lead to conflict or collusion) and/or the structure of multi-agent systems, which can create cascading failures, selection pressures, new security vulnerabilities, and a lack of shared information and trust.
- 53
- 5
- —
- —
| Label | Value |
|---|---|
| AI | 35 |
| Other | 15 |
| Human | 3 |
| Label | Value |
|---|---|
| Unintentional | 23 |
| Other | 15 |
| Intentional | 15 |
| Label | Value |
|---|---|
| Post-deployment | 44 |
| Other | 8 |
| Pre-deployment | 1 |
| Label | Value |
|---|---|
| Risk Category | 11 |
| Risk Sub-Category | 42 |
Risk entries
Browse and export all- Undesirable Dispositions from Competition
"Undesirable Dispositions from Competition. It is plausible that evolution selected for certain conflict-prone dispostions in humans, such as vengefulness, aggression, risk-seeking, selfishness, disho...
- Undesirable Dispositions from Human Data
"Undesirable Dispositions from Human Data. It is well-understood that models trained on human data – such as being pre-trained on human-written text or fine-tuned on human feedback – can exhibit human...
- Undesirable Capabilities
"Undesirable Capabilities. As agents interact, they iteratively exploit each other’s weaknesses, forc- ing them to address these weaknesses and gain new capabilities. This co-adaptation between agents...
- Destabilising Dynamics
"Destabilising dynamics (Section 3.4): systems that adapt in response to one another can produce dangerous feedback loops and unpredictability;"
- Feedback Loops
"Feedback Loops. One of the best-known historical examples to illustrate destabilising dynamics in the context of autonomous agents is the 2010 flash crash, in which algorithmic trading agents entered...
- Cyclic Behaviour
"Cyclic Behaviour. The dynamics described above are highly non-linear (small changes to the system’s state can result in large changes to its trajectory). Similar non-linear dynamics can emerge in mul...
- Chaos
"Chaos. Unlike the systems that tend towards fixed points or cycles described above, chaotic systems are inherently unpredictable and highly sensitive to initial conditions. While it might seem easy t...
- Phase Transitions
"Phase Transitions. Finally, small external changes to the system – such as the introduction of new agents or a distributional shift – can cause phase transitions, where the system undergoes an abrupt...
- Distributional Shift
"Distributional Shift. Individual ML systems can perform poorly in contexts different from those in which they were trained. A key source of these distributional shifts is the actions and adaptations...
- Commitment and Trust
"Commitment and trust (Section 3.5): difficulties in forming credible commitments, trust, or reputation can prevent mutual gains in AI-AI and human-AI interactions;"
- Inefficient Outcomes
"Inefficient Outcomes. Without careful planning and the appropriate safeguards, we may soon be entering a world overrun by increasingly competent and autonomous software agents, able to act with littl...
- Threats and Extortion
"Threats and Extortion. A natural solution to problems of trust is to provide some kind of com- mitment ability to AI agents, which can be used to bind them to more cooperative courses of action. Unfo...
- Rigidity and Mistaken Commitments
"Rigidity and Mistaken Commitments. Even when it is desirable to be able to make threats in order to deter socially harmful behaviour, doing so using AI agents effectively removes the human from the l...
- Emergent Agency
"Emergent agency (Section 3.6): qualitatively different goals or capabilities can emerge from the composition of innocuous independent systems or behaviours;"
- Emergent Capabilities
"Emergent Capabilities. Dangerous emergent capabilities could arise when a multi-agent system over- comes the safety-enhancing limitations of the individual systems, such as individual models’ narrow...
- Emergent Goals
"Emergent Goals. Ascribing goals to a system is not always straightforward. For our present purposes, it will suffice to adopt a Dennetian perspective (Dennett, 1971), ascribing goals and intentions o...
- Multi-Agent Security
"Multi-agent security (Section 3.7): multi-agent systems give rise to new kinds of security threats and vulnerabilities."
- Swarm Attacks
"Swarm Attacks. The need for multi-agent security is foreshadowed by attacks today that benefit from the use of many decentralised agents, such as distributed denial-of-service attacks (Cisco, 2023; Y...
- Heterogeneous Attacks
"Heterogeneous Attacks. A closely related risk is the possibility of multiple agents combining different affordances to overcome safeguards, for which there is already preliminary evidence (Jones et a...
- Social Engineering at Scale
"Social Engineering at Scale. Advanced AI agents will be more easily able to interact with large numbers of humans, and vice versa. This provides a wider attack surface for various forms of automated...
- Vulnerable AI Agents
"Vulnerable AI Agents. The use of AI agents as delegates or representatives of humans or organisa- tions also introduces the possibility of attacks on AI agents themselves. In other words, agents can...
- Cascading Security Failures
"Cascading Security Failures. Localised attacks in multi-agent systems can result in catastrophic macroscopic outcomes (Motter & Lai, 2002, see also Sections 3.2 and 3.4). These cascades can be hard t...
- Undetectable Threats
"Undetectable Threats. Cooperation and trust in many multi-agent systems relies crucially on the ability to detect (and then avoid or sanction) adversarial actions taken by others (Ostrom, 1990; Schne...
- Impact on Financial Stability
"The integration of general-purpose AI into high-frequency trading, market-making, or systemic risk management could exacerbate systemic risk by exhibiting unexpected behavioral patterns during market...
- Multi-agent collaboration capability
"Multiple autonomous AI agents able to establish collaborative relationships through explicit communication or implicit behavioral consistency, forming decentralized decision networks, jointly executi...