MIT AI Risk Repository · domain 7: AI system safety, failures, & limitations

7.6 Multi-agent risks

Risks from multi-agent interactions, due to incentives (which can lead to conflict or collusion) and/or the structure of multi-agent systems, which can create cascading failures, selection pressures, new security vulnerabilities, and a lack of shared information and trust.

Risk entries
53
Frameworks citing it
5
Recorded incidents
Incidents since 2020
Causal entity (risk entries)
Causal entity (risk entries) 35 0 AI: 35 AI 35 Other: 15 Other 15 Human: 3 Human 3
Causal entity (risk entries)
LabelValue
AI35
Other15
Human3
Intent (risk entries)
Intent (risk entries) 23 0 Unintentional: 23 Unintentional 23 Other: 15 Other 15 Intentional: 15 Intentional 15
Intent (risk entries)
LabelValue
Unintentional23
Other15
Intentional15
Timing (risk entries)
Timing (risk entries) 44 0 Post-deployment: 44 Post-deployment 44 Other: 8 Other 8 Pre-deployment: 1 Pre-deployment 1
Timing (risk entries)
LabelValue
Post-deployment44
Other8
Pre-deployment1
Entries by levelRisk categories, subcategories and additional evidence coded to this subdomain
Entries by level 42 0 Risk Category: 11 Risk Category 11 Risk Sub-Category: 42 Risk Sub-Category 42
Entries by level
LabelValue
Risk Category11
Risk Sub-Category42
  • Undesirable Dispositions from Competition

    "Undesirable Dispositions from Competition. It is plausible that evolution selected for certain conflict-prone dispostions in humans, such as vengefulness, aggression, risk-seeking, selfishness, disho...

    Multi-Agent Risks from Advanced AI (Hammond2025) · Other · Unintentional · Other

  • Undesirable Dispositions from Human Data

    "Undesirable Dispositions from Human Data. It is well-understood that models trained on human data – such as being pre-trained on human-written text or fine-tuned on human feedback – can exhibit human...

    Multi-Agent Risks from Advanced AI (Hammond2025) · Other · Unintentional · Post-deployment

  • Undesirable Capabilities

    "Undesirable Capabilities. As agents interact, they iteratively exploit each other’s weaknesses, forc- ing them to address these weaknesses and gain new capabilities. This co-adaptation between agents...

    Multi-Agent Risks from Advanced AI (Hammond2025) · AI · Intentional · Post-deployment

  • Destabilising Dynamics

    "Destabilising dynamics (Section 3.4): systems that adapt in response to one another can produce dangerous feedback loops and unpredictability;"

    Multi-Agent Risks from Advanced AI (Hammond2025) · AI · Unintentional · Post-deployment

  • Feedback Loops

    "Feedback Loops. One of the best-known historical examples to illustrate destabilising dynamics in the context of autonomous agents is the 2010 flash crash, in which algorithmic trading agents entered...

    Multi-Agent Risks from Advanced AI (Hammond2025) · AI · Unintentional · Post-deployment

  • Cyclic Behaviour

    "Cyclic Behaviour. The dynamics described above are highly non-linear (small changes to the system’s state can result in large changes to its trajectory). Similar non-linear dynamics can emerge in mul...

    Multi-Agent Risks from Advanced AI (Hammond2025) · AI · Unintentional · Post-deployment

  • Chaos

    "Chaos. Unlike the systems that tend towards fixed points or cycles described above, chaotic systems are inherently unpredictable and highly sensitive to initial conditions. While it might seem easy t...

    Multi-Agent Risks from Advanced AI (Hammond2025) · AI · Other · Other

  • Phase Transitions

    "Phase Transitions. Finally, small external changes to the system – such as the introduction of new agents or a distributional shift – can cause phase transitions, where the system undergoes an abrupt...

    Multi-Agent Risks from Advanced AI (Hammond2025) · Other · Unintentional · Post-deployment

  • Distributional Shift

    "Distributional Shift. Individual ML systems can perform poorly in contexts different from those in which they were trained. A key source of these distributional shifts is the actions and adaptations...

    Multi-Agent Risks from Advanced AI (Hammond2025) · AI · Unintentional · Post-deployment

  • Commitment and Trust

    "Commitment and trust (Section 3.5): difficulties in forming credible commitments, trust, or reputation can prevent mutual gains in AI-AI and human-AI interactions;"

    Multi-Agent Risks from Advanced AI (Hammond2025) · Other · Other · Post-deployment

  • Inefficient Outcomes

    "Inefficient Outcomes. Without careful planning and the appropriate safeguards, we may soon be entering a world overrun by increasingly competent and autonomous software agents, able to act with littl...

    Multi-Agent Risks from Advanced AI (Hammond2025) · AI · Unintentional · Post-deployment

  • Threats and Extortion

    "Threats and Extortion. A natural solution to problems of trust is to provide some kind of com- mitment ability to AI agents, which can be used to bind them to more cooperative courses of action. Unfo...

    Multi-Agent Risks from Advanced AI (Hammond2025) · AI · Intentional · Post-deployment

  • Rigidity and Mistaken Commitments

    "Rigidity and Mistaken Commitments. Even when it is desirable to be able to make threats in order to deter socially harmful behaviour, doing so using AI agents effectively removes the human from the l...

    Multi-Agent Risks from Advanced AI (Hammond2025) · Human · Unintentional · Post-deployment

  • Emergent Agency

    "Emergent agency (Section 3.6): qualitatively different goals or capabilities can emerge from the composition of innocuous independent systems or behaviours;"

    Multi-Agent Risks from Advanced AI (Hammond2025) · AI · Unintentional · Post-deployment

  • Emergent Capabilities

    "Emergent Capabilities. Dangerous emergent capabilities could arise when a multi-agent system over- comes the safety-enhancing limitations of the individual systems, such as individual models’ narrow...

    Multi-Agent Risks from Advanced AI (Hammond2025) · AI · Unintentional · Post-deployment

  • Emergent Goals

    "Emergent Goals. Ascribing goals to a system is not always straightforward. For our present purposes, it will suffice to adopt a Dennetian perspective (Dennett, 1971), ascribing goals and intentions o...

    Multi-Agent Risks from Advanced AI (Hammond2025) · AI · Unintentional · Post-deployment

  • Multi-Agent Security

    "Multi-agent security (Section 3.7): multi-agent systems give rise to new kinds of security threats and vulnerabilities."

    Multi-Agent Risks from Advanced AI (Hammond2025) · Other · Other · Other

  • Swarm Attacks

    "Swarm Attacks. The need for multi-agent security is foreshadowed by attacks today that benefit from the use of many decentralised agents, such as distributed denial-of-service attacks (Cisco, 2023; Y...

    Multi-Agent Risks from Advanced AI (Hammond2025) · Human · Intentional · Post-deployment

  • Heterogeneous Attacks

    "Heterogeneous Attacks. A closely related risk is the possibility of multiple agents combining different affordances to overcome safeguards, for which there is already preliminary evidence (Jones et a...

    Multi-Agent Risks from Advanced AI (Hammond2025) · AI · Intentional · Post-deployment

  • Social Engineering at Scale

    "Social Engineering at Scale. Advanced AI agents will be more easily able to interact with large numbers of humans, and vice versa. This provides a wider attack surface for various forms of automated...

    Multi-Agent Risks from Advanced AI (Hammond2025) · AI · Intentional · Post-deployment

  • Vulnerable AI Agents

    "Vulnerable AI Agents. The use of AI agents as delegates or representatives of humans or organisa- tions also introduces the possibility of attacks on AI agents themselves. In other words, agents can...

    Multi-Agent Risks from Advanced AI (Hammond2025) · Other · Intentional · Post-deployment

  • Cascading Security Failures

    "Cascading Security Failures. Localised attacks in multi-agent systems can result in catastrophic macroscopic outcomes (Motter & Lai, 2002, see also Sections 3.2 and 3.4). These cascades can be hard t...

    Multi-Agent Risks from Advanced AI (Hammond2025) · Other · Intentional · Post-deployment

  • Undetectable Threats

    "Undetectable Threats. Cooperation and trust in many multi-agent systems relies crucially on the ability to detect (and then avoid or sanction) adversarial actions taken by others (Ostrom, 1990; Schne...

    Multi-Agent Risks from Advanced AI (Hammond2025) · AI · Intentional · Post-deployment

  • Impact on Financial Stability

    "The integration of general-purpose AI into high-frequency trading, market-making, or systemic risk management could exacerbate systemic risk by exhibiting unexpected behavioral patterns during market...

    Frontier AI Risk Management Framework (v1.0) (Tse2025) · Other · Unintentional · Post-deployment

  • Multi-agent collaboration capability

    "Multiple autonomous AI agents able to establish collaborative relationships through explicit communication or implicit behavioral consistency, forming decentralized decision networks, jointly executi...

    Frontier AI Risk Management Framework (v1.0) (Tse2025) · AI · Intentional · Post-deployment