MIT AI Risk Repository · Risk Sub-Category · 63.10.04
Vulnerable AI Agents
Category: Multi-Agent Security
Description
"Vulnerable AI Agents. The use of AI agents as delegates or representatives of humans or organisa- tions also introduces the possibility of attacks on AI agents themselves. In other words, agents can be considered vulnerable extensions of their principals, introducing a novel attack surface (SecureWorks, 2023). Attacks on an AI agent could be used to extract private information about their principal (Wei & Liu, 2024; Wu et al., 2024a), or to manipulate the agent to take actions that the principal would find undesirable (Zhang et al., 2024a). This includes attacks that have direct relevance for
From Multi-Agent Risks from Advanced AI (Hammond2025), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Subdomain
- 7.6 Multi-agent risks
- Causal entity
- Other
- Intent
- Intentional
- Timing
- Post-deployment
Subdomain definition: Risks from multi-agent interactions, due to incentives (which can lead to conflict or collusion) and/or the structure of multi-agent systems, which can create cascading failures, selection pressures, new security vulnerabilities, and a lack of shared information and trust.
How other frameworks describe this risk
- Groups of LLM-Agents May Show Emergent Functionality
- Foundationality May Cause Correlated Failures
- Multi-Agent Safety Is Not Assured by Single-Agent Safety
- Collusion between LLM-Agents
- Financial instability due to model homogeneity
- Multi-agent collaboration capability
- Impact on Financial Stability
- Multi-agent collusion propensity: