MIT AI Risk Repository · Risk Sub-Category · 63.10.04

Vulnerable AI Agents

Category: Multi-Agent Security

Description

"Vulnerable AI Agents. The use of AI agents as delegates or representatives of humans or organisa- tions also introduces the possibility of attacks on AI agents themselves. In other words, agents can be considered vulnerable extensions of their principals, introducing a novel attack surface (SecureWorks, 2023). Attacks on an AI agent could be used to extract private information about their principal (Wei & Liu, 2024; Wu et al., 2024a), or to manipulate the agent to take actions that the principal would find undesirable (Zhang et al., 2024a). This includes attacks that have direct relevance for

From Multi-Agent Risks from Advanced AI (Hammond2025), as extracted by the MIT AI Risk Repository (CC BY 4.0).

Classification

Causal entity
Other

Subdomain definition: Risks from multi-agent interactions, due to incentives (which can lead to conflict or collusion) and/or the structure of multi-agent systems, which can create cascading failures, selection pressures, new security vulnerabilities, and a lack of shared information and trust.

How other frameworks describe this risk

Other entries from Hammond2025