MIT AI Risk Repository · Risk Sub-Category · 63.10.06
Undetectable Threats
Category: Multi-Agent Security
Description
"Undetectable Threats. Cooperation and trust in many multi-agent systems relies crucially on the ability to detect (and then avoid or sanction) adversarial actions taken by others (Ostrom, 1990; Schneier, 2012). Recent developments, however, have shown that AI agents are capable of both steganographic communication (Motwani et al., 2024; Schroeder de Witt et al., 2023b) and ‘illusory’ attacks (Franzmeyer et al., 2023), which are black-box undetectable and can even be hidden using white-box undetectable encrypted backdoors (Draguns et al., 2024). Similarly, in environments where agents learn fr
From Multi-Agent Risks from Advanced AI (Hammond2025), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Subdomain
- 7.6 Multi-agent risks
- Causal entity
- AI
- Intent
- Intentional
- Timing
- Post-deployment
Subdomain definition: Risks from multi-agent interactions, due to incentives (which can lead to conflict or collusion) and/or the structure of multi-agent systems, which can create cascading failures, selection pressures, new security vulnerabilities, and a lack of shared information and trust.
How other frameworks describe this risk
- Groups of LLM-Agents May Show Emergent Functionality
- Foundationality May Cause Correlated Failures
- Multi-Agent Safety Is Not Assured by Single-Agent Safety
- Collusion between LLM-Agents
- Financial instability due to model homogeneity
- Multi-agent collaboration capability
- Impact on Financial Stability
- Multi-agent collusion propensity: