MIT AI Risk Repository · Risk Sub-Category · 71.02.02
Malicious and Indirect
Category: User Intent
Description
"Benign intermediate for harmful end objective"
From Risks of AI Scientists: Prioritizing Safeguarding Over Autonomy (Tang2025), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Domain
- 4. Malicious actors
- Subdomain
- 4.0
- Causal entity
- Other
- Intent
- Intentional
- Timing
- Other