MIT AI Risk Repository · Risk Sub-Category · 71.02.02

Malicious and Indirect

Category: User Intent

Description

"Benign intermediate for harmful end objective"

From Risks of AI Scientists: Prioritizing Safeguarding Over Autonomy (Tang2025), as extracted by the MIT AI Risk Repository (CC BY 4.0).

Classification

Subdomain
4.0
Causal entity
Other
Timing
Other

How other frameworks describe this risk

Other entries from Tang2025