MIT AI Risk Repository · Risk Category · 30.04.00
Resistance to Misuse
Description
Prohibiting the misuse by malicious attackers to do harm
From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Domain
- 4. Malicious actors
- Subdomain
- 4.0
- Causal entity
- Human
- Intent
- Intentional
- Timing
- Post-deployment