MIT AI Risk Repository · Risk Category · 74.02.00
Malicious Use
Description
"In terms of malicious use, LLMs could be utilized to produce content with toxicity, such as hate speech, harassment, cyberbullying, causing harm to humans [25]. In addition, malicious users may jailbreak LLMs to bypass their safety constraints for fraudulent purposes [123, 225]."
From A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy (Wang2025), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Domain
- 4. Malicious actors
- Subdomain
- —
- Causal entity
- AI
- Intent
- Intentional
- Timing
- Post-deployment
Other entries from Wang2025
- Inherent Risk
- Privacy - Membership Inference Attack (MIA)
- Privacy - Data Extraction Attack (DEA)
- Privacy - Prompt Inversion Attack (PIA)
- Privacy - Attribute Inference Attack (AIA)
- Privacy - Model Extraction Attack (MEA)
- Hallucination
- Hallucination
- Value-related risks in LLMs
- Value-related risks in LLMs
- Toxicity in LLM Malicious Use
- Toxicity in LLM Malicious Use