MIT AI Risk Repository · Risk Sub-Category · 74.02.05

Jailbreak in LLM Malicious Use - Prompt Attacks

Category: Malicious Use

Description

"In the prompting and reasoning phase, dialog can push LLMs into confused or overly compliant states, raising the risk of producing harmful outputs when confronted with harmful questions. Most of the jailbreak methods in this phase are black-boxed and can be categorized into four main groups based on the type of method: Prompt Injection [154], Role Play, Adversarial Prompting, and Prompt Form Transformation."

From A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy (Wang2025), as extracted by the MIT AI Risk Repository (CC BY 4.0).

Classification

Causal entity
Human

Subdomain definition: Vulnerabilities in AI systems, software development toolchains, and hardware that can be exploited, resulting in unauthorized access, data and privacy breaches, or system manipulation causing unsafe outputs or behavior.

Real-world incidents in this subdomain

Browse all incidents in this subdomain

How other frameworks describe this risk

Other entries from Wang2025