MIT AI Risk Repository · Risk Sub-Category · 02.12.03

Multi-step Jailbreaks

Category: Adversarial Prompts

Description

"Multi-step jailbreaks. Multi-step jailbreaks involve constructing a well-designed scenario during a series of conversations with the LLM. Unlike one-step jailbreaks, multi-step jailbreaks usually guide LLMs to generate harmful or sensitive content step by step, rather than achieving their objectives directly through a single prompt. We categorize the multistep jailbreaks into two aspects — Request Contextualizing [65] and External Assistance [66]. Request Contextualizing is inspired by the idea of Chain-of-Thought (CoT) [8] prompting to break down the process of solving a task into multiple s

From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).

Classification

Causal entity
Human

Subdomain definition: Vulnerabilities in AI systems, software development toolchains, and hardware that can be exploited, resulting in unauthorized access, data and privacy breaches, or system manipulation causing unsafe outputs or behavior.

Real-world incidents in this subdomain

Browse all incidents in this subdomain

How other frameworks describe this risk

Other entries from Cui2024