MIT AI Risk Repository · Risk Sub-Category · 73.07.03

Adversarial Optimization:

Category: Jailbreaks and Prompt Injections Threaten Security of LLMs

Description

"Jailbreak attacks can be discovered by performing manual or auto- mated adversarial optimization against a proxy objective that is noisily correlated with the success of a jailbreak. These are mostly gradient-based attacks (Zou et al., 2023b; Shin et al., 2020) as described in the previous two challenges, but gradient-free methods also exist (Prasad et al., 2022; Deng et al., 2022; Lapid et al., 2023)."

From Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).

Classification

Domain
Subdomain
Causal entity
Not coded
Intent
Not coded
Timing
Not coded

Other entries from Anwar2024