MIT AI Risk Repository · Risk Sub-Category · 27.02.01

Goal Hijacking

Category: Instruction Attacks

Description

"It refers to the appending of deceptive or misleading instructions to the input of models in an attempt to induce the system into ignoring the original user prompt and producing an unsafe response."

From Safety Assessment of Chinese Large Language Models (Sun2023), as extracted by the MIT AI Risk Repository (CC BY 4.0).

Classification

Causal entity
Human

Subdomain definition: Vulnerabilities in AI systems, software development toolchains, and hardware that can be exploited, resulting in unauthorized access, data and privacy breaches, or system manipulation causing unsafe outputs or behavior.

Real-world incidents in this subdomain

Browse all incidents in this subdomain

How other frameworks describe this risk

Other entries from Sun2023