MIT AI Risk Repository · Risk Category · 62.19.00
Attacks on GPAIs/GPAI Failure Modes
Description
"This section catalogs the risk sources related to GPAI failure modes or attacks targeting GPAIs. Many of these apply mainly to LLM-based GPAIs, which share some common failure modes such as jailbreaks and trojans. These vulnerabilities often extend beyond GPAIs and fall into the broader field of adversarial machine learning. However, additional vulnerabilities may arise with the introduction of new modalities, longer context windows, or different encodings."
From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).