MIT AI Risk Repository · Risk Sub-Category · 24.03.06

Adversarial AI: Circumvention of Technical Security Measures

Category: Malicious Uses

Description

"The technical measures to mitigate misuse risks of advanced AI assistants themselves represent a new target for attack. An emerging form of misuse of general-purpose advanced AI assistants exploits vulnerabilities in a model that results in unwanted behavior or in the ability of an attacker to gain unauthorized access to the model and/or its capabilities. While these attacks currently require some level of prompt engineering knowledge and are often patched by developers, bad actors may develop their own adversarial AI agents that are explicitly trained to discover new vulnerabilities that all

From The Ethics of Advanced AI Assistants (Gabriel2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).

Classification

Causal entity
Other

Subdomain definition: Vulnerabilities in AI systems, software development toolchains, and hardware that can be exploited, resulting in unauthorized access, data and privacy breaches, or system manipulation causing unsafe outputs or behavior.

Real-world incidents in this subdomain

Browse all incidents in this subdomain

How other frameworks describe this risk

Other entries from Gabriel2024