MIT AI Risk Repository · Risk Sub-Category · 21.01.04

Adversarial attack

Category: Data-level risk

Description

"Recent advances have shown that a deep learning model with high predictive accuracy frequently misbehaves on adversarial examples [57,58]. In particular, a small perturbation to an input image, which is imperceptible to humans, could fool a well-trained deep learning model into making completely different predictions [23]."

From Towards risk-aware artificial intelligence and machine learning systems: An overview (Zhang2022), as extracted by the MIT AI Risk Repository (CC BY 4.0).

Classification

Causal entity
Human
Timing
Other

Subdomain definition: Vulnerabilities in AI systems, software development toolchains, and hardware that can be exploited, resulting in unauthorized access, data and privacy breaches, or system manipulation causing unsafe outputs or behavior.

Real-world incidents in this subdomain

Browse all incidents in this subdomain

How other frameworks describe this risk

Other entries from Zhang2022