MIT AI Risk Repository · domain 2: Privacy & Security

2.2 AI system security vulnerabilities and attacks

Vulnerabilities in AI systems, software development toolchains, and hardware that can be exploited, resulting in unauthorized access, data and privacy breaches, or system manipulation causing unsafe outputs or behavior.

Risk entries
112
Frameworks citing it
12
Recorded incidents
24
Incidents since 2020
20
Causal entity (risk entries)
Causal entity (risk entries) 87 0 Human: 87 Human 87 Other: 18 Other 18 AI: 6 AI 6 Not coded: 1 Not coded 1
Causal entity (risk entries)
LabelValue
Human87
Other18
AI6
Not coded1
Intent (risk entries)
Intent (risk entries) 83 0 Intentional: 83 Intentional 83 Unintentional: 16 Unintentional 16 Other: 12 Other 12 Not coded: 1 Not coded 1
Intent (risk entries)
LabelValue
Intentional83
Unintentional16
Other12
Not coded1
Timing (risk entries)
Timing (risk entries) 64 0 Post-deployment: 64 Post-deployment 64 Pre-deployment: 25 Pre-deployment 25 Other: 22 Other 22 Not coded: 1 Not coded 1
Timing (risk entries)
LabelValue
Post-deployment64
Pre-deployment25
Other22
Not coded1
Recorded incidents per yearIncident date; current year partial
Recorded incidents per year 7 0 2017: 3 2017 3 2019: 1 2019 1 2022: 2 2022 2 2023: 3 2023 3 2024: 3 2024 3 2025: 5 2025 5 2026: 7 2026 7
Recorded incidents per year
LabelValue
20173
20191
20222
20233
20243
20255
20267
Entries by levelRisk categories, subcategories and additional evidence coded to this subdomain
Entries by level 96 0 Risk Category: 16 Risk Category 16 Risk Sub-Category: 96 Risk Sub-Category 96
Entries by level
LabelValue
Risk Category16
Risk Sub-Category96
  • Privacy - Prompt Inversion Attack (PIA)

    "stealing the private prompting texts"

    A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy (Wang2025) · Human · Intentional · Post-deployment

  • Privacy - Attribute Inference Attack (AIA)

    "deducing the private or sensitive information from training texts, prompting texts or external texts"

    A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy (Wang2025) · Human · Intentional · Post-deployment

  • Privacy - Model Extraction Attack (MEA)

    "replicating the parameters of the LLM,"

    A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy (Wang2025) · Human · Intentional · Post-deployment

  • Jailbreak in LLM Malicious Use - Poisoning Training Data

    "In the data collecting and pre-training phase, malicious adversaries can Jailbreak LLMs through poisoning their training data to make the model to output harmful content."

    A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy (Wang2025) · Human · Intentional · Pre-deployment

  • Jailbreak in LLM Malicious Use - Backdoor Attack

    "However, there are still ones who can leave holes in the training dataset, making LLMs appear safe on average, but generate harmful content under other specific conditions. This kind of attack can be...

    A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy (Wang2025) · Human · Intentional · Pre-deployment

  • Jailbreak in LLM Malicious Use - White & Black Box Attacks

    "In the fine-tuning and alignment phase, elaborately- designed instruction datasets can be utilized to fine-tune LLMs to drive them to perform undesirable behaviors, such as generating harmful informa...

    A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy (Wang2025) · Human · Intentional · Pre-deployment

  • Jailbreak in LLM Malicious Use - Prompt Attacks

    "In the prompting and reasoning phase, dialog can push LLMs into confused or overly compliant states, raising the risk of producing harmful outputs when confronted with harmful questions. Most of the...

    A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy (Wang2025) · Human · Intentional · Post-deployment

  • Vulnerability of AI systems to attacks and misuse

    Governance of artificial intelligence: A risk and guideline-based integrative framework (Wirtz2022) · Other · Intentional · Other

  • On Purpose - Pre-Deployment

    "During the pre-deployment development stage, software may be subject to sabotage by someone with necessary access (a programmer, tester, even janitor) who for a number of possible reasons may alter s...

    Taxonomy of Pathways to Dangerous Artificial Intelligence (Yampolskiy2016) · Human · Intentional · Pre-deployment

  • Security risks (confidentiality)

    AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies (Zeng2024) · Other · Intentional · Post-deployment

  • Security risks (integrity)

    AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies (Zeng2024) · Other · Intentional · Other

  • Adversarial attack

    "Recent advances have shown that a deep learning model with high predictive accuracy frequently misbehaves on adversarial examples [57,58]. In particular, a small perturbation to an input image, which...

    Towards risk-aware artificial intelligence and machine learning systems: An overview (Zhang2022) · Human · Intentional · Other