MIT AI Risk Repository · domain 2: Privacy & Security

2.2 AI system security vulnerabilities and attacks

Vulnerabilities in AI systems, software development toolchains, and hardware that can be exploited, resulting in unauthorized access, data and privacy breaches, or system manipulation causing unsafe outputs or behavior.

Risk entries
112
Frameworks citing it
12
Recorded incidents
24
Incidents since 2020
20
Causal entity (risk entries)
Causal entity (risk entries) 87 0 Human: 87 Human 87 Other: 18 Other 18 AI: 6 AI 6 Not coded: 1 Not coded 1
Causal entity (risk entries)
LabelValue
Human87
Other18
AI6
Not coded1
Intent (risk entries)
Intent (risk entries) 83 0 Intentional: 83 Intentional 83 Unintentional: 16 Unintentional 16 Other: 12 Other 12 Not coded: 1 Not coded 1
Intent (risk entries)
LabelValue
Intentional83
Unintentional16
Other12
Not coded1
Timing (risk entries)
Timing (risk entries) 64 0 Post-deployment: 64 Post-deployment 64 Pre-deployment: 25 Pre-deployment 25 Other: 22 Other 22 Not coded: 1 Not coded 1
Timing (risk entries)
LabelValue
Post-deployment64
Pre-deployment25
Other22
Not coded1
Recorded incidents per yearIncident date; current year partial
Recorded incidents per year 7 0 2017: 3 2017 3 2019: 1 2019 1 2022: 2 2022 2 2023: 3 2023 3 2024: 3 2024 3 2025: 5 2025 5 2026: 7 2026 7
Recorded incidents per year
LabelValue
20173
20191
20222
20233
20243
20255
20267
Entries by levelRisk categories, subcategories and additional evidence coded to this subdomain
Entries by level 96 0 Risk Category: 16 Risk Category 16 Risk Sub-Category: 96 Risk Sub-Category 96
Entries by level
LabelValue
Risk Category16
Risk Sub-Category96
  • Vulnerabilities to jailbreaks exploiting long context windows (many- shot jailbreaking)

    "Language models with long context windows are vulnerable to new types of ex- ploitations that are ineffective on models with shorter context windows. While few-shot jailbreaking, which involves provi...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Other · Other · Post-deployment

  • Misuse of AI model by user-performed persuasion

    "AI models can be influenced to accept misinformation through persuasive conversations, even when their initial responses are factually correct. Multi-turn persuasion can be more effective than single...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Human · Intentional · Post-deployment

  • Non-decomissionability of models with open weights

    "If the model parameter weights are released or leaked in a security breach, the model cannot be decommissioned because the developer no longer has control over the publicly available model or its use...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Human · Other · Post-deployment

  • Interconnectivity with malicious external tools

    "The growing integration and interconnectivity with external tools and plugins increase the risk of exposure to malicious external inputs. This interconnectivity makes it easier for external tools to...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Human · Other · Post-deployment

  • Model weight leak

    "Model weights or access to them can be leaked when initial access is granted only to a select group of individuals, such as institutional researchers [209]. This risk can increase as more people gain...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Human · Intentional · Post-deployment

  • Technical vulnerabilities (Robustness - vulnerability to jailbreaking

    "Individuals can manipulate models into performing actions that violate the model’s usage restrictions—a phenomenon known as “jailbreaking.” These manipulations may result in causing the model to perf...

    Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024) · Human · Intentional · Post-deployment

  • Insufficient Security Measures

    Malicious entities can take advantage of weaknesses in AI algorithms to alter results, potentially resulting in tangible real-life impacts. Additionally, it’s vital to prioritize safeguarding privacy...

    Artificial Intelligence Trust, Risk and Security Management (AI TRiSM): Frameworks, Applications, Challenges and Future Research Directions (Habbal2024) · Human · Intentional · Post-deployment

  • Security - Robustness

    While AI safety focuses on threats emanating from generative AI systems, security centers on threats posed to these systems. The most extensively discussed issue in this context are jailbreaking risks...

    Mapping the Ethics of Generative AI: A Comprehensive Scoping Review (Hagendorff2024) · Human · Intentional · Other

  • Data poisoning

    "A type of adversarial attack where an adversary or malicious insider injects intentionally corrupted, false, misleading, or incorrect samples into the training or fine-tuning datasets."

    AI Risk Atlas (IBM2025) · Human · Intentional · Pre-deployment

  • Prompt injection attack

    "A prompt injection attack forces a generative model that takes a prompt as input to produce unexpected output by manipulating the structure, instructions, or information contained in its prompt."

    AI Risk Atlas (IBM2025) · Human · Intentional · Post-deployment

  • Extraction attack

    "An attribute inference attack is used to detect whether certain sensitive features can be inferred about individuals who participated in training a model. These attacks occur when an adversary has so...

    AI Risk Atlas (IBM2025) · Human · Intentional · Post-deployment

  • Evasion attack

    "Evasion attacks attempt to make a model output incorrect results by slightly perturbing the input data that is sent to the trained model."

    AI Risk Atlas (IBM2025) · Human · Intentional · Post-deployment

  • Prompt leaking

    "A prompt leak attack attempts to extract a model's system prompt (also known as the system message)."

    AI Risk Atlas (IBM2025) · Human · Intentional · Other

  • Jailbreaking

    "A jailbreaking attack attempts to break through the guardrails that are established in the model to perform restricted actions."

    AI Risk Atlas (IBM2025) · Human · Intentional · Post-deployment

  • Prompt priming

    "Because generative models tend to produce output like the input provided, the model can be prompted to reveal specific kinds of information. For example, adding personal information in the prompt inc...

    AI Risk Atlas (IBM2025) · Human · Intentional · Post-deployment

  • Membership inference attack

    "A membership inference attack repeatedly queries a model to determine whether a given input was part of the model’s training. More specifically, given a trained model and a data sample, an attacker s...

    AI Risk Atlas (IBM2025) · Human · Intentional · Post-deployment

  • Attribute inference attack

    "An attribute inference attack repeatedly queries a model to detect whether certain sensitive features can be inferred about individuals who participated in training a model. These attacks occur when...

    AI Risk Atlas (IBM2025) · Human · Intentional · Post-deployment

  • Harmful code generation

    "Models might generate code that causes harm or unintentionally affects other systems."

    AI Risk Atlas (IBM2025) · AI · Unintentional · Post-deployment

  • Prompt Attacks

    carefully controlled adversarial perturbation can flip a GPT model’s answer when used to classify text inputs. Furthermore, we find that by twisting the prompting question in a certain way, one can so...

    Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024) · Human · Intentional · Other

  • Poisoning Attacks

    fool the model by manipulating the training data, usually performed on classification models

    Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024) · Human · Intentional · Pre-deployment

  • Prompt injection

    "Prompt Injections are a form of Adversarial Input that involve manipulating the text instructions given to a GenAI system (Liu et al., 2023). Prompt Injections exploit loopholes in a model’s architec...

    Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data (Marchal2024) · Human · Intentional · Post-deployment

  • Adversarial input

    "Adversarial Inputs involve modifying individual input data to cause a model to malfunction. These modifications, which are often imperceptible to humans, exploit how the model makes decisions to prod...

    Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data (Marchal2024) · Human · Intentional · Post-deployment

  • Jailbreaking

    "Jailbreaking aims to bypass or remove restrictions and safety filters placed on a GenAI model completely (Chao et al., 2023; Shen et al., 2023). This gives the actor free rein to generate any output,...

    Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data (Marchal2024) · Human · Intentional · Post-deployment

  • Model extraction

    "Data Exfiltration goes beyond revealing private information, and involves illicitly obtaining the training data used to build a model that may be sensitive or proprietary. Model Extraction is the sam...

    Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data (Marchal2024) · Human · Intentional · Post-deployment

  • Steganography

    "Steganography is the practice of hiding coded messages in GenAI model outputs, which may allow malicious actors to communicate covertly.8"

    Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data (Marchal2024) · Human · Intentional · Post-deployment