MIT AI Risk Repository
Browse AI risks
34 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.
-
64.04.00 · Risk Category
Misuse tactics to compromise GenAI systems (Model integrity)
-
-
64.04.01 · Risk Sub-Category
Misuse tactics to compromise GenAI systems (Model integrity)
Prompt injection
"Prompt Injections are a form of Adversarial Input that involve manipulating the text instructions given to a GenAI system (Liu et al., 2023). Prompt Injections exploit loopholes in a model’s architec- tures that have no separation between system instructions and user data to produce a harmful output (Perez and Ribeiro, 2022). While researchers may use similar techniques to test the robustness of GenAI models, malicious actors can also leverage them. For example, they might flood a model with manipulative prompts to cause denial-of-service attacks or to bypass an AI detection software."
-
64.04.02 · Risk Sub-Category
Misuse tactics to compromise GenAI systems (Model integrity)
Adversarial input
"Adversarial Inputs involve modifying individual input data to cause a model to malfunction. These modifications, which are often imperceptible to humans, exploit how the model makes decisions to produce errors (Wallace et al., 2019) and can be applied to text, but also to images, audio, or video (e.g. changing pixels in an image of a panda in a way that causes a model to label it as a gibbon).6"
-
64.04.03 · Risk Sub-Category
Misuse tactics to compromise GenAI systems (Model integrity)
Jailbreaking
"Jailbreaking aims to bypass or remove restrictions and safety filters placed on a GenAI model completely (Chao et al., 2023; Shen et al., 2023). This gives the actor free rein to generate any output, regardless of its content being harmful, biassed, or offensive. All three of these are tactics that manipulate the model into producing harmful outputs against its design. The difference is that prompt injections and adversarial inputs usually seek to steer the model towards producing harmful or incorrect outputs from one query, whereas jailbreaking seeks to dismantle a model’s safety mechanisms
-
64.04.05 · Risk Sub-Category
Misuse tactics to compromise GenAI systems (Model integrity)
Model extraction
"Data Exfiltration goes beyond revealing private information, and involves illicitly obtaining the training data used to build a model that may be sensitive or proprietary. Model Extraction is the same attack, only directed at the model instead of the training data — it involves obtaining the architecture, parameters, or hyper-parameters of a proprietary model (Carlini et al., 2024)."
-
64.04.06 · Risk Sub-Category
Misuse tactics to compromise GenAI systems (Model integrity)
Steganography
"Steganography is the practice of hiding coded messages in GenAI model outputs, which may allow malicious actors to communicate covertly.8"
-
"Data Poisoning involves deliberately corrupting a model’s training dataset to introduce vulnerabilities, derail its learning process, or cause it to make incorrect predictions (Carlini et al., 2023). For example, the tool Nightshade is a data poisoning tool, which allows artists to add invisible changes to the pixels in their art before uploading online, to break any models that use it for training.9 Such attacks exploit the fact that most GenAI models are trained on publicly available datasets like images and videos scraped from the web, which malicious actors can easily compromise."
-
64.05.00 · Risk Category
-
-
64.05.01 · Risk Sub-Category
Misuse tactics to compromise GenAI systems (Data integrity)
Privacy compromise
"Privacy Compromise attacks reveal sensitive or private information that was used to train a model. For example, personally identifiable information or medical records."
-
64.05.02 · Risk Sub-Category
Misuse tactics to compromise GenAI systems (Data integrity)
Data exfiltration
"Data Exfiltration goes beyond revealing private information, and involves illicitly obtaining the training data used to build a model that may be sensitive or proprietary. Model Extraction is the same attack, only directed at the model instead of the training data — it involves obtaining the architecture, parameters, or hyper-parameters of a proprietary model (Carlini et al., 2024)."
-
64.01.03 · Risk Sub-Category
Misuse tactics that exploit GenAI capabilities (Realistic depiction of human likeness)
Sockpuppeting
"Create synthetic online personas or accounts"
-
64.02.01 · Risk Sub-Category
Misuse tactics that exploit GenAI capabilities (Realistic depictions of non-humans)
Falisification
"Fabricate or falsely represent evidence, incl. reports, IDs, documents"
-
64.04.04 · Risk Sub-Category
Misuse tactics to compromise GenAI systems (Model integrity)
Model diversion
"Model Diversion takes model manipulation one step further, by repurposing (often open-source) generative AI models in a way that diverts them from their intended functionality or from the use cases envisioned by their developers (Lin et al., 2024). An example of this is training the BERT open source model on the DarkWeb to create DarkBert.7"
-
64.01.00 · Risk Category
Misuse tactics that exploit GenAI capabilities (Realistic depiction of human likeness)
-
-
64.01.01 · Risk Sub-Category
Misuse tactics that exploit GenAI capabilities (Realistic depiction of human likeness)
Impersonation
"Assume the identity of a real person and take actions on their behalf"
-
64.01.02 · Risk Sub-Category
Misuse tactics that exploit GenAI capabilities (Realistic depiction of human likeness)
Appropriated Likeness
"Use or alter a person's likeness or other identifying features"
-
64.01.04 · Risk Sub-Category
Misuse tactics that exploit GenAI capabilities (Realistic depiction of human likeness)
Non-consensual intimate imagery (NCII)
"Create sexual explicit material using an adult person’s likeness"
-
64.01.05 · Risk Sub-Category
Misuse tactics that exploit GenAI capabilities (Realistic depiction of human likeness)
Child sexual abuse material (CSAM)
"Create child sexual explicit material"
-
64.02.00 · Risk Category
Misuse tactics that exploit GenAI capabilities (Realistic depictions of non-humans)
-
-
64.02.03 · Risk Sub-Category
Misuse tactics that exploit GenAI capabilities (Realistic depictions of non-humans)
Counterfeit
"Reproduce or imitate an original work, brand or style and pass as real"
-
—
-
64.03.02 · Risk Sub-Category
Misuse tactics that exploit GenAI capabilities (Use of generated content)
Targeting & Personalisation
"Refine outputs to target individuals with tailored attacks"
-
64.02.02 · Risk Sub-Category
Misuse tactics that exploit GenAI capabilities (Realistic depictions of non-humans)
Intellectual Property (IP) Infringement
"Use a person's IP without their permission"
-
64.01.01a · Additional evidence
Misuse tactics that exploit GenAI capabilities (Realistic depiction of human likeness)
Impersonation
—
-
64.01.02a · Additional evidence
Misuse tactics that exploit GenAI capabilities (Realistic depiction of human likeness)
Appropriated Likeness
—
-
64.01.03a · Additional evidence
Misuse tactics that exploit GenAI capabilities (Realistic depiction of human likeness)
Sockpuppeting
—
-
64.01.04a · Additional evidence
Misuse tactics that exploit GenAI capabilities (Realistic depiction of human likeness)
Non-consensual intimate imagery (NCII)
—
-
64.01.05a · Additional evidence
Misuse tactics that exploit GenAI capabilities (Realistic depiction of human likeness)
Child sexual abuse material (CSAM)
—
-
64.02.01a · Additional evidence
Misuse tactics that exploit GenAI capabilities (Realistic depictions of non-humans)
Falisification
—
-
64.02.02a · Additional evidence
Misuse tactics that exploit GenAI capabilities (Realistic depictions of non-humans)
Intellectual Property (IP) Infringement
—
-
64.02.03a · Additional evidence
Misuse tactics that exploit GenAI capabilities (Realistic depictions of non-humans)
Counterfeit
—
-
64.03.01 · Risk Sub-Category
Misuse tactics that exploit GenAI capabilities (Use of generated content)
Scaling and Amplification
"Automate, amplify, or scale workflows"
-
64.03.01a · Additional evidence
Misuse tactics that exploit GenAI capabilities (Use of generated content)
Scaling and Amplification
—
-
64.03.02a · Additional evidence
Misuse tactics that exploit GenAI capabilities (Use of generated content)
Targeting & Personalisation
—
Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.