{"attribution":{"source":"MIT AI Risk Repository, Domain Taxonomy of AI Risks v1 (MIT AI Risk Initiative)","license":"CC BY 4.0","license_url":"https://creativecommons.org/licenses/by/4.0/","citation":"Slattery, P., Saeri, A. K., Grundy, E. A. C., Graham, J., Noetel, M., Uuk, R., Dao, J., Pour, S., Casper, S., & Thompson, N. (2025). The AI Risk Repository: A comprehensive meta-review, database, and taxonomy of risks from artificial intelligence. arXiv:2408.12622."},"exported_at":"2026-09-12"}
{"rows":[{"ev_id":"64.04.01","quick_ref":"Marchal2024","paper_title":"Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data","level":"Risk Sub-Category","risk_category":"Misuse tactics to compromise GenAI systems (Model integrity) ","risk_subcategory":"Prompt injection ","description":"\"Prompt Injections are a form of Adversarial Input that involve manipulating the text instructions given to a GenAI system (Liu et al., 2023). Prompt Injections exploit loopholes in a model’s architec- tures that have no separation between system instructions and user data to produce a harmful output (Perez and Ribeiro, 2022). While researchers may use similar techniques to test the robustness of GenAI models, malicious actors can also leverage them. For example, they might flood a model with manipulative prompts to cause denial-of-service attacks or to bypass an AI detection software.\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"64.04.02","quick_ref":"Marchal2024","paper_title":"Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data","level":"Risk Sub-Category","risk_category":"Misuse tactics to compromise GenAI systems (Model integrity) ","risk_subcategory":"Adversarial input ","description":"\"Adversarial Inputs involve modifying individual input data to cause a model to malfunction. These modifications, which are often imperceptible to humans, exploit how the model makes decisions to produce errors (Wallace et al., 2019) and can be applied to text, but also to images, audio, or video (e.g. changing pixels in an image of a panda in a way that causes a model to label it as a gibbon).6\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"64.04.03","quick_ref":"Marchal2024","paper_title":"Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data","level":"Risk Sub-Category","risk_category":"Misuse tactics to compromise GenAI systems (Model integrity) ","risk_subcategory":"Jailbreaking ","description":"\"Jailbreaking aims to bypass or remove restrictions and safety filters placed on a GenAI model completely (Chao et al., 2023; Shen et al., 2023). This gives the actor free rein to generate any output, regardless of its content being harmful, biassed, or offensive. All three of these are tactics that manipulate the model into producing harmful outputs against its design. The difference is that prompt injections and adversarial inputs usually seek to steer the model towards producing harmful or incorrect outputs from one query, whereas jailbreaking seeks to dismantle a model’s safety mechanisms ","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"64.04.05","quick_ref":"Marchal2024","paper_title":"Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data","level":"Risk Sub-Category","risk_category":"Misuse tactics to compromise GenAI systems (Model integrity) ","risk_subcategory":"Model extraction ","description":"\"Data Exfiltration goes beyond revealing private information, and involves illicitly obtaining the training data used to build a model that may be sensitive or proprietary. Model Extraction is the same attack, only directed at the model instead of the training data — it involves obtaining the architecture, parameters, or hyper-parameters of a proprietary model (Carlini et al., 2024).\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"64.04.06","quick_ref":"Marchal2024","paper_title":"Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data","level":"Risk Sub-Category","risk_category":"Misuse tactics to compromise GenAI systems (Model integrity) ","risk_subcategory":"Steganography ","description":"\"Steganography is the practice of hiding coded messages in GenAI model outputs, which may allow malicious actors to communicate covertly.8\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"64.04.07","quick_ref":"Marchal2024","paper_title":"Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data","level":"Risk Sub-Category","risk_category":"Misuse tactics to compromise GenAI systems (Model integrity) ","risk_subcategory":"Poisoning ","description":"\"Data Poisoning involves deliberately corrupting a model’s training dataset to introduce vulnerabilities, derail its learning process, or cause it to make incorrect predictions (Carlini et al., 2023). For example, the tool Nightshade is a data poisoning tool, which allows artists to add invisible changes to the pixels in their art before uploading online, to break any models that use it for training.9 Such attacks exploit the fact that most GenAI models are trained on publicly available datasets like images and videos scraped from the web, which malicious actors can easily compromise.\"","entity":"Human","intent":"Intentional","timing":"Pre-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"64.05.00","quick_ref":"Marchal2024","paper_title":"Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data","level":"Risk Category","risk_category":"Misuse tactics to compromise GenAI systems (Data integrity) ","risk_subcategory":null,"description":"-","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"64.05.01","quick_ref":"Marchal2024","paper_title":"Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data","level":"Risk Sub-Category","risk_category":"Misuse tactics to compromise GenAI systems (Data integrity) ","risk_subcategory":"Privacy compromise ","description":"\"Privacy Compromise attacks reveal sensitive or private information that was used to train a model. For example, personally identifiable information or medical records.\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"64.05.02","quick_ref":"Marchal2024","paper_title":"Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data","level":"Risk Sub-Category","risk_category":"Misuse tactics to compromise GenAI systems (Data integrity) ","risk_subcategory":"Data exfiltration ","description":"\"Data Exfiltration goes beyond revealing private information, and involves illicitly obtaining the training data used to build a model that may be sensitive or proprietary. Model Extraction is the same attack, only directed at the model instead of the training data — it involves obtaining the architecture, parameters, or hyper-parameters of a proprietary model (Carlini et al., 2024).\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.2"}]}