{"attribution":{"source":"MIT AI Risk Repository, Domain Taxonomy of AI Risks v1 (MIT AI Risk Initiative)","license":"CC BY 4.0","license_url":"https://creativecommons.org/licenses/by/4.0/","citation":"Slattery, P., Saeri, A. K., Grundy, E. A. C., Graham, J., Noetel, M., Uuk, R., Dao, J., Pour, S., Casper, S., & Thompson, N. (2025). The AI Risk Repository: A comprehensive meta-review, database, and taxonomy of risks from artificial intelligence. arXiv:2408.12622."},"exported_at":"2026-09-12"}
{"rows":[{"ev_id":"65.08.01","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Training Data Risks (Robustness) ","risk_subcategory":"Data poisoning ","description":"\"A type of adversarial attack where an adversary or malicious insider injects intentionally corrupted, false, misleading, or incorrect samples into the training or fine-tuning datasets.\"","entity":"Human","intent":"Intentional","timing":"Pre-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"65.09.01","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Inference risks (Robustness) ","risk_subcategory":"Prompt injection attack ","description":"\"A prompt injection attack forces a generative model that takes a prompt as input to produce unexpected output by manipulating the structure, instructions, or information contained in its prompt.\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"65.09.02","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Inference risks (Robustness) ","risk_subcategory":"Extraction attack ","description":"\"An attribute inference attack is used to detect whether certain sensitive features can be inferred about individuals who participated in training a model. These attacks occur when an adversary has some prior knowledge about the training data and uses that knowledge to infer the sensitive data.\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"65.09.03","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Inference risks (Robustness) ","risk_subcategory":"Evasion attack ","description":"\"Evasion attacks attempt to make a model output incorrect results by slightly perturbing the input data that is sent to the trained model.\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"65.09.04","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Inference risks (Robustness) ","risk_subcategory":"Prompt leaking ","description":"\"A prompt leak attack attempts to extract a model's system prompt (also known as the system message).\"","entity":"Human","intent":"Intentional","timing":"Other","domain":2,"subdomain":"2.2"},{"ev_id":"65.10.01","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Inference risks (Multi-category) ","risk_subcategory":"Jailbreaking ","description":"\"A jailbreaking attack attempts to break through the guardrails that are established in the model to perform restricted actions.\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"65.10.02","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Inference risks (Multi-category) ","risk_subcategory":"Prompt priming ","description":"\"Because generative models tend to produce output like the input provided, the model can be prompted to reveal specific kinds of information. For example, adding personal information in the prompt increases its likelihood of generating similar kinds of personal information in its output. If personal data was included as part of the model’s training, there is a possibility it could be revealed.\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"65.11.01","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Inference risks (Privacy) ","risk_subcategory":"Membership inference attack ","description":"\"A membership inference attack repeatedly queries a model to determine whether a given input was part of the model’s training. More specifically, given a trained model and a data sample, an attacker samples the input space, observing outputs to deduce whether that sample was part of the model's training.\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"65.11.02","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Inference risks (Privacy) ","risk_subcategory":"Attribute inference attack ","description":"\"An attribute inference attack repeatedly queries a model to detect whether certain sensitive features can be inferred about individuals who participated in training a model. These attacks occur when an adversary has some prior knowledge about the training data and uses that knowledge to infer the sensitive data.\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"65.15.02","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Output risks (Value alignment)","risk_subcategory":"Harmful code generation ","description":"\"Models might generate code that causes harm or unintentionally affects other systems.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.2"}]}