{"attribution":{"source":"MIT AI Risk Repository, Domain Taxonomy of AI Risks v1 (MIT AI Risk Initiative)","license":"CC BY 4.0","license_url":"https://creativecommons.org/licenses/by/4.0/","citation":"Slattery, P., Saeri, A. K., Grundy, E. A. C., Graham, J., Noetel, M., Uuk, R., Dao, J., Pour, S., Casper, S., & Thompson, N. (2025). The AI Risk Repository: A comprehensive meta-review, database, and taxonomy of risks from artificial intelligence. arXiv:2408.12622."},"exported_at":"2026-09-11"}
{"rows":[{"ev_id":"73.07.00","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Category","risk_category":"Jailbreaks and Prompt Injections Threaten Security of LLMs","risk_subcategory":null,"description":"\"LLMs are not adversarially robust and are vulnerable to security failures such as jailbreaks and prompt-injection attacks. While a number of jailbreak attacks have been proposed in the literature, the lack of standardized evaluation makes it difficult to compare them. We also do not have efficient white-box methods to evaluate adver- sarial robustness. Multi-modal LLMs may further allow novel types of jailbreaks via additional modalities. Finally, the lack of robust privilege levels within the LLM input means that jailbreaking and prompt-injection attacks may be particularly hard to eliminate","entity":"Other","intent":"Other","timing":"Other","domain":2,"subdomain":"2.2"},{"ev_id":"73.07.01","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"Jailbreaks and Prompt Injections Threaten Security of LLMs","risk_subcategory":"Exploiting Limited Generalization of Safety Finetuning","description":"\"Safety tuning is performed over a much narrower distribution compared to the pretraining distribution. This leaves the model vulnerable to attacks that exploit gaps in the generalization of the safety training, e.g. using encoded text (Wei et al., 2023c) or low-resource languages (Deng et al., 2023a; Yong et al., 2023) (see also Section 3.2).\"","entity":"Other","intent":"Unintentional","timing":"Other","domain":2,"subdomain":"2.2"},{"ev_id":"73.07.02","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"Jailbreaks and Prompt Injections Threaten Security of LLMs","risk_subcategory":"“Model Psychology” Attacks","description":"\"LLMs are vulnerable to “psychological” tricks (Li et al., 2023e; Shen et al., 2023), which can be exploited by attackers. Examples include instructing the model to behave like a specific persona (Shah et al., 2023; Andreas, 2022), or employing various “social engineering” tricks crafted by humans (Wei et al., 2023c) or other LLMs (Perez et al., 2022b; Casper et al., 2023c).\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"73.07.04","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"Jailbreaks and Prompt Injections Threaten Security of LLMs","risk_subcategory":"Attacking LLMs via Additional Modalities a","description":"\"LLMs can now process modalities other than text, e.g. images or video frames (OpenAI, 2023c; Gemini Team, 2023). Several studies show that gradient-based attacks on multimodal models are easy and effective (Carlini et al., 2023a; Bailey et al., 2023; Qi et al., 2023b). These attacks manipulate images that are input to the model (via an appropriate encoding). GPT-4Vision (OpenAI, 2023c) is vulnerable to jailbreaks and exfiltration attacks through much simpler means as well, e.g. writing jailbreaking text in the image (Willison, 2023a; Gong et al., 2023). For indirect prompt injection, the atta","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"73.08.00","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Category","risk_category":"Vulnerability to Poisoning and Backdoors","risk_subcategory":null,"description":"\"The previous section explored jailbreaks and other forms of adversarial prompts as ways to elicit harmful capabilities acquired during pretraining. These methods make no assumptions about the training data. On the other hand, poisoning attacks (Biggio et al., 2012) perturb training data to introduce specific vulnerabilities, called backdoors, that can then be exploited at inference time by the adversary. This is a challenging problem in current large language models because they are trained on data gathered from untrusted sources (e.g. internet), which can easily be poisoned by an adversary (","entity":"Human","intent":"Intentional","timing":"Pre-deployment","domain":2,"subdomain":"2.2"}]}