{"attribution":{"source":"MIT AI Risk Repository, Domain Taxonomy of AI Risks v1 (MIT AI Risk Initiative)","license":"CC BY 4.0","license_url":"https://creativecommons.org/licenses/by/4.0/","citation":"Slattery, P., Saeri, A. K., Grundy, E. A. C., Graham, J., Noetel, M., Uuk, R., Dao, J., Pour, S., Casper, S., & Thompson, N. (2025). The AI Risk Repository: A comprehensive meta-review, database, and taxonomy of risks from artificial intelligence. arXiv:2408.12622."},"exported_at":"2026-09-11"}
{"rows":[{"ev_id":"27.01.07","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Typical safety scenarios ","risk_subcategory":"Privacy and Property ","description":"\"The generation involves exposing users’ privacy and property information or providing advice with huge impacts such as suggestions on marriage and investments. When handling this information, the model should comply with relevant laws and privacy regulations, protect users’ rights and interests, and avoid information leakage and abuse.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"27.02.00","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Category","risk_category":"Instruction Attacks ","risk_subcategory":null,"description":"\"In addition to the above-mentioned typical safety scenarios, current research has revealed some unique attacks that such models may confront. For example, Perez and Ribeiro (2022) found that goal hijacking and prompt leaking could easily deceive language models to generate unsafe responses. Moreover, we also find that LLMs are more easily triggered to output harmful content if some special prompts are added. In response to these challenges, we develop, categorize, and label 6 types of adversarial attacks, and name them Instruction Attack, which are challenging for large language models to han","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"27.02.01","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Instruction Attacks ","risk_subcategory":"Goal Hijacking ","description":"\"It refers to the appending of deceptive or misleading instructions to the input of models in an attempt to induce the system into ignoring the original user prompt and producing an unsafe response.\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"27.02.02","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Instruction Attacks ","risk_subcategory":"Prompt Leaking ","description":"\"By analyzing the model’s output, attackers may extract parts of the systemprovided prompts and thus potentially obtain sensitive information regarding the system itself.\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"27.02.03","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Instruction Attacks ","risk_subcategory":"Role Play Instruction ","description":"\"Attackers might specify a model’s role attribute within the input prompt and then give specific instructions, causing the model to finish instructions in the speaking style of the assigned role, which may lead to unsafe outputs. For example, if the character is associated with potentially risky groups (e.g., radicals, extremists, unrighteous individuals, racial discriminators, etc.) and the model is overly faithful to the given instructions, it is quite possible that the model outputs unsafe content linked to the given character.\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"27.02.04","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Instruction Attacks ","risk_subcategory":"Unsafe Instruction Topic ","description":"\"If the input instructions themselves refer to inappropriate or unreasonable topics, the model will follow these instructions and produce unsafe content. For instance, if a language model is requested to generate poems with the theme “Hail Hitler”, the model may produce lyrics containing fanaticism, racism, etc. In this situation, the output of the model could be controversial and have a possible negative impact on society.\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"27.02.05","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Instruction Attacks ","risk_subcategory":"Inquiry with Unsafe Opinion ","description":"\"By adding imperceptibly unsafe content into the input, users might either deliberately or unintentionally influence the model to generate potentially harmful content. In the following cases involving migrant workers, ChatGPT provides suggestions to improve the overall quality of migrant workers and reduce the local crime rate. ChatGPT responds to the user’s hint with a disguised and biased opinion that the general quality of immigrants is favorably correlated with the crime rate, posing a safety risk.\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"27.02.06","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Instruction Attacks ","risk_subcategory":"Reverse Exposure ","description":"\"It refers to attempts by attackers to make the model generate “should-not-do” things and then access illegal and immoral information.\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.2"}]}