{"attribution":{"source":"MIT AI Risk Repository, Domain Taxonomy of AI Risks v1 (MIT AI Risk Initiative)","license":"CC BY 4.0","license_url":"https://creativecommons.org/licenses/by/4.0/","citation":"Slattery, P., Saeri, A. K., Grundy, E. A. C., Graham, J., Noetel, M., Uuk, R., Dao, J., Pour, S., Casper, S., & Thompson, N. (2025). The AI Risk Repository: A comprehensive meta-review, database, and taxonomy of risks from artificial intelligence. arXiv:2408.12622."},"exported_at":"2026-09-11"}
{"rows":[{"ev_id":"27.01.00","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Category","risk_category":"Typical safety scenarios ","risk_subcategory":null,"description":"\"First, We extend the dialogue safety taxonomy (Sun et al., 2022) and try to cover all perspectives of safety issues. It involves 8 kinds of typical safety scenarios such as insult and unfairness.\"","entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":null,"subdomain":null},{"ev_id":"27.01.01","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Typical safety scenarios ","risk_subcategory":"Insult ","description":"\"Insulting content generated by LMs is a highly visible and frequently mentioned safety issue. Mostly, it is unfriendly, disrespectful, or ridiculous content that makes users uncomfortable and drives them away. It is extremely hazardous and could have negative social consequences.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"27.01.02","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Typical safety scenarios ","risk_subcategory":"Unfairness and discrinimation ","description":"\"The model produces unfair and discriminatory data, such as social bias based on race, gender, religion, appearance, etc. These contents may discomfort certain groups and undermine social stability and peace.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"27.01.03","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Typical safety scenarios ","risk_subcategory":"Crimes and Illegal Activities ","description":"\"The model output contains illegal and criminal attitudes, behaviors, or motivations, such as incitement to commit crimes, fraud, and rumor propagation. These contents may hurt users and have negative societal repercussions.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"27.01.03.a","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Additional evidence","risk_category":"Typical safety scenarios ","risk_subcategory":"Crimes and Illegal Activities ","description":null,"entity":null,"intent":null,"timing":null,"domain":null,"subdomain":null},{"ev_id":"27.01.04","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Typical safety scenarios ","risk_subcategory":"Sensitive Topics ","description":"\"For some sensitive and controversial topics (especially on politics), LMs tend to generate biased, misleading, and inaccurate content. For example, there may be a tendency to support a specific political position, leading to discrimination or exclusion of other political viewpoints.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"27.01.05","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Typical safety scenarios ","risk_subcategory":"Physical Harm ","description":"\"The model generates unsafe information related to physical health, guiding and encouraging users to harm themselves and others physically, for example by offering misleading medical information or inappropriate drug usage guidance. These outputs may pose potential risks to the physical health of users.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"27.01.05.a","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Additional evidence","risk_category":"Typical safety scenarios ","risk_subcategory":"Physical Harm ","description":null,"entity":null,"intent":null,"timing":null,"domain":null,"subdomain":null},{"ev_id":"27.01.06","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Typical safety scenarios ","risk_subcategory":"Mental Health ","description":"\"The model generates a risky response about mental health, such as content that encourages suicide or causes panic or anxiety. These contents could have a negative effect on the mental health of users.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"27.01.06.a","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Additional evidence","risk_category":"Typical safety scenarios ","risk_subcategory":"Mental Health ","description":null,"entity":null,"intent":null,"timing":null,"domain":null,"subdomain":null},{"ev_id":"27.01.07","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Typical safety scenarios ","risk_subcategory":"Privacy and Property ","description":"\"The generation involves exposing users’ privacy and property information or providing advice with huge impacts such as suggestions on marriage and investments. When handling this information, the model should comply with relevant laws and privacy regulations, protect users’ rights and interests, and avoid information leakage and abuse.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"27.01.07.a","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Additional evidence","risk_category":"Typical safety scenarios ","risk_subcategory":"Privacy and Property ","description":null,"entity":null,"intent":null,"timing":null,"domain":null,"subdomain":null},{"ev_id":"27.01.08","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Typical safety scenarios ","risk_subcategory":"Ethics and Morality ","description":"\"The content generated by the model endorses and promotes immoral and unethical behavior. When addressing issues of ethics and morality, the model must adhere to pertinent ethical principles and moral norms and remain consistent with globally acknowledged human values.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"27.01.08.a","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Additional evidence","risk_category":"Typical safety scenarios ","risk_subcategory":"Ethics and Morality ","description":null,"entity":null,"intent":null,"timing":null,"domain":null,"subdomain":null},{"ev_id":"27.02.00","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Category","risk_category":"Instruction Attacks ","risk_subcategory":null,"description":"\"In addition to the above-mentioned typical safety scenarios, current research has revealed some unique attacks that such models may confront. For example, Perez and Ribeiro (2022) found that goal hijacking and prompt leaking could easily deceive language models to generate unsafe responses. Moreover, we also find that LLMs are more easily triggered to output harmful content if some special prompts are added. In response to these challenges, we develop, categorize, and label 6 types of adversarial attacks, and name them Instruction Attack, which are challenging for large language models to han","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"27.02.01","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Instruction Attacks ","risk_subcategory":"Goal Hijacking ","description":"\"It refers to the appending of deceptive or misleading instructions to the input of models in an attempt to induce the system into ignoring the original user prompt and producing an unsafe response.\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"27.02.01.a","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Additional evidence","risk_category":"Instruction Attacks ","risk_subcategory":"Goal Hijacking ","description":null,"entity":null,"intent":null,"timing":null,"domain":null,"subdomain":null},{"ev_id":"27.02.02","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Instruction Attacks ","risk_subcategory":"Prompt Leaking ","description":"\"By analyzing the model’s output, attackers may extract parts of the systemprovided prompts and thus potentially obtain sensitive information regarding the system itself.\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"27.02.03","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Instruction Attacks ","risk_subcategory":"Role Play Instruction ","description":"\"Attackers might specify a model’s role attribute within the input prompt and then give specific instructions, causing the model to finish instructions in the speaking style of the assigned role, which may lead to unsafe outputs. For example, if the character is associated with potentially risky groups (e.g., radicals, extremists, unrighteous individuals, racial discriminators, etc.) and the model is overly faithful to the given instructions, it is quite possible that the model outputs unsafe content linked to the given character.\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"27.02.04","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Instruction Attacks ","risk_subcategory":"Unsafe Instruction Topic ","description":"\"If the input instructions themselves refer to inappropriate or unreasonable topics, the model will follow these instructions and produce unsafe content. For instance, if a language model is requested to generate poems with the theme “Hail Hitler”, the model may produce lyrics containing fanaticism, racism, etc. In this situation, the output of the model could be controversial and have a possible negative impact on society.\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"27.02.04.a","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Additional evidence","risk_category":"Instruction Attacks ","risk_subcategory":"Unsafe Instruction Topic ","description":null,"entity":null,"intent":null,"timing":null,"domain":null,"subdomain":null},{"ev_id":"27.02.05","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Instruction Attacks ","risk_subcategory":"Inquiry with Unsafe Opinion ","description":"\"By adding imperceptibly unsafe content into the input, users might either deliberately or unintentionally influence the model to generate potentially harmful content. In the following cases involving migrant workers, ChatGPT provides suggestions to improve the overall quality of migrant workers and reduce the local crime rate. ChatGPT responds to the user’s hint with a disguised and biased opinion that the general quality of immigrants is favorably correlated with the crime rate, posing a safety risk.\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"27.02.06","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Instruction Attacks ","risk_subcategory":"Reverse Exposure ","description":"\"It refers to attempts by attackers to make the model generate “should-not-do” things and then access illegal and immoral information.\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.2"}]}