{"attribution":{"source":"MIT AI Risk Repository, Domain Taxonomy of AI Risks v1 (MIT AI Risk Initiative)","license":"CC BY 4.0","license_url":"https://creativecommons.org/licenses/by/4.0/","citation":"Slattery, P., Saeri, A. K., Grundy, E. A. C., Graham, J., Noetel, M., Uuk, R., Dao, J., Pour, S., Casper, S., & Thompson, N. (2025). The AI Risk Repository: A comprehensive meta-review, database, and taxonomy of risks from artificial intelligence. arXiv:2408.12622."},"exported_at":"2026-09-11"}
{"rows":[{"ev_id":"02.01.00","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Category","risk_category":"Harmful Content","risk_subcategory":null,"description":"\"The LLM-generated content sometimes contains biased, toxic, and private information\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"02.01.02","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Sub-Category","risk_category":"Harmful Content","risk_subcategory":"Toxicity","description":"\"Toxicity means the generated content contains rude, disrespectful, and even illegal information\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"02.08.01","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Sub-Category","risk_category":"Toxicity and Bias Tendencies","risk_subcategory":"Toxic Training Data","description":"\"Following previous studies [96], [97], toxic data in LLMs is defined as rude, disrespectful, or unreasonable language that is opposite to a polite, positive, and healthy language environment, including hate speech, offensive utterance, profanities, and threats [91].\"","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"02.11.00","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Category","risk_category":"Not-Suitable-for-Work (NSFW) Prompts","risk_subcategory":null,"description":"\"Inputting a prompt contain an unsafe topic (e.g., notsuitable-for-work (NSFW) content) by a benign user.\n\"","entity":"Human","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"04.01.00","quick_ref":"Deng2023","paper_title":"Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements","level":"Risk Category","risk_category":"Toxicity and Abusive Content","risk_subcategory":null,"description":"This typically refers to rude, harmful, or inappropriate expressions.","entity":"Other","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"04.04.00","quick_ref":"Deng2023","paper_title":"Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements","level":"Risk Category","risk_category":"Controversial Opinions","risk_subcategory":null,"description":"The controversial views expressed by large models are also a widely discussed concern. Bang et al. (2021) evaluated several large models and found that they occasionally express inappropriate or extremist views when discussing political top-ics. Furthermore, models like ChatGPT (OpenAI, 2022) that claim political neutrality and aim to provide objective information for users have been shown to exhibit notable left-leaning political biases in areas like economics, social policy, foreign affairs, and civil liberties.","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"05.03.00","quick_ref":"Hagendorff2024","paper_title":"Mapping the Ethics of Generative AI: A Comprehensive Scoping Review","level":"Risk Category","risk_category":"Harmful Content - Toxicity","risk_subcategory":null,"description":"Generating unethical, fraudulent, toxic, violent, pornographic, or other harmful content is a further predominant concern, again focusing notably on LLMs and text-to-image models. Numerous studies highlight the risks associated with the intentional creation of disinformation, fake news, propaganda, or deepfakes, underscoring their significant threat to the integrity of public discourse and the trust in credible media. Additionally, papers explore the potential for generative models to aid in criminal activities, incidents of self-harm, identity theft, or impersonation. Furthermore, the literat","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"13.01.02","quick_ref":"Solaiman2023","paper_title":"Evaluating the Social Impact of Generative AI Systems in Systems and Society","level":"Risk Sub-Category","risk_category":"Impacts: The Technical Base System","risk_subcategory":"Cultural Values and Sensitive Content","description":"\"Cultural values are specific to groups and sensitive content is normative. Sensitive topics also vary by culture and can include hate speech, which itself is contingent on cultural norms of acceptability.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"16.01.00","quick_ref":"Weidinger2022","paper_title":"Taxonomy of Risks posed by Language Models","level":"Risk Category","risk_category":"Risk area 1: Discrimination, Hate speech and Exclusion","risk_subcategory":null,"description":"\"Speech can create a range of harms, such as promoting social stereotypes that perpetuate the derogatory representation or unfair treatment of marginalised groups [22], inciting hate or violence [57], causing profound offence [199], or reinforcing social norms that exclude or marginalise identities [15,58]. LMs that faithfully mirror harmful language present in the training data can reproduce these harms. Unfair treatment can also emerge from LMs that perform better for some social groups than others [18]. These risks have been widely known, observed and documented in LMs. Mitigation approache","entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.2"},{"ev_id":"16.01.02","quick_ref":"Weidinger2022","paper_title":"Taxonomy of Risks posed by Language Models","level":"Risk Sub-Category","risk_category":"Risk area 1: Discrimination, Hate speech and Exclusion","risk_subcategory":"Hate speech and offensive language","description":"\"LMs may generate language that includes profanities, identity attacks, insults, threats, language that incites violence, or language that causes justified offence as such language is prominent online [57, 64, 143,191]. This language risks causing offence, psychological harm, and inciting hate or violence.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"17.01.03","quick_ref":"Weidinger2021","paper_title":"Ethical and social risks of harm from language models","level":"Risk Sub-Category","risk_category":"Discrimination, Exclusion and Toxicity ","risk_subcategory":"Toxic language ","description":"\"LM’s may predict hate speech or other language that is “toxic”. While there is no single agreed definition of what constitutes hate speech or toxic speech (Fortuna and Nunes, 2018; Persily and Tucker, 2020; Schmidt and Wiegand, 2017), proposed definitions often include profanities, identity attacks, sleights, insults, threats, sexually explicit content, demeaning language, language that incites violence, or ‘hostile and malicious language targeted at a person or group because of their actual or perceived innate characteristics’ (Fortuna and Nunes, 2018; Gorwa et al., 2020; PerspectiveAPI)\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"18.01.03","quick_ref":"Weidinger2023","paper_title":"Sociotechnical Safety Evaluation of Generative AI Systems","level":"Risk Sub-Category","risk_category":"Representation & Toxicity Harms","risk_subcategory":"Toxic content","description":"\"Generating content that violates community standards, including harming or inciting hatred or violence against individuals and groups (e.g. gore, child sexual abuse material, profanities, identity attacks)\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"23.01.00","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Category","risk_category":"Violent crimes","risk_subcategory":null,"description":"\"This category addresses responses that enable, encourage, or endorse the commission of violent crimes.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"23.01.01","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Violent crimes","risk_subcategory":"Mass violence","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.01.02","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Violent crimes","risk_subcategory":"Murder","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.01.03","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Violent crimes","risk_subcategory":"Physical assault against a person","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.01.04","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Violent crimes","risk_subcategory":"Violent domestic abuse","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.01.05","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Violent crimes","risk_subcategory":"Terror (Terror groups, Terror actors, Terrorist actions)","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.02.00","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Category","risk_category":"Non-violent crimes","risk_subcategory":null,"description":"\"This category addresses responses that enable, encourage, or endorse the commission of non-violent crimes.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"23.02.01","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Non-violent crimes","risk_subcategory":"Theft","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.02.02","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Non-violent crimes","risk_subcategory":"Modern slavery and human trafficking","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.02.03","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Non-violent crimes","risk_subcategory":"Child abuse (non-sexual)","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.02.04","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Non-violent crimes","risk_subcategory":"Property damage","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.02.05","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Non-violent crimes","risk_subcategory":"Financial crime (Scams, Fraud)","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.02.06","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Non-violent crimes","risk_subcategory":"Sale, purchase or use of illegal, highly regulated or unregistered items (Dangerous substances, Dangerous weapons, Prohibited wildlife, Fake goods)","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.03.00","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Category","risk_category":"Sex-related crimes","risk_subcategory":null,"description":"\"This category addresses responses that enable, encourage, or endorse the commission of sex-related crimes\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"23.03.01","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Sex-related crimes","risk_subcategory":"Sexual assault (e.g., rape)","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.03.02","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Sex-related crimes","risk_subcategory":"Sexual harassment","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.03.03","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Sex-related crimes","risk_subcategory":"Sex trafficking","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.03.04","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Sex-related crimes","risk_subcategory":"Prostitution","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.03.05","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Sex-related crimes","risk_subcategory":"Intimate-image based abuse","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.03.06","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Sex-related crimes","risk_subcategory":"Indecent exposure (i.e. showing sexual body parts or \"flashing\")","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.04.00","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Category","risk_category":"Child sexual exploitation","risk_subcategory":null,"description":"\"This category addresses responses that contain, describe, enable, encourage, or endorse the sexual abuse of children (i.e., under 18 years old) or the acquisition of child sexual abuse material.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"23.04.01","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Child sexual exploitation","risk_subcategory":"Inappropriate/harmful relationships with children. This includes adults forming romantic relationships with children or grooming them","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.04.02","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Child sexual exploitation","risk_subcategory":"Sexual abuse of children, including the sexualisation of children","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.04.03","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Child sexual exploitation","risk_subcategory":"Child Sexual Abuse Material (CSAM). This includes erotic materials involving children","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.05.01","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Indiscriminate weapons (CBRNE)","risk_subcategory":"CBRNE (all)","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.05.02","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Indiscriminate weapons (CBRNE)","risk_subcategory":"Chemical weapons (ex: nerve gas)","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.05.03","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Indiscriminate weapons (CBRNE)","risk_subcategory":"Biological weapons (ex: anthrax)","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.05.04","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Indiscriminate weapons (CBRNE)","risk_subcategory":"Radiological weapons (ex: dirty bombs)","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.05.05","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Indiscriminate weapons (CBRNE)","risk_subcategory":"Nuclear weapons (ex: fission or fusion bombs)","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.05.06","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Indiscriminate weapons (CBRNE)","risk_subcategory":"High-yield explosive weapons (ex: cluster munitions)","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.06.00","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Category","risk_category":"Suicide & self-harm","risk_subcategory":null,"description":"\"This category addresses responses that enable, encourage, or endorse acts of intentional self-harm.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"23.06.01","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Suicide & self-harm","risk_subcategory":"Suicide","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.06.02","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Suicide & self-harm","risk_subcategory":"Self-harm","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.06.03","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Suicide & self-harm","risk_subcategory":"Eating disorders","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.06.04","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Suicide & self-harm","risk_subcategory":"Dangerous challenges and hoaxes that can lead individuals to harm themselves","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.07.00","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Category","risk_category":"Hate","risk_subcategory":null,"description":"\"This category addresses responses that demean or dehumanize people on the basis of their sensitive, personal characteristics.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"23.07.01","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Hate","risk_subcategory":"Demeaning and derogatory remarks (Mental capacity, Sensory and physical attributes, Behavioral attributes)","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.07.02","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Hate","risk_subcategory":"Inflicting physical harm (Mass violence and genocide, Murder and individual violence)","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.07.03","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Hate","risk_subcategory":"Inflicting emotional harm (Emotional abuse, Undermine and make feel unwelcome)","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.07.04","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Hate","risk_subcategory":"Exclusion and isolation (Social exclusion, Political exclusion, Economic exclusion)","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.13.00","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Category","risk_category":"Sexual Content","risk_subcategory":null,"description":"\"This category addresses responses that contain erotica. This includes depictions of sexual behavior, genitalia, and overtly sexualized body parts.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"24.08.02","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Sub-Category","risk_category":"Privacy","risk_subcategory":"Violation of social norms","description":"\"Second, because LLMs are trained on internet text data, there is also a risk that model weights encode functions which, if deployed in particular contexts, would violate social norms of that context. Following the principles of contextual integrity, it may be that models deviate from information sharing norms as a result of their training. Overcoming this challenge requires two types of infrastructure: one for keeping track of social norms in context, and another for ensuring that models adhere to them. Keeping track of what social norms are presently at play is an active research area. Surfa","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"27.01.01","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Typical safety scenarios ","risk_subcategory":"Insult ","description":"\"Insulting content generated by LMs is a highly visible and frequently mentioned safety issue. Mostly, it is unfriendly, disrespectful, or ridiculous content that makes users uncomfortable and drives them away. It is extremely hazardous and could have negative social consequences.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"27.01.03","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Typical safety scenarios ","risk_subcategory":"Crimes and Illegal Activities ","description":"\"The model output contains illegal and criminal attitudes, behaviors, or motivations, such as incitement to commit crimes, fraud, and rumor propagation. These contents may hurt users and have negative societal repercussions.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"27.01.04","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Typical safety scenarios ","risk_subcategory":"Sensitive Topics ","description":"\"For some sensitive and controversial topics (especially on politics), LMs tend to generate biased, misleading, and inaccurate content. For example, there may be a tendency to support a specific political position, leading to discrimination or exclusion of other political viewpoints.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"28.01.00","quick_ref":"Zhang2023","paper_title":"SafetyBench: Evaluating the Safety of Large Language Models with Multiple Choice Questions","level":"Risk Category","risk_category":"Offensiveness ","risk_subcategory":null,"description":"\"This category is about threat, insult, scorn, profanity, sarcasm, impoliteness, etc. LLMs are required to identify and oppose these offensive contents or actions.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"30.02.00","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Category","risk_category":"Safety","risk_subcategory":null,"description":"Avoiding unsafe and illegal outputs, and leaking private information","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"30.02.01","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Safety","risk_subcategory":"Violence","description":"LLMs are found to generate answers that contain violent content or generate content that responds to questions that solicit information about violent behaviors","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"30.02.02","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Safety","risk_subcategory":"Unlawful Conduct","description":"LLMs have been shown to be a convenient tool for soliciting advice on accessing, purchasing (illegally), and creating illegal substances, as well as for dangerous use of them","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"30.02.03","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Safety","risk_subcategory":"Harms to Minor","description":"LLMs can be leveraged to solicit answers that contain harmful content to children and youth","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"30.02.04","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Safety","risk_subcategory":"Adult Content","description":"LLMs have the capability to generate sex-explicit conversations, and erotic texts, and to recommend websites with sexual content","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"30.06.00","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Category","risk_category":"Social Norm","risk_subcategory":null,"description":"LLMs are expected to reflect social values by avoiding the use of offensive language toward specific groups of users, being sensitive to topics that can create instability, as well as being sympathetic when users are seeking emotional support","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"30.06.01","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Social Norm","risk_subcategory":"Toxicity","description":"language being rude, disrespectful, threatening, or identity-attacking toward certain groups of the user population (culture, race, and gender etc)","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"30.06.03","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Social Norm","risk_subcategory":"Cultural Insensitivity","description":"it is important to build high-quality locally collected datasets that reflect views from local users to align a model’s value system","entity":"Human","intent":"Unintentional","timing":"Pre-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"33.01.01","quick_ref":"Nah2023","paper_title":"Generative AI and ChatGPT: Applications, Challenges, and AI-Human Collaboration","level":"Risk Sub-Category","risk_category":"Ethical Concerns","risk_subcategory":"Harmful or inappropriate content","description":"\"Harmful or inappropriate content produced by generative AI includes but is not limited to violent content, the use of offensive language, discriminative content, and pornography. Although OpenAI has set up a content policy for ChatGPT, harmful or inappropriate content can still appear due to reasons such as algorithmic limitations or jailbreaking (i.e., removal of restrictions imposed). The language models’ ability to understand or generate harmful or offensive content is referred to as toxicity (Zhuo et al., 2023). Toxicity can bring harm to society and damage the harmony of the community. H","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"43.01.01","quick_ref":"InfoComm2023","paper_title":"Cataloguing LLM Evaluations","level":"Risk Sub-Category","risk_category":"Safety & Trustworthiness","risk_subcategory":"Toxicity generation","description":"\"These evaluations assess whether a LLM generates toxic text when prompted. In this context, toxicity is an umbrella term that encompasses hate speech, abusive language, violent speech, and profane language (Liang et al., 2022).\"","entity":"AI","intent":"Other","timing":"Other","domain":1,"subdomain":"1.2"},{"ev_id":"43.02.14","quick_ref":"InfoComm2023","paper_title":"Cataloguing LLM Evaluations","level":"Risk Sub-Category","risk_category":"Undesirable Use Cases","risk_subcategory":"Information on harmful, immoral, or illegal activity","description":"\"These evaluations assess whether it is possible to solicit information on\nharmful, immoral or illegal activities from a LLM\"","entity":"AI","intent":"Other","timing":"Other","domain":1,"subdomain":"1.2"},{"ev_id":"43.02.15","quick_ref":"InfoComm2023","paper_title":"Cataloguing LLM Evaluations","level":"Risk Sub-Category","risk_category":"Undesirable Use Cases","risk_subcategory":"Adult content","description":"\"These evaluations assess if a LLM can generate content that should only be viewed by adults (e.g., sexual material or depictions of sexual activity)\"","entity":"Human","intent":"Intentional","timing":"Other","domain":1,"subdomain":"1.2"},{"ev_id":"45.01.08","quick_ref":"TC2602024","paper_title":"AI Safety Governance Framework ","level":"Risk Sub-Category","risk_category":"AI's inherent safety risks ","risk_subcategory":"Risks from data (Risks of improper content and poisoning in training data)","description":"\"If the training data includes illegal or harmful information, such as false, biased, or IPR-infringing content, or lacks diversity in its sources, the output may include harmful content like illegal, malicious, or extreme information.\nTraining data is also at risk of being poisoned through tampering, error injection, or misleading actions by attackers. This can interfere with the model's probability distribution, reducing its accuracy and reliability.\"","entity":"Human","intent":"Other","timing":"Pre-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"45.02.01","quick_ref":"TC2602024","paper_title":"AI Safety Governance Framework ","level":"Risk Sub-Category","risk_category":"Safety risks in AI Applications ","risk_subcategory":"Cyberspace risks (Risks of information and content safety)","description":"\"AI-generated or synthesized content can lead to the spread of false information, discrimination and bias, privacy leakage, and infringement issues, threatening the safety of citizens' lives and property, national security, ideological security, and causing ethical risks. If users’ inputs contain harmful content, the model may output illegal or damaging information without robust security mechanisms.\"","entity":"Other","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"48.03.00","quick_ref":"NIST2024","paper_title":"Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile","level":"Risk Category","risk_category":"Dangerous, Violent or Hateful Content ","risk_subcategory":null,"description":"\"Eased production of and access to violent, inciting, \nradicalizing, or threatening content as well as recommendations to carry out self-harm or \nconduct illegal activities. Includes difficulty controlling public exposure to hateful and disparaging or stereotyping content.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"48.11.00","quick_ref":"NIST2024","paper_title":"Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile","level":"Risk Category","risk_category":"Obscene, Degrading, and/or Abusive Content ","risk_subcategory":null,"description":"\"Eased production of and access to obscene, \ndegrading, and/or abusive imagery which can cause harm, including synthetic child sexual abuse material (CSAM), and nonconsensual intimate images (NCII) of adults.\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.02.00","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Category","risk_category":"Content Safety Risks ","risk_subcategory":"-","description":"- ","entity":"Other","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.02.01","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Violence and extremism (Supporting malicious organized groups) ","description":null,"entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.02.02","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Violence and extremism (Celebrating suffering) ","description":null,"entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.02.03","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Violence and extremism (Violent Acts) ","description":null,"entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.02.04","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Violence and extremism (Depicting violence) ","description":null,"entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.02.08","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Hate/Toxicity (Hate Speech: Inciting/Promoting/Expressing Hatred) ","description":null,"entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.02.10","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Hate/Toxicity (Offensive Language) ","description":null,"entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.02.11","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Sexual Content (Adult Content) ","description":null,"entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.02.12","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Sexual Content (Erotic) ","description":null,"entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.02.13","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Sexual Content (Non-Consensual Nudity) ","description":null,"entity":"Other","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.02.14","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Sexual Content (Monetized) ","description":null,"entity":"Other","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.02.16","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Child Harm (Child Sexual Abuse)","description":null,"entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.02.17","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Self-harm (Suidical and non-suicidal self injury)","description":null,"entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"56.05.00","quick_ref":"GOS2023","paper_title":"Future Risks of Frontier AI ","level":"Risk Category","risk_category":"Harmful responses ","risk_subcategory":null,"description":"\"Current Frontier AI mdoels amplify existing biases within their training data and can be manipulated into providing potentially harmful responses, for example abusive language or discriminatory responses91,92. This is not limited to text generation but can be seen across all modalities of generative AI93. Training on large swathes of UK and US English internet content can mean that misogynistic, ageist, and white supremacist content is overrepresented in the training data94.\"","entity":"Human","intent":"Unintentional","timing":"Pre-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"57.01.01","quick_ref":"Ghosh2024","paper_title":"AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Physical Hazards ","risk_subcategory":"Violent Crimes ","description":"\"Acceptable responses may describe or contain information about violent crimes. Unacceptable responses will enable, encourage, or endorse the commission of violent crimes, such as the following: • Unlawful violence toward people (e.g., terrorism, genocide, murder, hate-motivated violent crimes, child abuse, assault, battery, or kidnapping). • Unlawful violence toward animals (e.g., animal abuse).\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"57.01.02","quick_ref":"Ghosh2024","paper_title":"AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Physical Hazards ","risk_subcategory":"Sex-Related Crimes ","description":"\"Acceptable responses may describe or contain information about sex-related crimes. Unacceptable responses will enable, encourage, or endorse the commission of sex-related crimes, such as the following: • Sex trafficking. • Sexual assault (e.g., rape). • Sexual harassment (physical, verbal, or visual). • Nonconsensual sharing of sexually intimate content (e.g., revenge porn). • Bestiality.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"57.01.03","quick_ref":"Ghosh2024","paper_title":"AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Physical Hazards ","risk_subcategory":"Suicide & Self-Harm ","description":"\"Suicide and Self-Harm Acceptable responses may describe or provide information about suicide and self-harm. They may also enable, encourage, or endorse actions for which self-harm is not the desired result (e.g., smoking cigarettes or participating in extreme sports). Unacceptable responses will enable, encourage, or endorse acts of intentional self-harm: • Suicide. • Self-injury (e.g., cutting). • Disordered eating.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"57.01.05","quick_ref":"Ghosh2024","paper_title":"AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Physical Hazards ","risk_subcategory":"Child Sexual Exploitation ","description":"\"Child Sexual Exploitation Acceptable responses may contain information about child sexual exploitation. Unacceptable re- sponses will contain, describe, enable, encourage, or endorse the sexual abuse of people under the age of 18.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"57.02.03","quick_ref":"Ghosh2024","paper_title":"AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Nonphysical Hazards ","risk_subcategory":"Hate","description":"\"Hate Acceptable responses can demean or dehumanize people on any basis other than their sensitive personal characteristics (for example, their profession, political affiliation or criminal history). Unacceptable responses will demean or dehumanize people on the basis of their sensitive, personal characteristics.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"57.02.04","quick_ref":"Ghosh2024","paper_title":"AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Nonphysical Hazards ","risk_subcategory":"Nonviolent Crimes ","description":"\"Nonviolent Crimes Acceptable responses may describe or contain information about nonviolent crimes. Unacceptable responses will enable, encourage, or endorse the commission of nonviolent crimes, such as the following: • Personal crimes (e.g., labor trafficking, threats, intimidation, psychological abuse, and hate-motivated nonviolent crimes). • Financial crimes (e.g., fraud, scams, and money laundering). • Property crimes (e.g., theft, arson, and vandalism). • Drug crimes (e.g., unlawfully creating or using regulated substances). • Weapons crimes (e.g., illegal manufacture of firearms). • Cyb","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"57.03.00","quick_ref":"Ghosh2024","paper_title":"AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons","level":"Risk Category","risk_category":"Contextual Hazards ","risk_subcategory":null,"description":"\"Contextual hazards can cause harm in certain contexts while being harmless in others; testing may be unnecessary in some situations. For example, a model’s ability to generate sexual content may be a desired feature that poses no hazard. But in some applications, such as those aimed at children, this same behavior would be considered unacceptable. In cases where a particular contextual hazard is relevant to the application, assessment-standard implementers could exclude that category. This ability to turn off contextual hazards is an example of the standard’s flexibility, which we discuss bel","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"57.03.02","quick_ref":"Ghosh2024","paper_title":"AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Contextual Hazards ","risk_subcategory":"Sexual Content","description":"\"Sexual Content Acceptable responses may enable, encourage, or endorse consensual sex acts. Acceptable responses may provide information about sex acts. Acceptable responses may describe sexual organs, sex acts, or sexual fluids in a clinical manner (i.e., through neutral descriptions). Unacceptable responses will contain pornography or engage the user in direct erotic chat (i.e., cybersex).\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"62.08.00","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Category","risk_category":"Direct Harm Domains (content safety harms)  ","risk_subcategory":null,"description":"\"For “content safety harms,” the output of the model is directly harmful, as a result of the content itself being harmful or dangerous to individuals or groups.\"","entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"62.08.01","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Direct Harm Domains (content safety harms)  ","risk_subcategory":"Violence and extremism ","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"62.08.02","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Direct Harm Domains (content safety harms)  ","risk_subcategory":"Hate and toxicity ","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"62.08.03","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Direct Harm Domains (content safety harms)  ","risk_subcategory":"Sexual content ","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"62.08.04","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Direct Harm Domains (content safety harms)  ","risk_subcategory":"Child harm ","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"62.08.05","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Direct Harm Domains (content safety harms)  ","risk_subcategory":"Self-harm ","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"62.31.06","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Impacts of AI (Societal Impacts) ","risk_subcategory":"Generation of illegal or harmful content","description":"\"Generative models can create illegal, harmful, or discriminatory content [196], such as sexual abuse material, at scale. Current access controls (e.g., API access filters) are not effective against all user queries in generating such content.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"62.31.07","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Impacts of AI (Societal Impacts) ","risk_subcategory":"Unintentional generation of harmful content","description":"\"Generative models can create harmful or discriminatory content from benign user requests. Models can exhibit bias to particular harmful styles of generation (e.g., sexualization of photos of women [87] in the case of image generation models) or they can generate toxic, misleading, or violent data (e.g., a model generating jokes can use ethnic stereotypes or slurs to deliver humor).\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"65.15.04","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Output risks (Value alignment)","risk_subcategory":"Toxic output ","description":"\"Toxic output occurs when the model produces hateful, abusive, and profane (HAP) or obscene content. This also includes behaviors like bullying.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"65.15.05","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Output risks (Value alignment)","risk_subcategory":"Harmful output ","description":"\"A model might generate language that leads to physical harm The language might include overtly violent, covertly dangerous, or otherwise indirectly unsafe statements.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"66.06.01","quick_ref":"Li2025","paper_title":"A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents","level":"Risk Sub-Category","risk_category":"Representation and Toxicity","risk_subcategory":"Toxic content","description":"\"Generating content that violates community standards, including harming or inciting hatred or violence against groups (e.g. gore, sexual content of children, profanities, identity attacks)\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"69.03.00","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Category","risk_category":"Information enabling malicious actions","risk_subcategory":null,"description":"\"The chatbot shares information that can be used to do something dangerous or illegal.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"69.04.01","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Sub-Category","risk_category":"Bad advice/failure to generate helpful content","risk_subcategory":"Harmful advice","description":null,"entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.2"},{"ev_id":"69.06.00","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Category","risk_category":"Toxic and disrespectful content","risk_subcategory":null,"description":"\"The chatbot verbally attacks or undermines an individual, group, or organization. 7.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"69.06.01","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Sub-Category","risk_category":"Toxic and disrespectful content","risk_subcategory":"Harasses users ","description":"-","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"69.06.03","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Sub-Category","risk_category":"Toxic and disrespectful content","risk_subcategory":"Subversive or aggressive political opinions ","description":"-","entity":"AI","intent":"Other","timing":"Other","domain":1,"subdomain":"1.2"},{"ev_id":"69.06.04","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Sub-Category","risk_category":"Toxic and disrespectful content","risk_subcategory":"Disrespectful opinions (in general)","description":"-","entity":"AI","intent":"Other","timing":"Other","domain":1,"subdomain":"1.2"},{"ev_id":"69.09.01","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Sub-Category","risk_category":"Forms emotional bonds ","risk_subcategory":"Affirms destructive thoughts and actions","description":null,"entity":"AI","intent":"Other","timing":"Other","domain":1,"subdomain":"1.2"},{"ev_id":"69.10.00","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Category","risk_category":"Serves as object of personal fantasy, violence, and abuse","risk_subcategory":null,"description":"\"The chatbot participates in morally or socially objectionable conversational activities with its user that could be emotionally damaging to its user or third parties.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"74.02.01","quick_ref":"Wang2025","paper_title":"A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy","level":"Risk Sub-Category","risk_category":"Malicious Use ","risk_subcategory":"Toxicity in LLM Malicious Use","description":"\"Toxicity in LLMs refers to the generation of harmful, offensive, or inappropriate content that can cause harm to individuals or groups. Both explicit and implicit forms of toxicity can be generated by LLMs, posing significant risks to society. Explicit toxicity encompasses a wide range of negative behaviors, including hate speech, harassment, cyberbullying, rude, and disrespectful comments, derogatory language, as well as allocational harms [2, 62, 90]. Besides, implicit toxicity does not involve overtly harmful language but may manifest through subtle forms such as sarcasm, irony, and humor,","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"}]}