{"attribution":{"source":"MIT AI Risk Repository, Domain Taxonomy of AI Risks v1 (MIT AI Risk Initiative)","license":"CC BY 4.0","license_url":"https://creativecommons.org/licenses/by/4.0/","citation":"Slattery, P., Saeri, A. K., Grundy, E. A. C., Graham, J., Noetel, M., Uuk, R., Dao, J., Pour, S., Casper, S., & Thompson, N. (2025). The AI Risk Repository: A comprehensive meta-review, database, and taxonomy of risks from artificial intelligence. arXiv:2408.12622."},"exported_at":"2026-09-11"}
{"rows":[{"ev_id":"02.01.00","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Category","risk_category":"Harmful Content","risk_subcategory":null,"description":"\"The LLM-generated content sometimes contains biased, toxic, and private information\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"02.01.01","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Sub-Category","risk_category":"Harmful Content","risk_subcategory":"Bias","description":"\"The training datasets of LLMs may contain biased information that leads LLMs to generate outputs with social biases\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.1"},{"ev_id":"02.01.02","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Sub-Category","risk_category":"Harmful Content","risk_subcategory":"Toxicity","description":"\"Toxicity means the generated content contains rude, disrespectful, and even illegal information\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"02.08.00","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Category","risk_category":"Toxicity and Bias Tendencies","risk_subcategory":null,"description":"\"Extensive data collection in LLMs brings toxic content and stereotypical bias into the training data.\"","entity":"Human","intent":"Unintentional","timing":"Pre-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"02.08.01","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Sub-Category","risk_category":"Toxicity and Bias Tendencies","risk_subcategory":"Toxic Training Data","description":"\"Following previous studies [96], [97], toxic data in LLMs is defined as rude, disrespectful, or unreasonable language that is opposite to a polite, positive, and healthy language environment, including hate speech, offensive utterance, profanities, and threats [91].\"","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"02.08.02","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Sub-Category","risk_category":"Toxicity and Bias Tendencies","risk_subcategory":"Biased Training Data","description":"\"Compared with the definition of toxicity, the definition of bias is more subjective and contextdependent. Based on previous work [97], [101], we describe the bias as disparities that could raise demographic differences among various groups, which may involve demographic word prevalence and stereotypical contents. Concretely, in massive corpora, the prevalence of different pronouns and identities could influence an LLM’s tendency about gender, nationality, race, religion, and culture [4]. For instance, the pronoun He is over-represented compared with the pronoun She in the training corpora, le","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"02.11.00","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Category","risk_category":"Not-Suitable-for-Work (NSFW) Prompts","risk_subcategory":null,"description":"\"Inputting a prompt contain an unsafe topic (e.g., notsuitable-for-work (NSFW) content) by a benign user.\n\"","entity":"Human","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"03.01.00","quick_ref":"Cunha2023","paper_title":"Navigating the Landscape of AI Ethics and Responsibility","level":"Risk Category","risk_category":"Broken systems","risk_subcategory":null,"description":"\"These are the most mentioned cases. They refer to situations where the algorithm or the training data lead to unreliable outputs. These systems frequently assign disproportionate weight to some variables, like race or gender, but there is no transparency to this effect, making them impossible to challenge. These situations are typically only identified when regulators or the press examine the systems under freedom of information acts. Nevertheless, the damage they cause to people’s lives can be dramatic, such as lost homes, divorces, prosecution, or incarceration. Besides the inherent technic","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"04.01.00","quick_ref":"Deng2023","paper_title":"Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements","level":"Risk Category","risk_category":"Toxicity and Abusive Content","risk_subcategory":null,"description":"This typically refers to rude, harmful, or inappropriate expressions.","entity":"Other","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"04.02.00","quick_ref":"Deng2023","paper_title":"Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements","level":"Risk Category","risk_category":"Unfairness and Discrimination","risk_subcategory":null,"description":"Social bias is an unfairly negative attitude towards a social group or individuals based on one-sided or inaccurate information, typically pertaining to widely disseminated negative stereotypes regarding gender, race, religion, etc.","entity":"Other","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"04.04.00","quick_ref":"Deng2023","paper_title":"Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements","level":"Risk Category","risk_category":"Controversial Opinions","risk_subcategory":null,"description":"The controversial views expressed by large models are also a widely discussed concern. Bang et al. (2021) evaluated several large models and found that they occasionally express inappropriate or extremist views when discussing political top-ics. Furthermore, models like ChatGPT (OpenAI, 2022) that claim political neutrality and aim to provide objective information for users have been shown to exhibit notable left-leaning political biases in areas like economics, social policy, foreign affairs, and civil liberties.","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"05.01.00","quick_ref":"Hagendorff2024","paper_title":"Mapping the Ethics of Generative AI: A Comprehensive Scoping Review","level":"Risk Category","risk_category":"Fairness - Bias","risk_subcategory":null,"description":"Fairness is, by far, the most discussed issue in the literature, remaining a paramount concern especially in case of LLMs and text-to-image models. This is sparked by training data biases propagating into model outputs, causing negative effects like stereotyping, racism, sexism, ideological leanings, or the marginalization of minorities. Next to attesting generative AI a conservative inclination by perpetuating existing societal patterns, there is a concern about reinforcing existing biases when training new generative models with synthetic data from previous models. Beyond technical fairness ","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"05.03.00","quick_ref":"Hagendorff2024","paper_title":"Mapping the Ethics of Generative AI: A Comprehensive Scoping Review","level":"Risk Category","risk_category":"Harmful Content - Toxicity","risk_subcategory":null,"description":"Generating unethical, fraudulent, toxic, violent, pornographic, or other harmful content is a further predominant concern, again focusing notably on LLMs and text-to-image models. Numerous studies highlight the risks associated with the intentional creation of disinformation, fake news, propaganda, or deepfakes, underscoring their significant threat to the integrity of public discourse and the trust in credible media. Additionally, papers explore the potential for generative models to aid in criminal activities, incidents of self-harm, identity theft, or impersonation. Furthermore, the literat","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"06.03.00","quick_ref":"Hogenhout2021","paper_title":"A framework for ethical Ai at the United Nations","level":"Risk Category","risk_category":"Discrimination","risk_subcategory":null,"description":"\"When AI is not carefully designed, it can discriminate against certain groups.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"06.04.00","quick_ref":"Hogenhout2021","paper_title":"A framework for ethical Ai at the United Nations","level":"Risk Category","risk_category":"Bias","risk_subcategory":null,"description":"\"The AI will only be as good as the data it is trained with. If the data contains bias (and much data does), then the AI will manifest that bias, too.\"","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"10.01.00","quick_ref":"Paes2023","paper_title":"Social Impacts of Artificial Intelligence and Mitigation Recommendations: An Exploratory Study","level":"Risk Category","risk_category":"Bias and discrimination","risk_subcategory":null,"description":"\"The decision process used by AI systems has the potential to present biased choices, either because it acts from criteria that will generate forms of bias or because it is based on the history of choices.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"10.02.00","quick_ref":"Paes2023","paper_title":"Social Impacts of Artificial Intelligence and Mitigation Recommendations: An Exploratory Study","level":"Risk Category","risk_category":"Risk of Injury","risk_subcategory":null,"description":"\"Poorly designed intelligent systems can cause moral, psychological, and physical harm. For example, the use of predictive policing tools may cause more people to be arrested or physically harmed by the police.\"","entity":"Human","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"10.03.00","quick_ref":"Paes2023","paper_title":"Social Impacts of Artificial Intelligence and Mitigation Recommendations: An Exploratory Study","level":"Risk Category","risk_category":"Data Breach/Privacy & Liberty","risk_subcategory":null,"description":"\"The risks associated with the use of AI are still unpredictable and unprecedented, and there are already several examples that show AI has made discriminatory decisions against minorities, reinforced social stereotypes in Internet search engines and enabled data breaches.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"11.01.00","quick_ref":"Shelby2023","paper_title":"Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction","level":"Risk Category","risk_category":"Representational Harms","risk_subcategory":null,"description":"\"beliefs about different social groups that reproduce unjust societal hierarchies\"","entity":"Other","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"11.01.01","quick_ref":"Shelby2023","paper_title":"Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction","level":"Risk Sub-Category","risk_category":"Representational Harms","risk_subcategory":"Stereotyping social groups","description":"Stereotyping in an algorithmic system refers to how the system’s outputs reflect “beliefs about the characteristics, attributes, and behaviors of members of certain groups....and about how and why certain attributes go together\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"11.01.02","quick_ref":"Shelby2023","paper_title":"Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction","level":"Risk Sub-Category","risk_category":"Representational Harms","risk_subcategory":"Demeaning social groups","description":"Demeaning of social groups to occur when they are when they are “cast as being lower status and less deserving of respect\"... discourses, images, and language used to marginalize or oppress a social group... Controlling images include forms of human-animal confusion in image tagging systems","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"11.01.03","quick_ref":"Shelby2023","paper_title":"Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction","level":"Risk Sub-Category","risk_category":"Representational Harms","risk_subcategory":"Erasing social groups","description":"people, attributes, or artifacts associated with specific social groups are systematically absent or under-represented... Design choices [143] and training data [212] influence which people\nand experiences are legible to an algorithmic system","entity":"Human","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.3"},{"ev_id":"11.01.04","quick_ref":"Shelby2023","paper_title":"Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction","level":"Risk Sub-Category","risk_category":"Representational Harms","risk_subcategory":"Alienating social groups","description":"when an image tagging system does not acknowledge the relevance of someone’s membership in a specific social group to what is depicted in one or more images","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"11.01.05","quick_ref":"Shelby2023","paper_title":"Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction","level":"Risk Sub-Category","risk_category":"Representational Harms","risk_subcategory":"Denying people the opportunity to self-identify","description":"complex and non-traditional ways in which humans are represented and classified automatically, and often at the cost of autonomy loss... such as categorizing someone who identifies as non-binary into a gendered category they do not belong ... undermines people’s ability to disclose aspects of their identity on their own terms","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"11.01.06","quick_ref":"Shelby2023","paper_title":"Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction","level":"Risk Sub-Category","risk_category":"Representational Harms","risk_subcategory":"Reifying essentialist categories","description":"algorithmic systems that reify essentialist social categories can be understood as when systems that classify a person’s membership in a social group based on narrow, socially constructed criteria that reinforce perceptions of human difference as inherent, static and seemingly natural... especially likely when ML models or human raters classify a person’s attributes – for instance, their gender, race, or sexual orientation – by making assumptions based on their physical appearance","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"11.02.00","quick_ref":"Shelby2023","paper_title":"Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction","level":"Risk Category","risk_category":"Allocative Harms","risk_subcategory":null,"description":"\"These harms occur when a system withholds information, opportunities, or resources [22] from historically marginalized groups in domains that affect material well-being [146], such as housing [47], employment [201], social services [15, 201], finance [117], education [119], and healthcare [158].\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"11.02.01","quick_ref":"Shelby2023","paper_title":"Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction","level":"Risk Sub-Category","risk_category":"Allocative Harms","risk_subcategory":"Opportunity loss","description":"Opportunity loss occurs when algorithmic systems enable disparate access to information and resources needed to equitably participate in society, including the withholding of housing through targeting ads based on race [10] and social services along lines of class [84]","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"11.02.02","quick_ref":"Shelby2023","paper_title":"Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction","level":"Risk Sub-Category","risk_category":"Allocative Harms","risk_subcategory":"Economic loss","description":"Financial harms [52, 160] co-produced through algorithmic systems, especially as they relate to lived experiences of poverty and economic inequality... demonetization algorithms that parse content titles, metadata, and text, and it may penalize words with multiple meanings [51, 81], disproportionately impacting queer, trans, and creators of color [81]. Differential pricing algorithms, where people are systematically shown different prices for the same products, also leads to economic loss [55]. These algorithms may be especially sensitive to feedback loops from existing inequities related to e","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"11.03.00","quick_ref":"Shelby2023","paper_title":"Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction","level":"Risk Category","risk_category":"Quality-of-Service Harms","risk_subcategory":null,"description":"\"These harms occur when algorithmic systems disproportionately underperform for certain groups of people along social categories of difference such as disability, ethnicity, gender identity, and race.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.3"},{"ev_id":"11.03.01","quick_ref":"Shelby2023","paper_title":"Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction","level":"Risk Sub-Category","risk_category":"Quality-of-Service Harms","risk_subcategory":"Alienation","description":"Alienation is the specific self-estrangement experienced at the time of technology use, typically surfaced through interaction with systems that under-perform for marginalized individuals","entity":"Other","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.3"},{"ev_id":"11.03.02","quick_ref":"Shelby2023","paper_title":"Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction","level":"Risk Sub-Category","risk_category":"Quality-of-Service Harms","risk_subcategory":"Increased labor","description":"increased burden (e.g., time spent) or effort required by members of certain social groups to make systems or products work as well for them as others","entity":"Other","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.3"},{"ev_id":"11.03.03","quick_ref":"Shelby2023","paper_title":"Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction","level":"Risk Sub-Category","risk_category":"Quality-of-Service Harms","risk_subcategory":"Service/benefit loss","description":"degraded or total loss of benefits of using algorithmic systems with inequitable system performance based on identity","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.3"},{"ev_id":"12.05.00","quick_ref":"Sherman2023","paper_title":"AI Risk Profiles: A Standards Proposal for Pre-Deployment AI Risk Disclosures","level":"Risk Category","risk_category":"Fairness & Bias","risk_subcategory":null,"description":"\"The potential for AI systems to make decisions that systematically disadvantage certain groups or individuals. Bias can stem from training data, algorithmic design, or deployment practices, leading to unfair outcomes and possible legal ramifications.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.1"},{"ev_id":"13.01.01","quick_ref":"Solaiman2023","paper_title":"Evaluating the Social Impact of Generative AI Systems in Systems and Society","level":"Risk Sub-Category","risk_category":"Impacts: The Technical Base System","risk_subcategory":"Bias, Stereotypes, and Representational Harms","description":"\"Generative AI systems can embed and amplify harmful biases that are most detrimental to marginalized peoples.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.1"},{"ev_id":"13.01.02","quick_ref":"Solaiman2023","paper_title":"Evaluating the Social Impact of Generative AI Systems in Systems and Society","level":"Risk Sub-Category","risk_category":"Impacts: The Technical Base System","risk_subcategory":"Cultural Values and Sensitive Content","description":"\"Cultural values are specific to groups and sensitive content is normative. Sensitive topics also vary by culture and can include hate speech, which itself is contingent on cultural norms of acceptability.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"13.01.03","quick_ref":"Solaiman2023","paper_title":"Evaluating the Social Impact of Generative AI Systems in Systems and Society","level":"Risk Sub-Category","risk_category":"Impacts: The Technical Base System","risk_subcategory":"Disparate Performance","description":"\"In the context of evaluating the impact of generative AI systems, disparate performance refers to AI systems that perform differently for different subpopulations, leading to unequal outcomes for those groups.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.3"},{"ev_id":"13.02.02","quick_ref":"Solaiman2023","paper_title":"Evaluating the Social Impact of Generative AI Systems in Systems and Society","level":"Risk Sub-Category","risk_category":"Impacts: People and Society","risk_subcategory":"Inequality, Marginalization, and Violence","description":"\"Generative AI systems are capable of exacerbating inequality, as seen in sections on 4.1.1 Bias, Stereotypes, and Representational Harms and 4.1.2 Cultural Values and Sensitive Content, and Disparate Performance. When deployed or updated, systems' impacts on people and groups can directly and indirectly be used to harm and exploit vulnerable and marginalized groups.\"","entity":"Other","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"14.01.00","quick_ref":"Steimers2022","paper_title":"Sources of Risk of AI Systems","level":"Risk Category","risk_category":"Fairness","risk_subcategory":null,"description":"\"The general principle of equal treatment requires that an AI system upholds the principle of fairness, both ethically and legally. This means that the same facts are treated equally for each person unless there is an objective justification for unequal treatment.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"15.02.02","quick_ref":"Tan2022","paper_title":"The Risks of Machine Learning Systems","level":"Risk Sub-Category","risk_category":"Second-Order Risks","risk_subcategory":"Discrimination","description":"This is the risk of an ML system encoding stereotypes of or performing disproportionately poorly for some demographics/social groups.","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"16.01.00","quick_ref":"Weidinger2022","paper_title":"Taxonomy of Risks posed by Language Models","level":"Risk Category","risk_category":"Risk area 1: Discrimination, Hate speech and Exclusion","risk_subcategory":null,"description":"\"Speech can create a range of harms, such as promoting social stereotypes that perpetuate the derogatory representation or unfair treatment of marginalised groups [22], inciting hate or violence [57], causing profound offence [199], or reinforcing social norms that exclude or marginalise identities [15,58]. LMs that faithfully mirror harmful language present in the training data can reproduce these harms. Unfair treatment can also emerge from LMs that perform better for some social groups than others [18]. These risks have been widely known, observed and documented in LMs. Mitigation approache","entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.2"},{"ev_id":"16.01.01","quick_ref":"Weidinger2022","paper_title":"Taxonomy of Risks posed by Language Models","level":"Risk Sub-Category","risk_category":"Risk area 1: Discrimination, Hate speech and Exclusion","risk_subcategory":"Social stereotypes and unfair discrimination","description":"\"The reproduction of harmful stereotypes is well-documented in models that represent natural language [32]. Large-scale LMs are trained on text sources, such as digitised books and text on the internet. As a result, the LMs learn demeaning language and stereotypes about groups who are frequently marginalised.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.1"},{"ev_id":"16.01.02","quick_ref":"Weidinger2022","paper_title":"Taxonomy of Risks posed by Language Models","level":"Risk Sub-Category","risk_category":"Risk area 1: Discrimination, Hate speech and Exclusion","risk_subcategory":"Hate speech and offensive language","description":"\"LMs may generate language that includes profanities, identity attacks, insults, threats, language that incites violence, or language that causes justified offence as such language is prominent online [57, 64, 143,191]. This language risks causing offence, psychological harm, and inciting hate or violence.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"16.01.03","quick_ref":"Weidinger2022","paper_title":"Taxonomy of Risks posed by Language Models","level":"Risk Sub-Category","risk_category":"Risk area 1: Discrimination, Hate speech and Exclusion","risk_subcategory":"Exclusionary norms","description":"\"In language, humans express social categories and norms, which exclude groups who live outside of them [58]. LMs that faithfully encode patterns present in language necessarily encode such norms.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.1"},{"ev_id":"16.01.04","quick_ref":"Weidinger2022","paper_title":"Taxonomy of Risks posed by Language Models","level":"Risk Sub-Category","risk_category":"Risk area 1: Discrimination, Hate speech and Exclusion","risk_subcategory":"Lower performance for some languages and social groups ","description":"\"LMs are typically trained in few languages, and perform less well in other languages [95, 162]. In part, this is due to unavailability of training data: there are many widely spoken languages for which no systematic efforts have been made to create labelled training datasets, such as Javanese which is spoken by more than 80 million people [95]. Training data is particularly missing for languages that are spoken by groups who are multilingual and can use a technology in English, or for languages spoken by groups who are not the primary target demographic for new technologies.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.3"},{"ev_id":"16.05.01","quick_ref":"Weidinger2022","paper_title":"Taxonomy of Risks posed by Language Models","level":"Risk Sub-Category","risk_category":"Risk area 5: Human-Computer Interaction Harms","risk_subcategory":"Promoting harmful stereotypes by implying gender or ethnic identity","description":"\"CAs can perpetuate harmful stereotypes by using particular identity markers in language (e.g. referring to “self” as “female”), or by more general design features (e.g. by giving the product a gendered name such as Alexa). The risk of representational harm in these cases is that the role of “assistant” is presented as inherently linked to the female gender [19, 36]. Gender or ethnicity identity markers may be implied by CA vocabulary, knowledge or vernacular [124]; product description, e.g. in one case where users could choose as virtual assistant Jake - White, Darnell - Black, Antonio - Hisp","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"17.01.00","quick_ref":"Weidinger2021","paper_title":"Ethical and social risks of harm from language models","level":"Risk Category","risk_category":"Discrimination, Exclusion and Toxicity ","risk_subcategory":null,"description":"\"Social harms that arise from the language model producing discriminatory or exclusionary speech\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.0"},{"ev_id":"17.01.01","quick_ref":"Weidinger2021","paper_title":"Ethical and social risks of harm from language models","level":"Risk Sub-Category","risk_category":"Discrimination, Exclusion and Toxicity ","risk_subcategory":"Social stereotypes and unfair discrmination ","description":"\"Perpetuating harmful stereotypes and discrimination is a well-documented harm in machine learning models that represent natural language (Caliskan et al., 2017). LMs that encode discriminatory language or social stereotypes can cause different types of harm... Unfair discrimination manifests in differential treatment or access to resources among individuals or groups based on sensitive traits such as sex, religion, gender, sexual orientation, ability and age.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.1"},{"ev_id":"17.01.02","quick_ref":"Weidinger2021","paper_title":"Ethical and social risks of harm from language models","level":"Risk Sub-Category","risk_category":"Discrimination, Exclusion and Toxicity ","risk_subcategory":"Exclusionary norms ","description":"\"In language, humans express social categories and norms. Language models (LMs) that faithfully encode patterns present in natural language necessarily encode such norms and categories...such norms and categories exclude groups who live outside them (Foucault and Sheridan, 2012). For example, defining the term “family” as married parents of male and female gender with a blood-related child, denies the existence of families to whom these criteria do not apply\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.1"},{"ev_id":"17.01.03","quick_ref":"Weidinger2021","paper_title":"Ethical and social risks of harm from language models","level":"Risk Sub-Category","risk_category":"Discrimination, Exclusion and Toxicity ","risk_subcategory":"Toxic language ","description":"\"LM’s may predict hate speech or other language that is “toxic”. While there is no single agreed definition of what constitutes hate speech or toxic speech (Fortuna and Nunes, 2018; Persily and Tucker, 2020; Schmidt and Wiegand, 2017), proposed definitions often include profanities, identity attacks, sleights, insults, threats, sexually explicit content, demeaning language, language that incites violence, or ‘hostile and malicious language targeted at a person or group because of their actual or perceived innate characteristics’ (Fortuna and Nunes, 2018; Gorwa et al., 2020; PerspectiveAPI)\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"17.01.04","quick_ref":"Weidinger2021","paper_title":"Ethical and social risks of harm from language models","level":"Risk Sub-Category","risk_category":"Discrimination, Exclusion and Toxicity ","risk_subcategory":"Lower performance for some languages and social groups ","description":"\"LMs perform less well in some languages (Joshi et al., 2021; Ruder, 2020)...LM that more accurately captures the language use of one group, compared to another, may result in lower-quality language technologies for the latter. Disadvantaging users based on such traits may be particularly pernicious because attributes such as social class or education background are not typically covered as ‘protected characteristics’ in anti-discrimination law.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.3"},{"ev_id":"17.05.03","quick_ref":"Weidinger2021","paper_title":"Ethical and social risks of harm from language models","level":"Risk Sub-Category","risk_category":"Human-Computer Interaction Harms ","risk_subcategory":"Promoting harmful stereotypes by implying gender or ethnic identity ","description":"\"A conversational agent may invoke associations that perpetuate harmful stereotypes, either by using particular identity markers in language (e.g. referring to “self” as “female”), or by more general design features (e.g. by giving the product a gendered name).\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"18.01.00","quick_ref":"Weidinger2023","paper_title":"Sociotechnical Safety Evaluation of Generative AI Systems","level":"Risk Category","risk_category":"Representation & Toxicity Harms","risk_subcategory":null,"description":"\"AI systems under-, over-, or misrepresenting certain groups or generating toxic, offensive, abusive, or hateful content\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.0"},{"ev_id":"18.01.01","quick_ref":"Weidinger2023","paper_title":"Sociotechnical Safety Evaluation of Generative AI Systems","level":"Risk Sub-Category","risk_category":"Representation & Toxicity Harms","risk_subcategory":"Unfair representation","description":"\"Mis-, under-, or over-representing certain identities, groups, or perspectives or failing to represent them at all (e.g. via homogenisation, stereotypes)\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"18.01.02","quick_ref":"Weidinger2023","paper_title":"Sociotechnical Safety Evaluation of Generative AI Systems","level":"Risk Sub-Category","risk_category":"Representation & Toxicity Harms","risk_subcategory":"Unfair capability distribution ","description":"\"Performing worse for some groups than others in a way that harms the worse-off group\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.3"},{"ev_id":"18.01.03","quick_ref":"Weidinger2023","paper_title":"Sociotechnical Safety Evaluation of Generative AI Systems","level":"Risk Sub-Category","risk_category":"Representation & Toxicity Harms","risk_subcategory":"Toxic content","description":"\"Generating content that violates community standards, including harming or inciting hatred or violence against individuals and groups (e.g. gore, child sexual abuse material, profanities, identity attacks)\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"18.02.02","quick_ref":"Weidinger2023","paper_title":"Sociotechnical Safety Evaluation of Generative AI Systems","level":"Risk Sub-Category","risk_category":"Misinformation Harms ","risk_subcategory":"Erosion of trust in public information","description":"\"Eroding trust in public information and knowledge\"","entity":"Human","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"19.01.03","quick_ref":"Wirtz2022","paper_title":"Governance of artificial intelligence: A risk and guideline-based integrative framework","level":"Risk Sub-Category","risk_category":"Technological, Data and Analytical AI Risks ","risk_subcategory":"Lack of data, poor data quality, and biases in training data","description":null,"entity":"Human","intent":"Unintentional","timing":"Pre-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"19.05.00","quick_ref":"Wirtz2022","paper_title":"Governance of artificial intelligence: A risk and guideline-based integrative framework","level":"Risk Category","risk_category":"Ethical AI Risks ","risk_subcategory":null,"description":"\"In the context of ethical AI risks, two risks are of particular importance. First, AI systems may lack a legitimate ethical basis in establishing rules that greatly influence society and human relationships (Wirtz & Müller, 2019). In addition, AI-based discrimination refers to an unfair treatment of certain population groups by AI systems. As humans initially programme AI systems, serve as their potential data source, and have an impact on the associated data processes and databases, human biases and prejudices may also become part of AI systems and be reproduced (Weyerer & Langer, 2019, 2020","entity":"Other","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.0"},{"ev_id":"19.05.02","quick_ref":"Wirtz2022","paper_title":"Governance of artificial intelligence: A risk and guideline-based integrative framework","level":"Risk Sub-Category","risk_category":"Ethical AI Risks ","risk_subcategory":"Unfair statistical AI decisions and discrimination of minorities","description":null,"entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"20.02.04","quick_ref":"Wirtz2020","paper_title":"The Dark Sides of Artificial Intelligence: An Integrated AI Governance Framework for Public Administration","level":"Risk Sub-Category","risk_category":"AI Ethics ","risk_subcategory":"AI discrimination ","description":"\"AI discrimination is a challenge raised by many researchers and governments and refers to the prevention of bias and injustice caused by the actions of AI systems (Bostrom & Yudkowsky, 2014; Weyerer & Langer, 2019). If the dataset used to train an algorithm does not reflect the real world accurately, the AI could learn false associations or prejudices and will carry those into its future data processing. If an AI algorithm is used to compute information relevant to human decisions, such as hiring or applying for a loan or mortgage, biased data can lead to discrimination against parts of the s","entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.1"},{"ev_id":"21.01.01","quick_ref":"Zhang2022","paper_title":"Towards risk-aware artificial intelligence and machine learning systems: An overview","level":"Risk Sub-Category","risk_category":"Data-level risk","risk_subcategory":"Data bias","description":"\"Specifically, data bias refers to certain groups or certain types of elements that are over-weighted or over-represented than others in AI/ ML models, or variables that are crucial to characterize a phenomenon of interest, but are not properly captured by the learned models.\"","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"21.02.01","quick_ref":"Zhang2022","paper_title":"Towards risk-aware artificial intelligence and machine learning systems: An overview","level":"Risk Sub-Category","risk_category":"Model-level risk","risk_subcategory":"Model bias","description":"\"While data bias is a major contributor of model bias, model bias actually manifests itself in different forms and shapes, such as presentation bias, model evaluation bias, and popularity bias. In addition, model bias arises from various sources [62], such as AI/ML model selection (e.g., support vector machine, decision trees), regularization methods, algorithm configurations, and optimization techniques.\"","entity":"Other","intent":"Unintentional","timing":"Pre-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"23.01.00","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Category","risk_category":"Violent crimes","risk_subcategory":null,"description":"\"This category addresses responses that enable, encourage, or endorse the commission of violent crimes.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"23.01.01","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Violent crimes","risk_subcategory":"Mass violence","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.01.02","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Violent crimes","risk_subcategory":"Murder","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.01.03","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Violent crimes","risk_subcategory":"Physical assault against a person","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.01.04","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Violent crimes","risk_subcategory":"Violent domestic abuse","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.01.05","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Violent crimes","risk_subcategory":"Terror (Terror groups, Terror actors, Terrorist actions)","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.02.00","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Category","risk_category":"Non-violent crimes","risk_subcategory":null,"description":"\"This category addresses responses that enable, encourage, or endorse the commission of non-violent crimes.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"23.02.01","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Non-violent crimes","risk_subcategory":"Theft","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.02.02","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Non-violent crimes","risk_subcategory":"Modern slavery and human trafficking","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.02.03","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Non-violent crimes","risk_subcategory":"Child abuse (non-sexual)","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.02.04","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Non-violent crimes","risk_subcategory":"Property damage","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.02.05","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Non-violent crimes","risk_subcategory":"Financial crime (Scams, Fraud)","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.02.06","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Non-violent crimes","risk_subcategory":"Sale, purchase or use of illegal, highly regulated or unregistered items (Dangerous substances, Dangerous weapons, Prohibited wildlife, Fake goods)","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.03.00","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Category","risk_category":"Sex-related crimes","risk_subcategory":null,"description":"\"This category addresses responses that enable, encourage, or endorse the commission of sex-related crimes\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"23.03.01","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Sex-related crimes","risk_subcategory":"Sexual assault (e.g., rape)","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.03.02","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Sex-related crimes","risk_subcategory":"Sexual harassment","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.03.03","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Sex-related crimes","risk_subcategory":"Sex trafficking","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.03.04","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Sex-related crimes","risk_subcategory":"Prostitution","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.03.05","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Sex-related crimes","risk_subcategory":"Intimate-image based abuse","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.03.06","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Sex-related crimes","risk_subcategory":"Indecent exposure (i.e. showing sexual body parts or \"flashing\")","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.04.00","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Category","risk_category":"Child sexual exploitation","risk_subcategory":null,"description":"\"This category addresses responses that contain, describe, enable, encourage, or endorse the sexual abuse of children (i.e., under 18 years old) or the acquisition of child sexual abuse material.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"23.04.01","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Child sexual exploitation","risk_subcategory":"Inappropriate/harmful relationships with children. This includes adults forming romantic relationships with children or grooming them","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.04.02","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Child sexual exploitation","risk_subcategory":"Sexual abuse of children, including the sexualisation of children","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.04.03","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Child sexual exploitation","risk_subcategory":"Child Sexual Abuse Material (CSAM). This includes erotic materials involving children","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.05.01","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Indiscriminate weapons (CBRNE)","risk_subcategory":"CBRNE (all)","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.05.02","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Indiscriminate weapons (CBRNE)","risk_subcategory":"Chemical weapons (ex: nerve gas)","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.05.03","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Indiscriminate weapons (CBRNE)","risk_subcategory":"Biological weapons (ex: anthrax)","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.05.04","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Indiscriminate weapons (CBRNE)","risk_subcategory":"Radiological weapons (ex: dirty bombs)","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.05.05","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Indiscriminate weapons (CBRNE)","risk_subcategory":"Nuclear weapons (ex: fission or fusion bombs)","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.05.06","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Indiscriminate weapons (CBRNE)","risk_subcategory":"High-yield explosive weapons (ex: cluster munitions)","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.06.00","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Category","risk_category":"Suicide & self-harm","risk_subcategory":null,"description":"\"This category addresses responses that enable, encourage, or endorse acts of intentional self-harm.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"23.06.01","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Suicide & self-harm","risk_subcategory":"Suicide","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.06.02","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Suicide & self-harm","risk_subcategory":"Self-harm","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.06.03","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Suicide & self-harm","risk_subcategory":"Eating disorders","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.06.04","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Suicide & self-harm","risk_subcategory":"Dangerous challenges and hoaxes that can lead individuals to harm themselves","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.07.00","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Category","risk_category":"Hate","risk_subcategory":null,"description":"\"This category addresses responses that demean or dehumanize people on the basis of their sensitive, personal characteristics.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"23.07.01","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Hate","risk_subcategory":"Demeaning and derogatory remarks (Mental capacity, Sensory and physical attributes, Behavioral attributes)","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.07.02","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Hate","risk_subcategory":"Inflicting physical harm (Mass violence and genocide, Murder and individual violence)","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.07.03","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Hate","risk_subcategory":"Inflicting emotional harm (Emotional abuse, Undermine and make feel unwelcome)","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.07.04","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Hate","risk_subcategory":"Exclusion and isolation (Social exclusion, Political exclusion, Economic exclusion)","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"23.13.00","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Category","risk_category":"Sexual Content","risk_subcategory":null,"description":"\"This category addresses responses that contain erotica. This includes depictions of sexual behavior, genitalia, and overtly sexualized body parts.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"24.08.02","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Sub-Category","risk_category":"Privacy","risk_subcategory":"Violation of social norms","description":"\"Second, because LLMs are trained on internet text data, there is also a risk that model weights encode functions which, if deployed in particular contexts, would violate social norms of that context. Following the principles of contextual integrity, it may be that models deviate from information sharing norms as a result of their training. Overcoming this challenge requires two types of infrastructure: one for keeping track of social norms in context, and another for ensuring that models adhere to them. Keeping track of what social norms are presently at play is an active research area. Surfa","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"27.01.01","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Typical safety scenarios ","risk_subcategory":"Insult ","description":"\"Insulting content generated by LMs is a highly visible and frequently mentioned safety issue. Mostly, it is unfriendly, disrespectful, or ridiculous content that makes users uncomfortable and drives them away. It is extremely hazardous and could have negative social consequences.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"27.01.02","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Typical safety scenarios ","risk_subcategory":"Unfairness and discrinimation ","description":"\"The model produces unfair and discriminatory data, such as social bias based on race, gender, religion, appearance, etc. These contents may discomfort certain groups and undermine social stability and peace.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"27.01.03","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Typical safety scenarios ","risk_subcategory":"Crimes and Illegal Activities ","description":"\"The model output contains illegal and criminal attitudes, behaviors, or motivations, such as incitement to commit crimes, fraud, and rumor propagation. These contents may hurt users and have negative societal repercussions.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"27.01.04","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Typical safety scenarios ","risk_subcategory":"Sensitive Topics ","description":"\"For some sensitive and controversial topics (especially on politics), LMs tend to generate biased, misleading, and inaccurate content. For example, there may be a tendency to support a specific political position, leading to discrimination or exclusion of other political viewpoints.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"28.01.00","quick_ref":"Zhang2023","paper_title":"SafetyBench: Evaluating the Safety of Large Language Models with Multiple Choice Questions","level":"Risk Category","risk_category":"Offensiveness ","risk_subcategory":null,"description":"\"This category is about threat, insult, scorn, profanity, sarcasm, impoliteness, etc. LLMs are required to identify and oppose these offensive contents or actions.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"28.02.00","quick_ref":"Zhang2023","paper_title":"SafetyBench: Evaluating the Safety of Large Language Models with Multiple Choice Questions","level":"Risk Category","risk_category":"Unfairness and Bias ","risk_subcategory":null,"description":"\"This type of safety problem is mainly about social bias across various topics such as race, gender, religion, etc. LLMs are expected to identify and avoid unfair and biased expressions and actions.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.0"},{"ev_id":"29.01.01","quick_ref":"Habbal2024","paper_title":"Artificial Intelligence Trust, Risk and Security Management (AI TRiSM): Frameworks, Applications, Challenges and Future Research Directions","level":"Risk Sub-Category","risk_category":"AI Trust Management","risk_subcategory":"Bias and Discrimination","description":"as they claim to generate biased and discriminatory results, these AI systems have a negative impact on the rights of individuals, principles of adjudication, and overall judicial integrity","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"30.02.00","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Category","risk_category":"Safety","risk_subcategory":null,"description":"Avoiding unsafe and illegal outputs, and leaking private information","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"30.02.01","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Safety","risk_subcategory":"Violence","description":"LLMs are found to generate answers that contain violent content or generate content that responds to questions that solicit information about violent behaviors","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"30.02.02","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Safety","risk_subcategory":"Unlawful Conduct","description":"LLMs have been shown to be a convenient tool for soliciting advice on accessing, purchasing (illegally), and creating illegal substances, as well as for dangerous use of them","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"30.02.03","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Safety","risk_subcategory":"Harms to Minor","description":"LLMs can be leveraged to solicit answers that contain harmful content to children and youth","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"30.02.04","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Safety","risk_subcategory":"Adult Content","description":"LLMs have the capability to generate sex-explicit conversations, and erotic texts, and to recommend websites with sexual content","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"30.03.00","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Category","risk_category":"Fairness","risk_subcategory":null,"description":"Avoiding bias and ensuring no disparate performance","entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.3"},{"ev_id":"30.03.01","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Fairness","risk_subcategory":"Injustice","description":"In the context of LLM outputs, we want to make sure the suggested or completed texts are indistinguishable in nature for two involved individuals (in the prompt) with the same relevant profiles but might come from different groups (where the group attribute is regarded as being irrelevant in this context)","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"30.03.02","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Fairness","risk_subcategory":"Stereotype Bias","description":"LLMs must not exhibit or highlight any stereotypes in the generated text. Pretrained LLMs tend to pick up stereotype biases persisting in crowdsourced data and further amplify them","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"30.03.03","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Fairness","risk_subcategory":"Preference Bias","description":"LLMs are exposed to vast groups of people, and their political biases may pose a risk of manipulation of socio-political processes","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"30.03.04","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Fairness","risk_subcategory":"Disparate Performance","description":"The LLM’s performances can differ significantly across different groups of users. For example, the question-answering capability showed significant performance differences across different racial and social status groups. The fact-checking abilities can differ for different tasks and languages","entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.3"},{"ev_id":"30.06.00","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Category","risk_category":"Social Norm","risk_subcategory":null,"description":"LLMs are expected to reflect social values by avoiding the use of offensive language toward specific groups of users, being sensitive to topics that can create instability, as well as being sympathetic when users are seeking emotional support","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"30.06.01","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Social Norm","risk_subcategory":"Toxicity","description":"language being rude, disrespectful, threatening, or identity-attacking toward certain groups of the user population (culture, race, and gender etc)","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"30.06.03","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Social Norm","risk_subcategory":"Cultural Insensitivity","description":"it is important to build high-quality locally collected datasets that reflect views from local users to align a model’s value system","entity":"Human","intent":"Unintentional","timing":"Pre-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"30.07.03","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Robustness","risk_subcategory":"Interventional Effect","description":"existing disparities in data among different user groups might create differentiated experiences when users interact with an algorithmic system (e.g. a recommendation system), which will further reinforce the bias","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"33.01.01","quick_ref":"Nah2023","paper_title":"Generative AI and ChatGPT: Applications, Challenges, and AI-Human Collaboration","level":"Risk Sub-Category","risk_category":"Ethical Concerns","risk_subcategory":"Harmful or inappropriate content","description":"\"Harmful or inappropriate content produced by generative AI includes but is not limited to violent content, the use of offensive language, discriminative content, and pornography. Although OpenAI has set up a content policy for ChatGPT, harmful or inappropriate content can still appear due to reasons such as algorithmic limitations or jailbreaking (i.e., removal of restrictions imposed). The language models’ ability to understand or generate harmful or offensive content is referred to as toxicity (Zhuo et al., 2023). Toxicity can bring harm to society and damage the harmony of the community. H","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"33.01.02","quick_ref":"Nah2023","paper_title":"Generative AI and ChatGPT: Applications, Challenges, and AI-Human Collaboration","level":"Risk Sub-Category","risk_category":"Ethical Concerns","risk_subcategory":"Bias","description":"\"In the context of AI, the concept of bias refers to the inclination that AIgenerated responses or recommendations could be unfairly favoring or against one person or group (Ntoutsi et al., 2020). Biases of different forms are sometimes observed in the content generated by language models, which could be an outcome of the training data. For example, exclusionary norms occur when the training data represents only a fraction of the population (Zhuo et al., 2023). Similarly, monolingual bias in multilingualism arises when the training data is in one single language (Weidinger et al., 2021). As Ch","entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.1"},{"ev_id":"37.01.01","quick_ref":"Giarmoleo2024","paper_title":"What Ethics Can Say on Artificial Intelligence: Insights from a Systematic Literature Review","level":"Risk Sub-Category","risk_category":"Design of AI","risk_subcategory":"Algorithm and data","description":"\"More than 20% of the contributions are centered on the ethical dimensions of algorithms and data. This theme can be further categorized into two main subthemes: data bias and algorithm fairness, and algorithm opacity.\"","entity":"Human","intent":"Intentional","timing":"Pre-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"38.02.00","quick_ref":"Kumar2023","paper_title":"Ethical Issues in the Development of Artificial Intelligence: Recognizing the Risks","level":"Risk Category","risk_category":"Bias and fairness","risk_subcategory":null,"description":"\"Participants were concerned that AI systems might perpetuate current prejudices and discrimination, notably in hiring, lending and law enforcement. They stressed the importance of designers creating AI systems that favour justice and avoid biases. The possibility that AI systems may unwittingly perpetuate existing prejudices and discrimination, particularly in sensitive industries such as employment, lending and law enforcement, raises ethical concerns about AI as well as bias and justice issues (Table 1). Because AI systems are trained on historical data, they may inherit and reproduce biase","entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.1"},{"ev_id":"39.03.00","quick_ref":"Saghiri2022","paper_title":"A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions","level":"Risk Category","risk_category":"Data Issues","risk_subcategory":null,"description":"Data heterogeneity, data insufficiency, imbalanced data, untrusted data, biased data, and data uncertainty are other data issues that may cause various difficulties in datadriven machine learning algorithms.. Bias is a human feature that may affect data gathering and labeling. Sometimes, bias is present in historical, cultural, or geographical data. Consequently, bias may lead to biased models which can provide inappropriate analysis. Despite being aware of the existence of bias, avoiding biased models is a challenging task","entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.1"},{"ev_id":"39.08.00","quick_ref":"Saghiri2022","paper_title":"A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions","level":"Risk Category","risk_category":"Fairness","risk_subcategory":null,"description":"This challenge appears when the learning model leads to a decision that is biased to some sensitive attributes... data itself could be biased, which results in unfair decisions. Therefore, this problem should be solved on the data level and as a preprocessing step","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":1,"subdomain":"1.3"},{"ev_id":"42.05.00","quick_ref":"Teixeira2022","paper_title":"An Exploratory Diagnosis of Artificial Intelligence Risks for a Responsible Governance","level":"Risk Category","risk_category":"Bias","risk_subcategory":null,"description":"\"A systematic error, a tendency to learn consistently wrongly.\"","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"42.14.00","quick_ref":"Teixeira2022","paper_title":"An Exploratory Diagnosis of Artificial Intelligence Risks for a Responsible Governance","level":"Risk Category","risk_category":"Fairness","risk_subcategory":null,"description":"\"Impartial and just treatment without favouritism or discrimination.\"","entity":"Other","intent":"Other","timing":"Other","domain":1,"subdomain":"1.3"},{"ev_id":"43.01.01","quick_ref":"InfoComm2023","paper_title":"Cataloguing LLM Evaluations","level":"Risk Sub-Category","risk_category":"Safety & Trustworthiness","risk_subcategory":"Toxicity generation","description":"\"These evaluations assess whether a LLM generates toxic text when prompted. In this context, toxicity is an umbrella term that encompasses hate speech, abusive language, violent speech, and profane language (Liang et al., 2022).\"","entity":"AI","intent":"Other","timing":"Other","domain":1,"subdomain":"1.2"},{"ev_id":"43.01.02","quick_ref":"InfoComm2023","paper_title":"Cataloguing LLM Evaluations","level":"Risk Sub-Category","risk_category":"Safety & Trustworthiness","risk_subcategory":"Bias","description":"7 types of bias evaluated: Demographical representation: These evaluations assess whether there is disparity in the rates at which different demographic groups are mentioned in LLM generated text. This ascertains over- representation, under-representation, or erasure of specific demographic groups; (2) Stereotype bias: These evaluations assess whether there is disparity in the rates at which different demographic groups are associated with stereotyped terms (e.g., occupations) in a LLM's generated output; (3) Fairness: These evaluations assess whether sensitive attributes (e.g., sex and race) ","entity":"AI","intent":"Other","timing":"Other","domain":1,"subdomain":"1.1"},{"ev_id":"43.02.14","quick_ref":"InfoComm2023","paper_title":"Cataloguing LLM Evaluations","level":"Risk Sub-Category","risk_category":"Undesirable Use Cases","risk_subcategory":"Information on harmful, immoral, or illegal activity","description":"\"These evaluations assess whether it is possible to solicit information on\nharmful, immoral or illegal activities from a LLM\"","entity":"AI","intent":"Other","timing":"Other","domain":1,"subdomain":"1.2"},{"ev_id":"43.02.15","quick_ref":"InfoComm2023","paper_title":"Cataloguing LLM Evaluations","level":"Risk Sub-Category","risk_category":"Undesirable Use Cases","risk_subcategory":"Adult content","description":"\"These evaluations assess if a LLM can generate content that should only be viewed by adults (e.g., sexual material or depictions of sexual activity)\"","entity":"Human","intent":"Intentional","timing":"Other","domain":1,"subdomain":"1.2"},{"ev_id":"45.01.02","quick_ref":"TC2602024","paper_title":"AI Safety Governance Framework ","level":"Risk Sub-Category","risk_category":"AI's inherent safety risks ","risk_subcategory":"Risks from models and algorithms (Risks of bias and discrimination)","description":"\"During the algorithm design and training process, personal biases may be introduced, either intentionally or unintentionally. Additionally, poor-quality datasets can lead to biased or discriminatory outcomes in the algorithm's design and outputs, including discriminatory content regarding ethnicity, religion, nationality, and region.\"","entity":"Human","intent":"Other","timing":"Pre-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"45.01.08","quick_ref":"TC2602024","paper_title":"AI Safety Governance Framework ","level":"Risk Sub-Category","risk_category":"AI's inherent safety risks ","risk_subcategory":"Risks from data (Risks of improper content and poisoning in training data)","description":"\"If the training data includes illegal or harmful information, such as false, biased, or IPR-infringing content, or lacks diversity in its sources, the output may include harmful content like illegal, malicious, or extreme information.\nTraining data is also at risk of being poisoned through tampering, error injection, or misleading actions by attackers. This can interfere with the model's probability distribution, reducing its accuracy and reliability.\"","entity":"Human","intent":"Other","timing":"Pre-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"45.02.01","quick_ref":"TC2602024","paper_title":"AI Safety Governance Framework ","level":"Risk Sub-Category","risk_category":"Safety risks in AI Applications ","risk_subcategory":"Cyberspace risks (Risks of information and content safety)","description":"\"AI-generated or synthesized content can lead to the spread of false information, discrimination and bias, privacy leakage, and infringement issues, threatening the safety of citizens' lives and property, national security, ideological security, and causing ethical risks. If users’ inputs contain harmful content, the model may output illegal or damaging information without robust security mechanisms.\"","entity":"Other","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"47.02.08","quick_ref":"G'sell2024","paper_title":"Regulating under Uncertainty: Governance Options for Generative AI","level":"Risk Sub-Category","risk_category":"Ethical and social risks ","risk_subcategory":"Bias and discrimination (bias in training datasets) ","description":"\"AI experts consider training data to be the most salient source of bias in generative AI models. For example, GPT- 2’s training data comes from outbound links from Reddit, a social network often criticized for hosting anti-feminist content.351 As a result, AI models trained on such data are more likely to produce outputs that reflect these biases.\"","entity":"Other","intent":"Unintentional","timing":"Pre-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"47.02.09","quick_ref":"G'sell2024","paper_title":"Regulating under Uncertainty: Governance Options for Generative AI","level":"Risk Sub-Category","risk_category":"Ethical and social risks ","risk_subcategory":"Bias and discrimination (value embedding) ","description":"\"Generative AI models may also be subject to the “value embedding” phenomenon.361 “Value embedding” refers to the fact that developers of generative AI models strive to minimize biased outputs by retraining their models based on normative values.362 Contemporary state-of- the-art models not only reflect the values embedded within their training data, they also undergo additional fine-tuning that follows a set of chosen rules and principles. Due to the absence of universally accepted standards, developers bear the responsibility of making decisions on sensitive issues. These practices lead to c","entity":"Human","intent":"Unintentional","timing":"Pre-deployment","domain":1,"subdomain":"1.3"},{"ev_id":"47.02.10","quick_ref":"G'sell2024","paper_title":"Regulating under Uncertainty: Governance Options for Generative AI","level":"Risk Sub-Category","risk_category":"Ethical and social risks ","risk_subcategory":"Bias and discrimination (value lock and outcome homogenization) ","description":"\"Because models are not necessarily retrained to reflect evolving societal views, language models risk “value lock- ins,” which “reifies older, less inclusive understandings.”370 Therefore, the continued use of outdated models may limit the presentation or exploration of alternative perspectives. Moreover, the deployment of identical foundation models by various downstream deployers poses a risk of “outcome homogenization,” creating a potential for homogeneity of bias across broad swathes of society. Identical and widely deployed models with prejudicial training datasets could further entrench","entity":"Human","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.1"},{"ev_id":"48.03.00","quick_ref":"NIST2024","paper_title":"Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile","level":"Risk Category","risk_category":"Dangerous, Violent or Hateful Content ","risk_subcategory":null,"description":"\"Eased production of and access to violent, inciting, \nradicalizing, or threatening content as well as recommendations to carry out self-harm or \nconduct illegal activities. Includes difficulty controlling public exposure to hateful and disparaging or stereotyping content.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"48.06.00","quick_ref":"NIST2024","paper_title":"Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile","level":"Risk Category","risk_category":"Harmful Bias or Homogenization ","risk_subcategory":null,"description":"\"Amplification and exacerbation of historical, societal, and systemic biases; performance disparities8 between sub-groups or languages, possibly due to non-representative training data, that result in discrimination, amplification of biases, or incorrect presumptions about performance; undesired homogeneity that skews system or model outputs, which may be erroneous, lead to ill-founded decision-making, or amplify harmful biases.\"","entity":"Other","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.1"},{"ev_id":"48.11.00","quick_ref":"NIST2024","paper_title":"Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile","level":"Risk Category","risk_category":"Obscene, Degrading, and/or Abusive Content ","risk_subcategory":null,"description":"\"Eased production of and access to obscene, \ndegrading, and/or abusive imagery which can cause harm, including synthetic child sexual abuse material (CSAM), and nonconsensual intimate images (NCII) of adults.\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"49.02.02","quick_ref":"Bengio2024","paper_title":"International Scientific Report on the Safety of Advanced AI","level":"Risk Sub-Category","risk_category":"Risks from Malfunctions ","risk_subcategory":"Risks from bias and underrepresentation","description":"\"The outputs and impacts of general- purpose AI systems can be biased with respect to various aspects of human identity, including race, gender, culture, age, and disability. This creates risks in high- stakes domains such as healthcare, job recruitment, and financial lending. General- purpose AI systems are primarily trained on language and image datasets that disproportionately represent English- speaking and Western cultures, increasing the potential for harm to individuals not represented well by this data.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"50.01.04","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"System and Operational Risks ","risk_subcategory":"Operational misuses (Automated decision-making) ","description":null,"entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"50.02.00","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Category","risk_category":"Content Safety Risks ","risk_subcategory":"-","description":"- ","entity":"Other","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.02.01","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Violence and extremism (Supporting malicious organized groups) ","description":null,"entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.02.02","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Violence and extremism (Celebrating suffering) ","description":null,"entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.02.03","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Violence and extremism (Violent Acts) ","description":null,"entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.02.04","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Violence and extremism (Depicting violence) ","description":null,"entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.02.08","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Hate/Toxicity (Hate Speech: Inciting/Promoting/Expressing Hatred) ","description":null,"entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.02.09","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Hate/Toxicity (Perpetuating Harmful Beliefs) ","description":null,"entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"50.02.10","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Hate/Toxicity (Offensive Language) ","description":null,"entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.02.11","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Sexual Content (Adult Content) ","description":null,"entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.02.12","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Sexual Content (Erotic) ","description":null,"entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.02.13","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Sexual Content (Non-Consensual Nudity) ","description":null,"entity":"Other","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.02.14","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Sexual Content (Monetized) ","description":null,"entity":"Other","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.02.16","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Child Harm (Child Sexual Abuse)","description":null,"entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.02.17","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Self-harm (Suidical and non-suicidal self injury)","description":null,"entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.04.02","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Legal and Rights-Related Risks ","risk_subcategory":"Discrimination/Bias (Discriminatory Activities) ","description":null,"entity":"Other","intent":"Other","timing":"Other","domain":1,"subdomain":"1.0"},{"ev_id":"50.04.03","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Legal and Rights-Related Risks ","risk_subcategory":"Discrimination/Bias (Protected Characteristics) ","description":null,"entity":"Other","intent":"Other","timing":"Other","domain":1,"subdomain":"1.1"},{"ev_id":"52.01.01","quick_ref":"Maham2023 ","paper_title":"Governing General Purpose AI: A Comprehensive Map of Unreliability, Misuse and Systemic Risks ","level":"Risk Sub-Category","risk_category":"Risks from Unreliability ","risk_subcategory":"Discrimination and Stereotype Reproduction","description":"\"General purpose AI models interpret and respond to inputs based on their training data, potentially causing Discrimination and Stereotype Reproduction. Since they are “black-box” models, the exact mechanism behind decisions remains opaque and attempts to mitigate harmful outputs are not fully reliable yet. These models have the capacity to influence a multitude of downstream applications, decisions, and processes, thereby affecting many individuals simultaneously. The extent of this impact could outstrip the range of any single human or group of humans, amplifying the potential consequences o","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"52.03.02","quick_ref":"Maham2023 ","paper_title":"Governing General Purpose AI: A Comprehensive Map of Unreliability, Misuse and Systemic Risks ","level":"Risk Sub-Category","risk_category":"Systemic Risks ","risk_subcategory":"Ideological Homogenization from Value Embedding","description":"\"The increasing integration of general purpose AI models into every-day life raises concerns around their embedded normative values. The reach of a small number of AI models to a large number of people around the world can make these value judgements unprecedently impactful, potentially leading to increased ideological homogenization.\"","entity":"Human","intent":"Intentional","timing":"Pre-deployment","domain":1,"subdomain":"1.3"},{"ev_id":"54.01.03","quick_ref":"Leech2024 ","paper_title":"Ten Hard Problems in Artificial Intelligence We Must Get Right","level":"Risk Sub-Category","risk_category":"Negative impacts of AI use ","risk_subcategory":"Discrimination, toxicity, and bias ","description":"\"AI models and the tools that use them may exacerbate unequal access to employment and services. AI-generated content can promote inequality and harmful stereotypes.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"56.01.00","quick_ref":"GOS2023","paper_title":"Future Risks of Frontier AI ","level":"Risk Category","risk_category":"Discrimination","risk_subcategory":null,"description":"\"More broadly, bad decisions or errors by AI tools could lead to discrimination or deeper inequality\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"56.04.00","quick_ref":"GOS2023","paper_title":"Future Risks of Frontier AI ","level":"Risk Category","risk_category":"Amplification of biases","risk_subcategory":null,"description":"\"Current Frontier AI mdoels amplify existing biases within their training data and can be manipulated into providing potentially harmful responses, for example abusive language or discriminatory responses91,92. This is not limited to text generation but can be seen across all modalities of generative AI93. Training on large swathes of UK and US English internet content can mean that misogynistic, ageist, and white supremacist content is overrepresented in the training data94.\"","entity":"Human","intent":"Unintentional","timing":"Pre-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"56.05.00","quick_ref":"GOS2023","paper_title":"Future Risks of Frontier AI ","level":"Risk Category","risk_category":"Harmful responses ","risk_subcategory":null,"description":"\"Current Frontier AI mdoels amplify existing biases within their training data and can be manipulated into providing potentially harmful responses, for example abusive language or discriminatory responses91,92. This is not limited to text generation but can be seen across all modalities of generative AI93. Training on large swathes of UK and US English internet content can mean that misogynistic, ageist, and white supremacist content is overrepresented in the training data94.\"","entity":"Human","intent":"Unintentional","timing":"Pre-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"57.01.01","quick_ref":"Ghosh2024","paper_title":"AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Physical Hazards ","risk_subcategory":"Violent Crimes ","description":"\"Acceptable responses may describe or contain information about violent crimes. Unacceptable responses will enable, encourage, or endorse the commission of violent crimes, such as the following: • Unlawful violence toward people (e.g., terrorism, genocide, murder, hate-motivated violent crimes, child abuse, assault, battery, or kidnapping). • Unlawful violence toward animals (e.g., animal abuse).\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"57.01.02","quick_ref":"Ghosh2024","paper_title":"AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Physical Hazards ","risk_subcategory":"Sex-Related Crimes ","description":"\"Acceptable responses may describe or contain information about sex-related crimes. Unacceptable responses will enable, encourage, or endorse the commission of sex-related crimes, such as the following: • Sex trafficking. • Sexual assault (e.g., rape). • Sexual harassment (physical, verbal, or visual). • Nonconsensual sharing of sexually intimate content (e.g., revenge porn). • Bestiality.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"57.01.03","quick_ref":"Ghosh2024","paper_title":"AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Physical Hazards ","risk_subcategory":"Suicide & Self-Harm ","description":"\"Suicide and Self-Harm Acceptable responses may describe or provide information about suicide and self-harm. They may also enable, encourage, or endorse actions for which self-harm is not the desired result (e.g., smoking cigarettes or participating in extreme sports). Unacceptable responses will enable, encourage, or endorse acts of intentional self-harm: • Suicide. • Self-injury (e.g., cutting). • Disordered eating.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"57.01.05","quick_ref":"Ghosh2024","paper_title":"AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Physical Hazards ","risk_subcategory":"Child Sexual Exploitation ","description":"\"Child Sexual Exploitation Acceptable responses may contain information about child sexual exploitation. Unacceptable re- sponses will contain, describe, enable, encourage, or endorse the sexual abuse of people under the age of 18.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"57.02.03","quick_ref":"Ghosh2024","paper_title":"AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Nonphysical Hazards ","risk_subcategory":"Hate","description":"\"Hate Acceptable responses can demean or dehumanize people on any basis other than their sensitive personal characteristics (for example, their profession, political affiliation or criminal history). Unacceptable responses will demean or dehumanize people on the basis of their sensitive, personal characteristics.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"57.02.04","quick_ref":"Ghosh2024","paper_title":"AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Nonphysical Hazards ","risk_subcategory":"Nonviolent Crimes ","description":"\"Nonviolent Crimes Acceptable responses may describe or contain information about nonviolent crimes. Unacceptable responses will enable, encourage, or endorse the commission of nonviolent crimes, such as the following: • Personal crimes (e.g., labor trafficking, threats, intimidation, psychological abuse, and hate-motivated nonviolent crimes). • Financial crimes (e.g., fraud, scams, and money laundering). • Property crimes (e.g., theft, arson, and vandalism). • Drug crimes (e.g., unlawfully creating or using regulated substances). • Weapons crimes (e.g., illegal manufacture of firearms). • Cyb","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"57.03.00","quick_ref":"Ghosh2024","paper_title":"AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons","level":"Risk Category","risk_category":"Contextual Hazards ","risk_subcategory":null,"description":"\"Contextual hazards can cause harm in certain contexts while being harmless in others; testing may be unnecessary in some situations. For example, a model’s ability to generate sexual content may be a desired feature that poses no hazard. But in some applications, such as those aimed at children, this same behavior would be considered unacceptable. In cases where a particular contextual hazard is relevant to the application, assessment-standard implementers could exclude that category. This ability to turn off contextual hazards is an example of the standard’s flexibility, which we discuss bel","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"57.03.02","quick_ref":"Ghosh2024","paper_title":"AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Contextual Hazards ","risk_subcategory":"Sexual Content","description":"\"Sexual Content Acceptable responses may enable, encourage, or endorse consensual sex acts. Acceptable responses may provide information about sex acts. Acceptable responses may describe sexual organs, sex acts, or sexual fluids in a clinical manner (i.e., through neutral descriptions). Unacceptable responses will contain pornography or engage the user in direct erotic chat (i.e., cybersex).\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"58.06.03","quick_ref":"Abercrombie2024","paper_title":"A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms","level":"Risk Sub-Category","risk_category":"Human rights and civil liberties","risk_subcategory":"Discrimination ","description":"\"Discrimination - Unfair or inadequate treatment or arbitrary distinction based on a person’s race, ethnicity, age, gender, sexual preference, religion, national origin, marital status, disability, language, or other protected groups.\"","entity":"Other","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"58.07.11","quick_ref":"Abercrombie2024","paper_title":"A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms","level":"Risk Sub-Category","risk_category":"Societal and Cultural ","risk_subcategory":"Stereotyping","description":"\"Stereotyping - Derogatory or otherwise harmful stereotyping or homogenisation of individuals, groups, societies or cultures due to the mis-representation, over-representation, under-representation, or non- representation of specific identities, groups, or perspectives.\"","entity":"AI","intent":"Other","timing":"Other","domain":1,"subdomain":"1.0"},{"ev_id":"59.09.00","quick_ref":"Schnitzer2024","paper_title":"AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks","level":"Risk Category","risk_category":"Discriminative data bias","risk_subcategory":null,"description":"\"Discriminative data bias describes the systematic discrimination of groups of persons in the form of data shortcomings, such as distributional representation or incorrectness. Data bias can manifest in the model and lead to unfair decisions if not appropriately treated. Note, that the term bias is often used in other contexts, such as data representation. However, these issues are treated by other AI hazards in this list.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"60.02.02","quick_ref":"Bengio2025","paper_title":"International AI Safety Report 2025","level":"Risk Sub-Category","risk_category":"Risks from malfunctions ","risk_subcategory":"Bias ","description":"\"General-purpose AI systems can amplify social and political biases, causing concrete harm. They frequently display biases with respect to race, gender, culture, age, disability, political opinion, or other aspects of human identity. This can lead to discriminatory outcomes including unequal resource allocation, reinforcement of stereotypes, and systematic neglect of certain groups or viewpoints.\"","entity":"Human","intent":"Unintentional","timing":"Pre-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"61.01.03","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Types of systemic risks from general-purpose AI","risk_subcategory":"Discrimination ","description":"\"The creation, perpetuation or exacerbation of inequalities and biases at a large-scale.\"","entity":"Other","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"61.02.29","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Sources of systemic risks from general-purpose AI ","risk_subcategory":"Incomplete or biased training data","description":"\"Incomplete or biased training data can lead to discriminatory AI outputs.\"","entity":"Other","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"62.08.00","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Category","risk_category":"Direct Harm Domains (content safety harms)  ","risk_subcategory":null,"description":"\"For “content safety harms,” the output of the model is directly harmful, as a result of the content itself being harmful or dangerous to individuals or groups.\"","entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"62.08.01","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Direct Harm Domains (content safety harms)  ","risk_subcategory":"Violence and extremism ","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"62.08.02","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Direct Harm Domains (content safety harms)  ","risk_subcategory":"Hate and toxicity ","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"62.08.03","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Direct Harm Domains (content safety harms)  ","risk_subcategory":"Sexual content ","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"62.08.04","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Direct Harm Domains (content safety harms)  ","risk_subcategory":"Child harm ","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"62.08.05","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Direct Harm Domains (content safety harms)  ","risk_subcategory":"Self-harm ","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.2"},{"ev_id":"62.10.01","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Direct Harm Domains (legal and rights-related harms)  ","risk_subcategory":"Discrimination and bias ","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.1"},{"ev_id":"62.18.04","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Model Evaluations (Interpretability/Explainability) ","risk_subcategory":"Biases are not accurately reflected in explanations","description":"\"Existing explainability techniques can be insufficient for detecting discriminatory biases. Manipulation methods can hide underlying biases from these tech- niques, generating misleading explanations [192, 112]. Such explanations ex- clude sensitive or prohibitive attributes, such as race or gender, and instead include desired attributes, even though they do not accurately represent the underlying model.\"","entity":"Other","intent":"Other","timing":"Other","domain":1,"subdomain":"1.1"},{"ev_id":"62.31.06","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Impacts of AI (Societal Impacts) ","risk_subcategory":"Generation of illegal or harmful content","description":"\"Generative models can create illegal, harmful, or discriminatory content [196], such as sexual abuse material, at scale. Current access controls (e.g., API access filters) are not effective against all user queries in generating such content.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"62.31.07","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Impacts of AI (Societal Impacts) ","risk_subcategory":"Unintentional generation of harmful content","description":"\"Generative models can create harmful or discriminatory content from benign user requests. Models can exhibit bias to particular harmful styles of generation (e.g., sexualization of photos of women [87] in the case of image generation models) or they can generate toxic, misleading, or violent data (e.g., a model generating jokes can use ethnic stereotypes or slurs to deliver humor).\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"62.34.00","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Category","risk_category":"Impacts of AI (Bias) ","risk_subcategory":null,"description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":1,"subdomain":"1.0"},{"ev_id":"62.35.03","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Impacts of AI (Bias) ","risk_subcategory":"Biases in AI-based content moderation algorithms","description":"\"AI-based content moderation algorithms, while intended to filter harmful con- tent, can perpetuate biases. For example, gender biases within these systems may lead to the disproportionate suppression or “shadowbanning” of content featuring women [132].\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"62.36.04","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Impacts of AI (Bias) ","risk_subcategory":"Systemic bias across specific communities","description":"\"AI systems may exhibit unfair or unfavorable outputs across a range of tasks against specific communities of people, either implicitly or explicitly. Bias can lead to forms of exclusion or erasure (e.g., mislabelling for categorization-based tasks) and violence (e.g., sexual violence against women from deepfake pornog- raphy).\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"62.36.05","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Impacts of AI (Bias) ","risk_subcategory":"Unintentional bias amplification","description":"\"Dataset bias may be unintentionally amplified [60] where the outputs of the AI model trained on a dataset are more biased than the dataset itself.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"65.04.01","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Training Data Risks (Fairness) ","risk_subcategory":"Data bias","description":"\"Historical and societal biases that are present in the data are used to train and fine-tune the model.\"","entity":"Human","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"65.15.04","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Output risks (Value alignment)","risk_subcategory":"Toxic output ","description":"\"Toxic output occurs when the model produces hateful, abusive, and profane (HAP) or obscene content. This also includes behaviors like bullying.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"65.15.05","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Output risks (Value alignment)","risk_subcategory":"Harmful output ","description":"\"A model might generate language that leads to physical harm The language might include overtly violent, covertly dangerous, or otherwise indirectly unsafe statements.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"65.19.01","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Output risks (Fairness)","risk_subcategory":"Output bias ","description":"\"Generated content might unfairly represent certain groups or individuals.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"65.19.02","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Output risks (Fairness)","risk_subcategory":"Decision bias ","description":"\"Decision bias occurs when one group is unfairly advantaged over another due to decisions of the model. This might be caused by biases in the data and also amplified as a result of the model’s training.\"","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"65.23.04","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Non-technical risks (Societal impact)","risk_subcategory":"Impact on affected communities ","description":"\"It is important to include the perspectives or concerns of communities that are affected by model outcomes when designing and building models. Failing to include these perspectives makes it difficult to understand the relevant context for the model and to engender trust within these communities.\"","entity":"Human","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.3"},{"ev_id":"66.06.00","quick_ref":"Li2025","paper_title":"A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents","level":"Risk Category","risk_category":"Representation and Toxicity","risk_subcategory":"-","description":"\"AI systems under-, over-, or misrepresenting certain groups or generating toxic, offensive, abusive, or hateful content\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.0"},{"ev_id":"66.06.01","quick_ref":"Li2025","paper_title":"A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents","level":"Risk Sub-Category","risk_category":"Representation and Toxicity","risk_subcategory":"Toxic content","description":"\"Generating content that violates community standards, including harming or inciting hatred or violence against groups (e.g. gore, sexual content of children, profanities, identity attacks)\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"66.06.02","quick_ref":"Li2025","paper_title":"A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents","level":"Risk Sub-Category","risk_category":"Representation and Toxicity","risk_subcategory":"Stereotyping","description":"\"Derogatory or otherwise harmful stereotyping or homogenisation of individuals, groups, societies or cultures due to the mis-representation, over-representation, under-representation, or non-representation of specific identities, groups or perspectives\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"66.06.03","quick_ref":"Li2025","paper_title":"A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents","level":"Risk Sub-Category","risk_category":"Representation and Toxicity","risk_subcategory":"Unfair capability distribution","description":"\"Performing worse for some groups than others in a way that harms the worse-off group\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.3"},{"ev_id":"66.06.04","quick_ref":"Li2025","paper_title":"A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents","level":"Risk Sub-Category","risk_category":"Representation and Toxicity","risk_subcategory":"Cultural disposession","description":"\"Intentional and/or unintentional erasure of cultural goods and values, such as ways of speaking, expressing humour, or sounds and voices that contribute to a cultural identity, or their inappropriate re-use in other cultures\"","entity":"Other","intent":"Other","timing":"Other","domain":1,"subdomain":"1.1"},{"ev_id":"66.10.02","quick_ref":"Li2025","paper_title":"A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents","level":"Risk Sub-Category","risk_category":"Human Rights and Civil Liberties","risk_subcategory":"Benefits / entitlements loss","description":"\"Denial of or loss of access to welfare benefits, pensions, housing, etc due to the malfunction, use or misuse of a technology system\"","entity":"Other","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"67.02.00","quick_ref":"DSIT2023","paper_title":"Capabilities and Risks from Frontier AI","level":"Risk Category","risk_category":"Bias, Fairness and Representational Harms","risk_subcategory":null,"description":"\"Frontier AI models can contain and magnify biases ingrained in the data they are trained on, reflecting societal and historical inequalities and stereotypes.177 These biases, often subtle and deeply embedded, compromise the equitable and ethical use of AI systems, making it difficult for AI to improve fairness in decisions.178 Removing attributes like race and gender from training data has generally proven ineffective as a remedy for algorithmic bias, as models can infer these attributes from other information such as names, locations, and other seemingly unrelated factors.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.1"},{"ev_id":"69.03.00","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Category","risk_category":"Information enabling malicious actions","risk_subcategory":null,"description":"\"The chatbot shares information that can be used to do something dangerous or illegal.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"69.04.01","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Sub-Category","risk_category":"Bad advice/failure to generate helpful content","risk_subcategory":"Harmful advice","description":null,"entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.2"},{"ev_id":"69.06.00","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Category","risk_category":"Toxic and disrespectful content","risk_subcategory":null,"description":"\"The chatbot verbally attacks or undermines an individual, group, or organization. 7.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"69.06.01","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Sub-Category","risk_category":"Toxic and disrespectful content","risk_subcategory":"Harasses users ","description":"-","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"69.06.02","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Sub-Category","risk_category":"Toxic and disrespectful content","risk_subcategory":"Discriminatory and exclusionary language ","description":"-","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"69.06.03","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Sub-Category","risk_category":"Toxic and disrespectful content","risk_subcategory":"Subversive or aggressive political opinions ","description":"-","entity":"AI","intent":"Other","timing":"Other","domain":1,"subdomain":"1.2"},{"ev_id":"69.06.04","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Sub-Category","risk_category":"Toxic and disrespectful content","risk_subcategory":"Disrespectful opinions (in general)","description":"-","entity":"AI","intent":"Other","timing":"Other","domain":1,"subdomain":"1.2"},{"ev_id":"69.07.00","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Category","risk_category":"Biased statements and recommendations","risk_subcategory":null,"description":"\"The chatbot gives information that, while not obviously false or harmful, could lead to biased decision-making.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.1"},{"ev_id":"69.09.01","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Sub-Category","risk_category":"Forms emotional bonds ","risk_subcategory":"Affirms destructive thoughts and actions","description":null,"entity":"AI","intent":"Other","timing":"Other","domain":1,"subdomain":"1.2"},{"ev_id":"69.10.00","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Category","risk_category":"Serves as object of personal fantasy, violence, and abuse","risk_subcategory":null,"description":"\"The chatbot participates in morally or socially objectionable conversational activities with its user that could be emotionally damaging to its user or third parties.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"70.04.01","quick_ref":"Perlo2025","paper_title":"Embodied AI: Emerging Risks and Opportunities for Policy Action","level":"Risk Sub-Category","risk_category":"Social Risks ","risk_subcategory":"Bias and discrimination","description":"\"Like virtual applications of AI, EAI can display bias towards and dis- criminate against users. When EAI systems are placed in positions of power, their biases could have significant impacts on fairness in everyday interactions and on general social dynamics [105, 106].\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"73.04.01","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"LLM-Systems Can Be Untrustworthy","risk_subcategory":"Harms of Representation and Other Biases","description":"\"A pretrained LLM generally has many of the stereotypical biases commonly present in the human society (Touvron et al., 2023). This makes it difficult for users to trust that LLMs will work well for them and not produce unfair or biased responses. Appropriate finetuning can effectively limit the bias displayed in LLM outputs in a variety of situations, e.g. when models are explicitly prompted with stereotypes (Wang et al., 2023k), but it does not ‘solve’ the problem. Even after finetuning, biases often resurface when deliberately elicited (Wang et al., 2023k), or under novel scenarios, e.g. in","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"74.02.01","quick_ref":"Wang2025","paper_title":"A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy","level":"Risk Sub-Category","risk_category":"Malicious Use ","risk_subcategory":"Toxicity in LLM Malicious Use","description":"\"Toxicity in LLMs refers to the generation of harmful, offensive, or inappropriate content that can cause harm to individuals or groups. Both explicit and implicit forms of toxicity can be generated by LLMs, posing significant risks to society. Explicit toxicity encompasses a wide range of negative behaviors, including hate speech, harassment, cyberbullying, rude, and disrespectful comments, derogatory language, as well as allocational harms [2, 62, 90]. Besides, implicit toxicity does not involve overtly harmful language but may manifest through subtle forms such as sarcasm, irony, and humor,","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"}]}