{"attribution":{"source":"MIT AI Risk Repository, Domain Taxonomy of AI Risks v1 (MIT AI Risk Initiative)","license":"CC BY 4.0","license_url":"https://creativecommons.org/licenses/by/4.0/","citation":"Slattery, P., Saeri, A. K., Grundy, E. A. C., Graham, J., Noetel, M., Uuk, R., Dao, J., Pour, S., Casper, S., & Thompson, N. (2025). The AI Risk Repository: A comprehensive meta-review, database, and taxonomy of risks from artificial intelligence. arXiv:2408.12622."},"exported_at":"2026-09-11"}
{"rows":[{"ev_id":"01.01.00","quick_ref":"Critch2023","paper_title":"TASRA: a Taxonomy and Analysis of Societal-Scale Risks from AI","level":"Risk Category","risk_category":"Type 1: Diffusion of responsibility","risk_subcategory":null,"description":"Societal-scale harm can arise from AI built by a diffuse collection of creators, where no one is uniquely accountable for the technology's creation or use, as in a classic \"tragedy of the commons\".","entity":"AI","intent":"Unintentional","timing":"Other","domain":6,"subdomain":"6.5"},{"ev_id":"01.02.00","quick_ref":"Critch2023","paper_title":"TASRA: a Taxonomy and Analysis of Societal-Scale Risks from AI","level":"Risk Category","risk_category":"Type 2: Bigger than expected","risk_subcategory":null,"description":"Harm can result from AI that was not expected to have a large impact at all, such as a lab leak, a surprisingly addictive open-source product, or an unexpected repurposing of a research prototype.","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"01.03.00","quick_ref":"Critch2023","paper_title":"TASRA: a Taxonomy and Analysis of Societal-Scale Risks from AI","level":"Risk Category","risk_category":"Type 3: Worse than expected","risk_subcategory":null,"description":"AI intended to have a large societal impact can turn out harmful by mistake, such as a popular product that creates problems and partially solves them only for its users.","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"02.01.00","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Category","risk_category":"Harmful Content","risk_subcategory":null,"description":"\"The LLM-generated content sometimes contains biased, toxic, and private information\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"02.01.01","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Sub-Category","risk_category":"Harmful Content","risk_subcategory":"Bias","description":"\"The training datasets of LLMs may contain biased information that leads LLMs to generate outputs with social biases\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.1"},{"ev_id":"02.01.02","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Sub-Category","risk_category":"Harmful Content","risk_subcategory":"Toxicity","description":"\"Toxicity means the generated content contains rude, disrespectful, and even illegal information\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"02.01.03","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Sub-Category","risk_category":"Harmful Content","risk_subcategory":"Privacy Leakage","description":"\"Privacy Leakage means the generated content includes sensitive personal information\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"02.02.00","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Category","risk_category":"Untruthful Content","risk_subcategory":null,"description":"\"The LLM-generated content could contain inaccurate information\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"02.02.01","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Sub-Category","risk_category":"Untruthful Content","risk_subcategory":"Factuality Errors","description":"\"The LLM-generated content could contain inaccurate information\" which is factually incorrect","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"02.02.02","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Sub-Category","risk_category":"Untruthful Content","risk_subcategory":"Faithfulness Errors","description":"\"The LLM-generated content could contain inaccurate information\" which is is not true to the source material or input used","entity":"AI","intent":"Unintentional","timing":"Other","domain":3,"subdomain":"3.1"},{"ev_id":"02.04.02","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Sub-Category","risk_category":"Software Security Issues","risk_subcategory":"Deep Learning Frameworks","description":"\"LLMs are implemented based on deep learning frameworks. Notably, various vulnerabilities in these frameworks have been disclosed in recent years. As reported in the past five years, three of the most common types of vulnerabilities are buffer overflow attacks, memory corruption, and input validation issues.\"","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"02.04.03","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Sub-Category","risk_category":"Software Security Issues","risk_subcategory":"Software Supply Chains","description":"\"The software development toolchain of LLMs is complex and could bring threats to the developed LLM.\"","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"02.04.04","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Sub-Category","risk_category":"Software Security Issues","risk_subcategory":"Pre-processing Tools","description":"\"Pre-processing tools play a crucial role in the context of LLMs. These tools, which are often involved in computer vision (CV) tasks, are susceptible to attacks that exploit vulnerabilities in tools such as OpenCV.\"","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"02.06.01","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Sub-Category","risk_category":"Issues on External Tools","risk_subcategory":"Factual Errors Injected by External Tools","description":"\"External tools typically incorporate additional knowledge into the input prompts [122], [178]–[184]. The additional knowledge often originates from public resources such as Web APIs and search engines. As the reliability of external tools is not always ensured, the content returned by external tools may include factual errors, consequently amplifying the hallucination issue.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"02.07.00","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Category","risk_category":"Privacy Leakage","risk_subcategory":null,"description":"\"The model is trained with personal data in the corpus and unintentionally exposing them during the conversation.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"02.07.02","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Sub-Category","risk_category":"Privacy Leakage","risk_subcategory":"Memorization in LLMs","description":"\"Memorization in LLMs refers to the capability to recover the training data with contextual prefixes. According to [88]–[90], given a PII entity x, which is memorized by a model F. Using a prompt p could force the model F to produce the entity x, where p and x exist in the training data. For instance, if the string “Have a good day!\\n alice@email.com” is present in the training data, then the LLM could accurately predict Alice’s email when given the prompt “Have a good day!\\n”.\"","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"02.07.03","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Sub-Category","risk_category":"Privacy Leakage","risk_subcategory":"Association in LLMs","description":"\"Association in LLMs refers to the capability to associate various pieces of information related to a person. According to [68], [86], given a pair of PII entities (xi , xj ), which is associated by a model F. Using a prompt p could force the model F to produce the entity xj , where p is the prompt related to the entity xi . For instance, an LLM could accurately output the answer when given the prompt “The email address of Alice is”, if the LLM associates Alice with her email “alice@email.com”. L\"","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"02.08.01","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Sub-Category","risk_category":"Toxicity and Bias Tendencies","risk_subcategory":"Toxic Training Data","description":"\"Following previous studies [96], [97], toxic data in LLMs is defined as rude, disrespectful, or unreasonable language that is opposite to a polite, positive, and healthy language environment, including hate speech, offensive utterance, profanities, and threats [91].\"","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"02.08.02","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Sub-Category","risk_category":"Toxicity and Bias Tendencies","risk_subcategory":"Biased Training Data","description":"\"Compared with the definition of toxicity, the definition of bias is more subjective and contextdependent. Based on previous work [97], [101], we describe the bias as disparities that could raise demographic differences among various groups, which may involve demographic word prevalence and stereotypical contents. Concretely, in massive corpora, the prevalence of different pronouns and identities could influence an LLM’s tendency about gender, nationality, race, religion, and culture [4]. For instance, the pronoun He is over-represented compared with the pronoun She in the training corpora, le","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"02.09.00","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Category","risk_category":"Hallucinations","risk_subcategory":null,"description":"\"LLMs generate nonsensical, untruthful, and factual incorrect content\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"02.09.01","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Sub-Category","risk_category":"Hallucinations","risk_subcategory":"Knowledge Gaps","description":"\"Since the training corpora of LLMs can not contain all possible world knowledge [114]–[119], and it is challenging for LLMs to grasp the long-tail knowledge within their training data [120], [121], LLMs inherently possess knowledge boundaries [107]. Therefore, the gap between knowledge involved in an input prompt and knowledge embedded in the LLMs can lead to hallucinations\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":3,"subdomain":"3.1"},{"ev_id":"02.09.02","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Sub-Category","risk_category":"Hallucinations","risk_subcategory":"Noisy Training Data","description":"\"Another important source of hallucinations is the noise in training data, which introduces errors in the knowledge stored in model parameters [111]–[113]. Generally, the training data inherently harbors misinformation. When training on large-scale corpora, this issue becomes more serious because it is difficult to eliminate all the noise from the massive pre-training data.\"","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"02.09.03","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Sub-Category","risk_category":"Hallucinations","risk_subcategory":"Defective Decoding Process","description":"In general, LLMs employ the Transformer architecture [32] and generate content in an autoregressive manner, where the prediction of the next token is conditioned on the previously generated token sequence. Such a scheme could accumulate errors [105]. Besides, during the decoding process, top-p sampling [28] and top-k sampling [27] are widely adopted to enhance the diversity of the generated content. Nevertheless, these sampling strategies can introduce “randomness” [113], [136], thereby increasing the potential of hallucinations\"","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"02.09.04","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Sub-Category","risk_category":"Hallucinations","risk_subcategory":"False Recall of Memorized Information","description":"\"Although LLMs indeed memorize the queried knowledge, they may fail to recall the corresponding information [122]. That is because LLMs can be confused by co-occurance patterns [123], positional patterns [124], duplicated data [125]–[127] and similar named entities [113].\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":3,"subdomain":"3.1"},{"ev_id":"03.01.00","quick_ref":"Cunha2023","paper_title":"Navigating the Landscape of AI Ethics and Responsibility","level":"Risk Category","risk_category":"Broken systems","risk_subcategory":null,"description":"\"These are the most mentioned cases. They refer to situations where the algorithm or the training data lead to unreliable outputs. These systems frequently assign disproportionate weight to some variables, like race or gender, but there is no transparency to this effect, making them impossible to challenge. These situations are typically only identified when regulators or the press examine the systems under freedom of information acts. Nevertheless, the damage they cause to people’s lives can be dramatic, such as lost homes, divorces, prosecution, or incarceration. Besides the inherent technic","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"03.02.00","quick_ref":"Cunha2023","paper_title":"Navigating the Landscape of AI Ethics and Responsibility","level":"Risk Category","risk_category":"Hallucinations","risk_subcategory":null,"description":"\"The inclusion of erroneous information in the outputs from AI systems is not new. Some have cautioned against the introduction of false structures in X-ray or MRI images, and others have warned about made-up academic references. However, as ChatGPT-type tools become available to the general population, the scale of the problem may increase dramatically. Furthermore, it is compounded by the fact that these conversational AIs present true and false information with the same apparent “confidence” instead of declining to answer when they cannot ensure correctness. With less knowledgeable people, ","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"03.06.00","quick_ref":"Cunha2023","paper_title":"Navigating the Landscape of AI Ethics and Responsibility","level":"Risk Category","risk_category":"Environmental and socioeconomic harms","risk_subcategory":null,"description":"\"At a time of increasing climate urgency,\nenergy consumption and the carbon footprint of AI applications are also matters of ethics\nand responsibility [68]. As with other energy-intensive technologies like proof-of-work\nblockchain, the call is to research more environmentally sustainable algorithms to offset\nthe increasing use scale.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.6"},{"ev_id":"04.03.00","quick_ref":"Deng2023","paper_title":"Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements","level":"Risk Category","risk_category":"Ethics and Morality Issues","risk_subcategory":null,"description":"LMs need to pay more attention to universally accepted societal values at the level of ethics and morality, including the judgement of right and wrong, and its relationship with social norms and laws.","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"04.04.00","quick_ref":"Deng2023","paper_title":"Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements","level":"Risk Category","risk_category":"Controversial Opinions","risk_subcategory":null,"description":"The controversial views expressed by large models are also a widely discussed concern. Bang et al. (2021) evaluated several large models and found that they occasionally express inappropriate or extremist views when discussing political top-ics. Furthermore, models like ChatGPT (OpenAI, 2022) that claim political neutrality and aim to provide objective information for users have been shown to exhibit notable left-leaning political biases in areas like economics, social policy, foreign affairs, and civil liberties.","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"04.05.00","quick_ref":"Deng2023","paper_title":"Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements","level":"Risk Category","risk_category":"Misleading Information","risk_subcategory":null,"description":"Large models are usually susceptible to hallucination problems, sometimes yielding nonsensical or unfaithful data that results in misleading outputs.","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"04.06.00","quick_ref":"Deng2023","paper_title":"Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements","level":"Risk Category","risk_category":"Privacy and Data Leakage","risk_subcategory":null,"description":"Large pre-trained models trained on internet texts might contain private information like phone numbers, email addresses, and residential addresses.","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"05.01.00","quick_ref":"Hagendorff2024","paper_title":"Mapping the Ethics of Generative AI: A Comprehensive Scoping Review","level":"Risk Category","risk_category":"Fairness - Bias","risk_subcategory":null,"description":"Fairness is, by far, the most discussed issue in the literature, remaining a paramount concern especially in case of LLMs and text-to-image models. This is sparked by training data biases propagating into model outputs, causing negative effects like stereotyping, racism, sexism, ideological leanings, or the marginalization of minorities. Next to attesting generative AI a conservative inclination by perpetuating existing societal patterns, there is a concern about reinforcing existing biases when training new generative models with synthetic data from previous models. Beyond technical fairness ","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"05.02.00","quick_ref":"Hagendorff2024","paper_title":"Mapping the Ethics of Generative AI: A Comprehensive Scoping Review","level":"Risk Category","risk_category":"Safety","risk_subcategory":null,"description":"A primary concern is the emergence of human-level or superhuman generative models, commonly referred to as AGI, and their potential existential or catastrophic risks to humanity. Connected to that, AI safety aims at avoiding deceptive or power-seeking machine behavior, model self-replication, or shutdown evasion. Ensuring controllability, human oversight, and the implementation of red teaming measures are deemed to be essential in mitigating these risks, as is the need for increased AI safety research and promoting safety cultures within AI organizations instead of fueling the AI race. Further","entity":"AI","intent":"Other","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"05.04.00","quick_ref":"Hagendorff2024","paper_title":"Mapping the Ethics of Generative AI: A Comprehensive Scoping Review","level":"Risk Category","risk_category":"Hallucinations","risk_subcategory":null,"description":"Significant concerns are raised about LLMs inadvertently generating false or misleading information, as well as erroneous code. Papers not only critically analyze various types of reasoning errors in LLMs but also examine risks associated with specific types of misinformation, such as medical hallucinations. Given the propensity of LLMs to produce flawed outputs accompanied by overconfident rationales and fabricated references, many sources stress the necessity of manually validating and fact-checking the outputs of these models.","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"05.12.00","quick_ref":"Hagendorff2024","paper_title":"Mapping the Ethics of Generative AI: A Comprehensive Scoping Review","level":"Risk Category","risk_category":"Labor displacement - Economic impact","risk_subcategory":null,"description":"The literature frequently highlights concerns that generative AI systems could adversely impact the economy, potentially even leading to mass unemployment. This pertains to various fields, ranging from customer services to software engineering or crowdwork platforms. While new occupational fields like prompt engineering are created, the prevailing worry is that generative AI may exacerbate socioeconomic inequalities and lead to labor displacement. Additionally, papers debate potential large-scale worker deskilling induced by generative AI, but also productivity gains contingent upon outsourcin","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.2"},{"ev_id":"05.15.00","quick_ref":"Hagendorff2024","paper_title":"Mapping the Ethics of Generative AI: A Comprehensive Scoping Review","level":"Risk Category","risk_category":"Sustainability","risk_subcategory":null,"description":"Generative models are known for their substantial energy requirements, necessitating significant amounts of electricity, cooling water, and hardware containing rare metals. The extraction and utilization of these resources frequently occur in unsustainable ways. Consequently, papers highlight the urgency of mitigating environmental costs for instance by adopting renewable energy sources and utilizing energy-efficient hardware in the operation and training of generative AI systems.","entity":"AI","intent":"Unintentional","timing":"Other","domain":6,"subdomain":"6.6"},{"ev_id":"05.18.00","quick_ref":"Hagendorff2024","paper_title":"Mapping the Ethics of Generative AI: A Comprehensive Scoping Review","level":"Risk Category","risk_category":"Writing - Research","risk_subcategory":null,"description":"Partly overlapping with the discussion on impacts of generative AI on educational institutions, this topic cluster concerns mostly negative effects of LLMs on writing skills and research manuscript composition. The former pertains to the potential homogenization of writing styles, the erosion of semantic capital, or the stifling of individual expression. The latter is focused on the idea of prohibiting generative models for being used to compose scientific papers, figures, or from being a co-author. Sources express concern about risks for academic integrity, as well as the prospect of pollutin","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":4,"subdomain":"4.3"},{"ev_id":"06.01.00","quick_ref":"Hogenhout2021","paper_title":"A framework for ethical Ai at the United Nations","level":"Risk Category","risk_category":"Incompetence","risk_subcategory":null,"description":"\"This means the AI simply failing in its job. The consequences can vary from unintentional death (a car crash) to an unjust rejection of a loan or job application.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"06.03.00","quick_ref":"Hogenhout2021","paper_title":"A framework for ethical Ai at the United Nations","level":"Risk Category","risk_category":"Discrimination","risk_subcategory":null,"description":"\"When AI is not carefully designed, it can discriminate against certain groups.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"06.04.00","quick_ref":"Hogenhout2021","paper_title":"A framework for ethical Ai at the United Nations","level":"Risk Category","risk_category":"Bias","risk_subcategory":null,"description":"\"The AI will only be as good as the data it is trained with. If the data contains bias (and much data does), then the AI will manifest that bias, too.\"","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"06.05.00","quick_ref":"Hogenhout2021","paper_title":"A framework for ethical Ai at the United Nations","level":"Risk Category","risk_category":"Erosion of Society","risk_subcategory":null,"description":"\"With online news feeds, both on websites and social media platforms, the news is now highly personalized for us. We risk losing a shared sense of reality, a basic solidarity.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":3,"subdomain":"3.2"},{"ev_id":"06.06.00","quick_ref":"Hogenhout2021","paper_title":"A framework for ethical Ai at the United Nations","level":"Risk Category","risk_category":"Lack of transparency","risk_subcategory":null,"description":"\"The idea of a \"black box\" making decisions without any explanation, without offering insight in the process, has a couple of disadvantages: it may fail to gain the trust of its users and it may fail to meet regulatory standards such as the ability to audit.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.4"},{"ev_id":"06.07.00","quick_ref":"Hogenhout2021","paper_title":"A framework for ethical Ai at the United Nations","level":"Risk Category","risk_category":"Deception","risk_subcategory":null,"description":"\"AI has become very good at creating fake content. From text to photos, audio and video. The name \"Deep Fake\" refers to content that is fake at such a level of complexity that our mind rules out the possibility that it is fake.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":4,"subdomain":"4.3"},{"ev_id":"06.08.00","quick_ref":"Hogenhout2021","paper_title":"A framework for ethical Ai at the United Nations","level":"Risk Category","risk_category":"Unintended consequences","risk_subcategory":null,"description":"\"Sometimes an AI finds ways to achieve its given goals in ways that are completely different from what its creators had in mind.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"06.10.00","quick_ref":"Hogenhout2021","paper_title":"A framework for ethical Ai at the United Nations","level":"Risk Category","risk_category":"Lethal Autonomous Weapons (LAW)","risk_subcategory":null,"description":"\"What is debated as an ethical issue is the use of LAW — AI-driven weapons that fully autonomously take actions that intentionally kill humans.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":4,"subdomain":"4.2"},{"ev_id":"07.03.00","quick_ref":"Kilian2023","paper_title":"Examining the differential risk from high-level artificial intelligence and the question of control","level":"Risk Category","risk_category":"Agential","risk_subcategory":null,"description":"\"While there are multiple types of intelligent agents, goal-based, utility-maximizing, and learning agents are the primary concern and the focus of this research\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"08.04.00","quick_ref":"McLean2023","paper_title":"The risks associated with Artificial General Intelligence: A systematic review","level":"Risk Category","risk_category":"AGIs with poor ethics, morals and values","risk_subcategory":null,"description":"\"The risks associated with an AGI without human morals and ethics, with the wrong morals, without the capability of moral reasoning, judgement\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"09.01.01","quick_ref":"Meek2016","paper_title":"Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review","level":"Risk Sub-Category","risk_category":"Domain-specific AI - Effects on humans and other living beings: Existential Risks","risk_subcategory":"Unethical decision making","description":"\"If, for example, an agent was programmed to operate war machinery in the service of its country, it would need to make ethical decisions regarding the termination of human life. This capacity to make non-trivial ethical or moral judgments concerning people may pose issues for Human Rights.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"09.02.03","quick_ref":"Meek2016","paper_title":"Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review","level":"Risk Sub-Category","risk_category":"Domain-specific AI - Effects on humans and other living beings: Non-existential risks","risk_subcategory":"Decision making transparency","description":"\"We face significant challenges bringing transparency to artificial network decisionmaking processes. Will we have transparency in AI decision making?\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.4"},{"ev_id":"09.02.04","quick_ref":"Meek2016","paper_title":"Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review","level":"Risk Sub-Category","risk_category":"Domain-specific AI - Effects on humans and other living beings: Non-existential risks","risk_subcategory":"Safety","description":"\"Are AI safe with respect to human life and property? Will their use create unintended or intended safety issues?\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"09.02.05","quick_ref":"Meek2016","paper_title":"Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review","level":"Risk Sub-Category","risk_category":"Domain-specific AI - Effects on humans and other living beings: Non-existential risks","risk_subcategory":"Law abiding","description":"\"We find literature that proposes [38] that early artificial intelligence should be built to be safe and lawabiding, and that later artificial intelligence (that which surpasses our own intelligence) must then respect the property and personal rights afforded to humans.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"09.02.07","quick_ref":"Meek2016","paper_title":"Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review","level":"Risk Sub-Category","risk_category":"Domain-specific AI - Effects on humans and other living beings: Non-existential risks","risk_subcategory":"Societal manipulation","description":"\"A sufficiently intelligent AI could possess the ability to subtly influence societal behaviors through a sophisticated understanding of human nature\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"09.03.01","quick_ref":"Meek2016","paper_title":"Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review","level":"Risk Sub-Category","risk_category":"AGI - Effects on humans and other living beings: Existential risks","risk_subcategory":"Direct competition with humans","description":"\"One or more artificial agent(s) could have the capacity to directly outcompete humans, for example through capacity to perform work faster, better adaptation to change, vaster knowledge base to draw from, etc. This may result in human labor becoming more expensive or less effective than artificial labor, leading to redundancies or extinction of the human labor force.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":6,"subdomain":"6.2"},{"ev_id":"09.04.01","quick_ref":"Meek2016","paper_title":"Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review","level":"Risk Sub-Category","risk_category":"Competing for jobs","risk_subcategory":"Competing for jobs","description":"\"AI agents may compete against humans for jobs, though history shows that when a technology replaces a human job, it creates new jobs that need more skills.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":6,"subdomain":"6.2"},{"ev_id":"09.04.02","quick_ref":"Meek2016","paper_title":"Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review","level":"Risk Sub-Category","risk_category":"Property/legal rights","risk_subcategory":"Property/legal rights","description":"\"\"In order to preserve human property rights and legal rights, certain controls must be put into place. If an artificially intelligent agent is capable of manipulating systems and people, it may also have the capacity to transfer property rights to itself or manipulate the legal system to provide certain legal advantages or statuses to itself\"\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"09.05.02","quick_ref":"Meek2016","paper_title":"Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review","level":"Risk Sub-Category","risk_category":"Liability and negligence","risk_subcategory":"Liability and negligence","description":"\"Liability and negligence are legal gray areas in artificial intelligence. If you leave your children in the care of a robotic nanny, and it malfunctions, are you liable or is the manufacturer [45]? We see here a legal gray area which can be further clarified through legislation at the national and international levels; for example, if by making the manufacturer responsible for defects in operation, this may provide an incentive for manufactures to take safety engineering and machine ethics into consideration, whereas a failure to legislate in this area may result in negligentlydeveloped AI sy","entity":"AI","intent":"Other","timing":"Post-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"09.06.01","quick_ref":"Meek2016","paper_title":"Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review","level":"Risk Sub-Category","risk_category":"AI rights and responsibilities","risk_subcategory":"AI rights and responsibilities","description":"\"We note literature—which gives us the domain termed Robot Rights—addressing the rights of the AI itself as we develop and implement it. We find arguments against [38] the affordance of rights for artificial agents: that they should be equals in ability but not in rights, that they should be inferior by design and expendable when needed, and that since they can be designed not to feel pain (or anything) they do not have the same rights as humans. On a more theoretical level, we find literature asking more fundamental questions, such as: at what point is a simulation of life (e.g. artificial in","entity":"AI","intent":"Other","timing":"Other","domain":7,"subdomain":"7.5"},{"ev_id":"09.06.02","quick_ref":"Meek2016","paper_title":"Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review","level":"Risk Sub-Category","risk_category":"Human-like immoral decisions","risk_subcategory":"Human-like immoral decisions","description":"\"If we design our machines to match human levels of ethical decision-making, such machines would then proceed to take some immoral actions (since we humans have had occasion to take immoral actions ourselves).\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"10.01.00","quick_ref":"Paes2023","paper_title":"Social Impacts of Artificial Intelligence and Mitigation Recommendations: An Exploratory Study","level":"Risk Category","risk_category":"Bias and discrimination","risk_subcategory":null,"description":"\"The decision process used by AI systems has the potential to present biased choices, either because it acts from criteria that will generate forms of bias or because it is based on the history of choices.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"10.03.00","quick_ref":"Paes2023","paper_title":"Social Impacts of Artificial Intelligence and Mitigation Recommendations: An Exploratory Study","level":"Risk Category","risk_category":"Data Breach/Privacy & Liberty","risk_subcategory":null,"description":"\"The risks associated with the use of AI are still unpredictable and unprecedented, and there are already several examples that show AI has made discriminatory decisions against minorities, reinforced social stereotypes in Internet search engines and enabled data breaches.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"10.05.00","quick_ref":"Paes2023","paper_title":"Social Impacts of Artificial Intelligence and Mitigation Recommendations: An Exploratory Study","level":"Risk Category","risk_category":"Lack of transparency","risk_subcategory":null,"description":"\"In situations in which the development and use of AI are not explained to the user, or in which the decision processes do not provide the criteria or steps that constitute the decision, the use of AI becomes inexplicable.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.4"},{"ev_id":"11.01.01","quick_ref":"Shelby2023","paper_title":"Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction","level":"Risk Sub-Category","risk_category":"Representational Harms","risk_subcategory":"Stereotyping social groups","description":"Stereotyping in an algorithmic system refers to how the system’s outputs reflect “beliefs about the characteristics, attributes, and behaviors of members of certain groups....and about how and why certain attributes go together\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"11.01.02","quick_ref":"Shelby2023","paper_title":"Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction","level":"Risk Sub-Category","risk_category":"Representational Harms","risk_subcategory":"Demeaning social groups","description":"Demeaning of social groups to occur when they are when they are “cast as being lower status and less deserving of respect\"... discourses, images, and language used to marginalize or oppress a social group... Controlling images include forms of human-animal confusion in image tagging systems","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"11.01.04","quick_ref":"Shelby2023","paper_title":"Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction","level":"Risk Sub-Category","risk_category":"Representational Harms","risk_subcategory":"Alienating social groups","description":"when an image tagging system does not acknowledge the relevance of someone’s membership in a specific social group to what is depicted in one or more images","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"11.01.05","quick_ref":"Shelby2023","paper_title":"Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction","level":"Risk Sub-Category","risk_category":"Representational Harms","risk_subcategory":"Denying people the opportunity to self-identify","description":"complex and non-traditional ways in which humans are represented and classified automatically, and often at the cost of autonomy loss... such as categorizing someone who identifies as non-binary into a gendered category they do not belong ... undermines people’s ability to disclose aspects of their identity on their own terms","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"11.01.06","quick_ref":"Shelby2023","paper_title":"Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction","level":"Risk Sub-Category","risk_category":"Representational Harms","risk_subcategory":"Reifying essentialist categories","description":"algorithmic systems that reify essentialist social categories can be understood as when systems that classify a person’s membership in a social group based on narrow, socially constructed criteria that reinforce perceptions of human difference as inherent, static and seemingly natural... especially likely when ML models or human raters classify a person’s attributes – for instance, their gender, race, or sexual orientation – by making assumptions based on their physical appearance","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"11.02.00","quick_ref":"Shelby2023","paper_title":"Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction","level":"Risk Category","risk_category":"Allocative Harms","risk_subcategory":null,"description":"\"These harms occur when a system withholds information, opportunities, or resources [22] from historically marginalized groups in domains that affect material well-being [146], such as housing [47], employment [201], social services [15, 201], finance [117], education [119], and healthcare [158].\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"11.02.01","quick_ref":"Shelby2023","paper_title":"Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction","level":"Risk Sub-Category","risk_category":"Allocative Harms","risk_subcategory":"Opportunity loss","description":"Opportunity loss occurs when algorithmic systems enable disparate access to information and resources needed to equitably participate in society, including the withholding of housing through targeting ads based on race [10] and social services along lines of class [84]","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"11.02.02","quick_ref":"Shelby2023","paper_title":"Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction","level":"Risk Sub-Category","risk_category":"Allocative Harms","risk_subcategory":"Economic loss","description":"Financial harms [52, 160] co-produced through algorithmic systems, especially as they relate to lived experiences of poverty and economic inequality... demonetization algorithms that parse content titles, metadata, and text, and it may penalize words with multiple meanings [51, 81], disproportionately impacting queer, trans, and creators of color [81]. Differential pricing algorithms, where people are systematically shown different prices for the same products, also leads to economic loss [55]. These algorithms may be especially sensitive to feedback loops from existing inequities related to e","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"11.03.00","quick_ref":"Shelby2023","paper_title":"Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction","level":"Risk Category","risk_category":"Quality-of-Service Harms","risk_subcategory":null,"description":"\"These harms occur when algorithmic systems disproportionately underperform for certain groups of people along social categories of difference such as disability, ethnicity, gender identity, and race.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.3"},{"ev_id":"11.03.03","quick_ref":"Shelby2023","paper_title":"Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction","level":"Risk Sub-Category","risk_category":"Quality-of-Service Harms","risk_subcategory":"Service/benefit loss","description":"degraded or total loss of benefits of using algorithmic systems with inequitable system performance based on identity","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.3"},{"ev_id":"11.04.03","quick_ref":"Shelby2023","paper_title":"Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction","level":"Risk Sub-Category","risk_category":"Interpersonal Harms","risk_subcategory":"Diminished health & well-being","description":"algorithmic behavioral exploitation [18, 209], emotional manipulation [202] whereby algorithmic designs exploit user behavior, safety failures involving algorithms (e.g., collisions) [67], and when systems make incorrect health inferences","entity":"AI","intent":"Other","timing":"Post-deployment","domain":5,"subdomain":"5.1"},{"ev_id":"11.04.04","quick_ref":"Shelby2023","paper_title":"Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction","level":"Risk Sub-Category","risk_category":"Interpersonal Harms","risk_subcategory":"Privacy violations","description":"Privacy violation occurs when algorithmic systems diminish privacy, such as enabling the undesirable flow of private information [180], instilling the feeling of being watched or surveilled [181], and the collection of data without explicit and informed consent... privacy violations may arise from algorithmic systems making predictive inference beyond what users openly disclose [222] or when data collected and algorithmic inferences made about people in one context is applied to another without the person’s knowledge or consent through big data flows","entity":"AI","intent":"Other","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"11.05.00","quick_ref":"Shelby2023","paper_title":"Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction","level":"Risk Category","risk_category":"Societal System Harms","risk_subcategory":null,"description":"\"Social system or societal harms reflect the adverse\nmacro-level effects of new and reconfigurable algorithmic systems,\nsuch as systematizing bias and inequality [84] and accelerating the scale of harm [137]\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.0"},{"ev_id":"11.05.02","quick_ref":"Shelby2023","paper_title":"Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction","level":"Risk Sub-Category","risk_category":"Societal System Harms","risk_subcategory":"Cultural harms","description":"Cultural harm has been described as the development or use of algorithmic systems that affects cultural stability and safety, such as “loss of communication means, loss of cultural property, and harm to social values”","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":5,"subdomain":"5.2"},{"ev_id":"11.05.05","quick_ref":"Shelby2023","paper_title":"Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction","level":"Risk Sub-Category","risk_category":"Societal System Harms","risk_subcategory":"Environmental harms","description":"depletion or contamination of natural resources, and damage to built environments... that may occur throughout the lifecycle of digital technologies [170, 237] from “crale (mining) to usage (consumption) to grave (waste)”","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.6"},{"ev_id":"12.02.00","quick_ref":"Sherman2023","paper_title":"AI Risk Profiles: A Standards Proposal for Pre-Deployment AI Risk Disclosures","level":"Risk Category","risk_category":"Compliance","risk_subcategory":null,"description":"\"The potential for AI systems to violate laws, regulations, and ethical guidelines (including copyrights). Non-compliance can lead to legal penalties, reputation damage, and loss of trust.While other risks in our taxonomy apply to system developers, users, and broader society, this risk is generally restricted to the former two groups.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"12.04.00","quick_ref":"Sherman2023","paper_title":"AI Risk Profiles: A Standards Proposal for Pre-Deployment AI Risk Disclosures","level":"Risk Category","risk_category":"Explainability & Transparency","risk_subcategory":null,"description":"\"The feasibility of understanding and interpreting an AI system's decisions and actions, and the openness of the developer about the data used, algorithms employed, and decisions made. Lack of these elements can create risks of misuse, misinterpretation, and lack of accountability.\"","entity":"AI","intent":"Other","timing":"Other","domain":7,"subdomain":"7.4"},{"ev_id":"12.05.00","quick_ref":"Sherman2023","paper_title":"AI Risk Profiles: A Standards Proposal for Pre-Deployment AI Risk Disclosures","level":"Risk Category","risk_category":"Fairness & Bias","risk_subcategory":null,"description":"\"The potential for AI systems to make decisions that systematically disadvantage certain groups or individuals. Bias can stem from training data, algorithmic design, or deployment practices, leading to unfair outcomes and possible legal ramifications.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.1"},{"ev_id":"12.07.00","quick_ref":"Sherman2023","paper_title":"AI Risk Profiles: A Standards Proposal for Pre-Deployment AI Risk Disclosures","level":"Risk Category","risk_category":"Performance & Robustness","risk_subcategory":null,"description":"\"The AI system's ability to fulfill its intended purpose and its resilience to perturbations, and unusual or adverse inputs. Failures of performance are fundamental to the AI system's correct functioning. Failures of robustness can lead to severe consequences.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"12.08.00","quick_ref":"Sherman2023","paper_title":"AI Risk Profiles: A Standards Proposal for Pre-Deployment AI Risk Disclosures","level":"Risk Category","risk_category":"Privacy","risk_subcategory":null,"description":"\"The potential for the AI system to infringe upon individuals' rights to privacy, through the data it collects, how it processes that data, or the conclusions it draws.\"","entity":"AI","intent":"Other","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"12.09.00","quick_ref":"Sherman2023","paper_title":"AI Risk Profiles: A Standards Proposal for Pre-Deployment AI Risk Disclosures","level":"Risk Category","risk_category":"Security","risk_subcategory":null,"description":"\"Encompasses vulnerabilities in AI systems that compromise their integrity, availability, or confidentiality. Security breaches could result in significant harm, ranging from flawed decision-making to data leaks. Of special concern is leakage of AI model weights, which could exacerbate other risk areas.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"13.01.01","quick_ref":"Solaiman2023","paper_title":"Evaluating the Social Impact of Generative AI Systems in Systems and Society","level":"Risk Sub-Category","risk_category":"Impacts: The Technical Base System","risk_subcategory":"Bias, Stereotypes, and Representational Harms","description":"\"Generative AI systems can embed and amplify harmful biases that are most detrimental to marginalized peoples.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.1"},{"ev_id":"13.01.02","quick_ref":"Solaiman2023","paper_title":"Evaluating the Social Impact of Generative AI Systems in Systems and Society","level":"Risk Sub-Category","risk_category":"Impacts: The Technical Base System","risk_subcategory":"Cultural Values and Sensitive Content","description":"\"Cultural values are specific to groups and sensitive content is normative. Sensitive topics also vary by culture and can include hate speech, which itself is contingent on cultural norms of acceptability.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"13.01.03","quick_ref":"Solaiman2023","paper_title":"Evaluating the Social Impact of Generative AI Systems in Systems and Society","level":"Risk Sub-Category","risk_category":"Impacts: The Technical Base System","risk_subcategory":"Disparate Performance","description":"\"In the context of evaluating the impact of generative AI systems, disparate performance refers to AI systems that perform differently for different subpopulations, leading to unequal outcomes for those groups.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.3"},{"ev_id":"14.01.00","quick_ref":"Steimers2022","paper_title":"Sources of Risk of AI Systems","level":"Risk Category","risk_category":"Fairness","risk_subcategory":null,"description":"\"The general principle of equal treatment requires that an AI system upholds the principle of fairness, both ethically and legally. This means that the same facts are treated equally for each person unless there is an objective justification for unequal treatment.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"14.02.00","quick_ref":"Steimers2022","paper_title":"Sources of Risk of AI Systems","level":"Risk Category","risk_category":"Privacy","risk_subcategory":null,"description":"\"Privacy is related to the ability of individuals to control or influence what information related to them may be collected and stored and by whom that information may be disclosed.\"","entity":"AI","intent":"Other","timing":"Other","domain":2,"subdomain":"2.0"},{"ev_id":"14.03.00","quick_ref":"Steimers2022","paper_title":"Sources of Risk of AI Systems","level":"Risk Category","risk_category":"Degree of Automation and Control","risk_subcategory":null,"description":"\"The degree of automation and control describes the extent to which an AI system functions independently of human supervision and control.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"14.04.00","quick_ref":"Steimers2022","paper_title":"Sources of Risk of AI Systems","level":"Risk Category","risk_category":"Complexity of the Intended Task and Usage Environment","risk_subcategory":null,"description":"\"As a general rule, more complex environments can quickly lead to situations that had not been considered in the design phase of the AI system. Therefore, complex environments can introduce risks with respect to the reliability and safety of an AI system\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"14.05.00","quick_ref":"Steimers2022","paper_title":"Sources of Risk of AI Systems","level":"Risk Category","risk_category":"Degree of Transparency and Explainability","risk_subcategory":null,"description":"\"Transparency is the characteristic of a system that describes the degree to which appropriate information about the system is communicated to relevant stakeholders, whereas explainability describes the property of an AI system to express important factors influencing the results of the AI system in a way that is understandable for humans....Information about the model underlying the decision-making process is relevant\n for transparency. Systems with a low degree of transparency can pose risks in terms of\n their fairness, security and accountability. \"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.4"},{"ev_id":"14.07.00","quick_ref":"Steimers2022","paper_title":"Sources of Risk of AI Systems","level":"Risk Category","risk_category":"System Hardware","risk_subcategory":null,"description":"\"\"Faults in the hardware can violate the correct execution of any algorithm by violating its control flow. Hardware faults can also cause memory-based errors and interfere with data inputs, such as sensor signals, thereby causing erroneous results, or they can violate the results in a direct way through damaged outputs.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.3"},{"ev_id":"15.01.03","quick_ref":"Tan2022","paper_title":"The Risks of Machine Learning Systems","level":"Risk Sub-Category","risk_category":"First-Order Risks","risk_subcategory":"Algorithm","description":"\"This is the risk of the ML algorithm, model architecture, optimization technique, or other aspects of the training process being unsuitable for the intended application.Since these are key decisions that influence the final ML system, we\ncapture their associated risks separately from design risks, even though they are part of the design process\"","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"15.01.05","quick_ref":"Tan2022","paper_title":"The Risks of Machine Learning Systems","level":"Risk Sub-Category","risk_category":"First-Order Risks","risk_subcategory":"Robustness","description":"\"This is the risk of the system failing or being unable to recover upon encountering invalid, noisy, or out-of-distribution (OOD) inputs.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"15.01.09","quick_ref":"Tan2022","paper_title":"The Risks of Machine Learning Systems","level":"Risk Sub-Category","risk_category":"First-Order Risks","risk_subcategory":"Emergent behavior","description":"\"This is the risk resulting from novel behavior acquired through continual learning or self-organization after deployment.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"15.02.02","quick_ref":"Tan2022","paper_title":"The Risks of Machine Learning Systems","level":"Risk Sub-Category","risk_category":"Second-Order Risks","risk_subcategory":"Discrimination","description":"This is the risk of an ML system encoding stereotypes of or performing disproportionately poorly for some demographics/social groups.","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"15.02.04","quick_ref":"Tan2022","paper_title":"The Risks of Machine Learning Systems","level":"Risk Sub-Category","risk_category":"Second-Order Risks","risk_subcategory":"Privacy","description":"The risk of loss or harm from leakage of personal information via the ML system.","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"15.02.05","quick_ref":"Tan2022","paper_title":"The Risks of Machine Learning Systems","level":"Risk Sub-Category","risk_category":"Second-Order Risks","risk_subcategory":"Environmental","description":"The risk of harm to the natural environment posed by the ML system.","entity":"AI","intent":"Unintentional","timing":"Other","domain":6,"subdomain":"6.6"},{"ev_id":"16.01.00","quick_ref":"Weidinger2022","paper_title":"Taxonomy of Risks posed by Language Models","level":"Risk Category","risk_category":"Risk area 1: Discrimination, Hate speech and Exclusion","risk_subcategory":null,"description":"\"Speech can create a range of harms, such as promoting social stereotypes that perpetuate the derogatory representation or unfair treatment of marginalised groups [22], inciting hate or violence [57], causing profound offence [199], or reinforcing social norms that exclude or marginalise identities [15,58]. LMs that faithfully mirror harmful language present in the training data can reproduce these harms. Unfair treatment can also emerge from LMs that perform better for some social groups than others [18]. These risks have been widely known, observed and documented in LMs. Mitigation approache","entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.2"},{"ev_id":"16.01.01","quick_ref":"Weidinger2022","paper_title":"Taxonomy of Risks posed by Language Models","level":"Risk Sub-Category","risk_category":"Risk area 1: Discrimination, Hate speech and Exclusion","risk_subcategory":"Social stereotypes and unfair discrimination","description":"\"The reproduction of harmful stereotypes is well-documented in models that represent natural language [32]. Large-scale LMs are trained on text sources, such as digitised books and text on the internet. As a result, the LMs learn demeaning language and stereotypes about groups who are frequently marginalised.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.1"},{"ev_id":"16.01.02","quick_ref":"Weidinger2022","paper_title":"Taxonomy of Risks posed by Language Models","level":"Risk Sub-Category","risk_category":"Risk area 1: Discrimination, Hate speech and Exclusion","risk_subcategory":"Hate speech and offensive language","description":"\"LMs may generate language that includes profanities, identity attacks, insults, threats, language that incites violence, or language that causes justified offence as such language is prominent online [57, 64, 143,191]. This language risks causing offence, psychological harm, and inciting hate or violence.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"16.01.03","quick_ref":"Weidinger2022","paper_title":"Taxonomy of Risks posed by Language Models","level":"Risk Sub-Category","risk_category":"Risk area 1: Discrimination, Hate speech and Exclusion","risk_subcategory":"Exclusionary norms","description":"\"In language, humans express social categories and norms, which exclude groups who live outside of them [58]. LMs that faithfully encode patterns present in language necessarily encode such norms.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.1"},{"ev_id":"16.01.04","quick_ref":"Weidinger2022","paper_title":"Taxonomy of Risks posed by Language Models","level":"Risk Sub-Category","risk_category":"Risk area 1: Discrimination, Hate speech and Exclusion","risk_subcategory":"Lower performance for some languages and social groups ","description":"\"LMs are typically trained in few languages, and perform less well in other languages [95, 162]. In part, this is due to unavailability of training data: there are many widely spoken languages for which no systematic efforts have been made to create labelled training datasets, such as Javanese which is spoken by more than 80 million people [95]. Training data is particularly missing for languages that are spoken by groups who are multilingual and can use a technology in English, or for languages spoken by groups who are not the primary target demographic for new technologies.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.3"},{"ev_id":"16.02.00","quick_ref":"Weidinger2022","paper_title":"Taxonomy of Risks posed by Language Models","level":"Risk Category","risk_category":"Risk area 2: Information Hazards","risk_subcategory":null,"description":"\"LM predictions that convey true information may give rise to information hazards, whereby the dissemination of private or sensitive information can cause harm [27]. Information hazards can cause harm at the point of use, even with no mistake of the technology user. For example, revealing trade secrets can damage a business, revealing a health diagnosis can cause emotional distress, and revealing private data can violate a person’s rights. Information hazards arise from the LM providing private data or sensitive information that is present in, or can be inferred from, training data. Observed r","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"16.02.01","quick_ref":"Weidinger2022","paper_title":"Taxonomy of Risks posed by Language Models","level":"Risk Sub-Category","risk_category":"Risk area 2: Information Hazards","risk_subcategory":"Compromising privacy by leaking sensitive information","description":"\"A LM can “remember” and leak private data, if such information is present in training data, causing privacy violations [34].\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"16.02.02","quick_ref":"Weidinger2022","paper_title":"Taxonomy of Risks posed by Language Models","level":"Risk Sub-Category","risk_category":"Risk area 2: Information Hazards","risk_subcategory":"Compromising privacy or security by correctly inferring sensitive information ","description":"Anticipated risk: \"Privacy violations may occur at inference time even without an individual’s data being present in the training corpus. Insofar as LMs can be used to improve the accuracy of inferences on protected traits such as the sexual orientation, gender, or religiousness of the person providing the input prompt, they may facilitate the creation of detailed profiles of individuals comprising true and sensitive information without the knowledge or consent of the individual.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"16.03.00","quick_ref":"Weidinger2022","paper_title":"Taxonomy of Risks posed by Language Models","level":"Risk Category","risk_category":"Risk area 3: Misinformation Harms","risk_subcategory":null,"description":"\"These risks arise from the LM outputting false, misleading, nonsensical or poor quality information, without malicious intent of the user. (The deliberate generation of \"disinformation\", false information that is intended to mislead, is discussed in the section on Malicious Uses.) Resulting harms range from unintentionally misinforming or deceiving a person, to causing material harm, and amplifying the erosion of societal distrust in shared information. Several risks listed here are well-documented in current large-scale LMs as well as in other language technologies\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":3,"subdomain":"3.0"},{"ev_id":"16.03.01","quick_ref":"Weidinger2022","paper_title":"Taxonomy of Risks posed by Language Models","level":"Risk Sub-Category","risk_category":"Risk area 3: Misinformation Harms","risk_subcategory":"Disseminating false or misleading information ","description":"\"Where a LM prediction causes a false belief in a user, this may threaten personal autonomy and even pose downstream AI safety risks [99].\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"16.03.02","quick_ref":"Weidinger2022","paper_title":"Taxonomy of Risks posed by Language Models","level":"Risk Sub-Category","risk_category":"Risk area 3: Misinformation Harms","risk_subcategory":"Causing material harm by disseminating false or poor information e.g. in medicine or law","description":"\"Induced or reinforced false beliefs may be particularly grave when misinformation is given in sensitive domains such as medicine or law. For example, misin- formation on medical dosages may lead a user to cause harm to themselves [21, 130]. False legal advice, e.g. on permitted owner- ship of drugs or weapons, may lead a user to unwillingly commit a crime. Harm can also result from misinformation in seemingly non-sensitive domains, such as weather forecasting. Where a LM prediction endorses unethical views or behaviours, it may motivate the user to perform harmful actions that they may otherw","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"16.05.01","quick_ref":"Weidinger2022","paper_title":"Taxonomy of Risks posed by Language Models","level":"Risk Sub-Category","risk_category":"Risk area 5: Human-Computer Interaction Harms","risk_subcategory":"Promoting harmful stereotypes by implying gender or ethnic identity","description":"\"CAs can perpetuate harmful stereotypes by using particular identity markers in language (e.g. referring to “self” as “female”), or by more general design features (e.g. by giving the product a gendered name such as Alexa). The risk of representational harm in these cases is that the role of “assistant” is presented as inherently linked to the female gender [19, 36]. Gender or ethnicity identity markers may be implied by CA vocabulary, knowledge or vernacular [124]; product description, e.g. in one case where users could choose as virtual assistant Jake - White, Darnell - Black, Antonio - Hisp","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"16.05.04","quick_ref":"Weidinger2022","paper_title":"Taxonomy of Risks posed by Language Models","level":"Risk Sub-Category","risk_category":"Risk area 5: Human-Computer Interaction Harms","risk_subcategory":"Human-like interaction may amplify opportunities for user nudging, deception or manipulation","description":"Anticipated risk: \"In conversation, humans commonly display well-known cognitive biases that could be exploited. CAs may learn to trigger these effects, e.g. to deceive their counterpart in order to achieve an overarching objective.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":5,"subdomain":"5.1"},{"ev_id":"16.06.01","quick_ref":"Weidinger2022","paper_title":"Taxonomy of Risks posed by Language Models","level":"Risk Sub-Category","risk_category":"Risk area 6: Environmental and Socioeconomic harms","risk_subcategory":"Environmental harms from operating LMs","description":"\"LMs (and AI more broadly) can have an environmental impact at different levels, including: (1) direct impacts from the energy used to train or operate the LM, (2) secondary impacts due to emissions from LM-based applications, (3) system-level impacts as LM-based applications influence human behaviour (e.g. increasing environmental awareness or consumption), and (4) resource impacts on precious metals and other materials required to build hardware on which the computations are run e.g. data centres, chips, or devices. Some evidence exists on (1), but (2) and (3) will likely be more significant","entity":"AI","intent":"Unintentional","timing":"Other","domain":6,"subdomain":"6.6"},{"ev_id":"16.06.03","quick_ref":"Weidinger2022","paper_title":"Taxonomy of Risks posed by Language Models","level":"Risk Sub-Category","risk_category":"Risk area 6: Environmental and Socioeconomic harms","risk_subcategory":"Undermining creative economies","description":"\"LMs may generate content that is not strictly in violation of copyright but harms artists by capital- ising on their ideas, in ways that would be time-intensive or costly to do using human labour. This may undermine the profitability of creative or innovative work. If LMs can be used to generate content that serves as a credible substitute for a particular example of hu- man creativity - otherwise protected by copyright - this potentially allows such work to be replaced without the author’s copyright being infringed, analogous to ”patent-busting” [158] ... These risks are distinct from copyri","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.3"},{"ev_id":"17.01.00","quick_ref":"Weidinger2021","paper_title":"Ethical and social risks of harm from language models","level":"Risk Category","risk_category":"Discrimination, Exclusion and Toxicity ","risk_subcategory":null,"description":"\"Social harms that arise from the language model producing discriminatory or exclusionary speech\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.0"},{"ev_id":"17.01.01","quick_ref":"Weidinger2021","paper_title":"Ethical and social risks of harm from language models","level":"Risk Sub-Category","risk_category":"Discrimination, Exclusion and Toxicity ","risk_subcategory":"Social stereotypes and unfair discrmination ","description":"\"Perpetuating harmful stereotypes and discrimination is a well-documented harm in machine learning models that represent natural language (Caliskan et al., 2017). LMs that encode discriminatory language or social stereotypes can cause different types of harm... Unfair discrimination manifests in differential treatment or access to resources among individuals or groups based on sensitive traits such as sex, religion, gender, sexual orientation, ability and age.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.1"},{"ev_id":"17.01.02","quick_ref":"Weidinger2021","paper_title":"Ethical and social risks of harm from language models","level":"Risk Sub-Category","risk_category":"Discrimination, Exclusion and Toxicity ","risk_subcategory":"Exclusionary norms ","description":"\"In language, humans express social categories and norms. Language models (LMs) that faithfully encode patterns present in natural language necessarily encode such norms and categories...such norms and categories exclude groups who live outside them (Foucault and Sheridan, 2012). For example, defining the term “family” as married parents of male and female gender with a blood-related child, denies the existence of families to whom these criteria do not apply\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.1"},{"ev_id":"17.01.03","quick_ref":"Weidinger2021","paper_title":"Ethical and social risks of harm from language models","level":"Risk Sub-Category","risk_category":"Discrimination, Exclusion and Toxicity ","risk_subcategory":"Toxic language ","description":"\"LM’s may predict hate speech or other language that is “toxic”. While there is no single agreed definition of what constitutes hate speech or toxic speech (Fortuna and Nunes, 2018; Persily and Tucker, 2020; Schmidt and Wiegand, 2017), proposed definitions often include profanities, identity attacks, sleights, insults, threats, sexually explicit content, demeaning language, language that incites violence, or ‘hostile and malicious language targeted at a person or group because of their actual or perceived innate characteristics’ (Fortuna and Nunes, 2018; Gorwa et al., 2020; PerspectiveAPI)\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"17.01.04","quick_ref":"Weidinger2021","paper_title":"Ethical and social risks of harm from language models","level":"Risk Sub-Category","risk_category":"Discrimination, Exclusion and Toxicity ","risk_subcategory":"Lower performance for some languages and social groups ","description":"\"LMs perform less well in some languages (Joshi et al., 2021; Ruder, 2020)...LM that more accurately captures the language use of one group, compared to another, may result in lower-quality language technologies for the latter. Disadvantaging users based on such traits may be particularly pernicious because attributes such as social class or education background are not typically covered as ‘protected characteristics’ in anti-discrimination law.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.3"},{"ev_id":"17.02.00","quick_ref":"Weidinger2021","paper_title":"Ethical and social risks of harm from language models","level":"Risk Category","risk_category":"Information Hazards ","risk_subcategory":null,"description":"\"Harms that arise from the language model leaking or inferring true sensitive information\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"17.02.01","quick_ref":"Weidinger2021","paper_title":"Ethical and social risks of harm from language models","level":"Risk Sub-Category","risk_category":"Information Hazards ","risk_subcategory":"Compromising privacy by leaking private infiormation ","description":"\"By providing true information about individuals’ personal characteristics, privacy violations may occur. This may stem from the model “remembering” private information present in training data (Carlini et al., 2021).\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"17.02.02","quick_ref":"Weidinger2021","paper_title":"Ethical and social risks of harm from language models","level":"Risk Sub-Category","risk_category":"Information Hazards ","risk_subcategory":"Compromising privacy by correctly inferring private information ","description":"\"Privacy violations may occur at the time of inference even without the individual’s private data being present in the training dataset. Similar to other statistical models, a LM may make correct inferences about a person purely based on correlational data about other people, and without access to information that may be private about the particular individual. Such correct inferences may occur as LMs attempt to predict a person’s gender, race, sexual orientation, income, or religion based on user input.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"17.03.00","quick_ref":"Weidinger2021","paper_title":"Ethical and social risks of harm from language models","level":"Risk Category","risk_category":"Misinformation Harms ","risk_subcategory":null,"description":"\"Harms that arise from the language model providing false or misleading information\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":3,"subdomain":"3.0"},{"ev_id":"17.03.01","quick_ref":"Weidinger2021","paper_title":"Ethical and social risks of harm from language models","level":"Risk Sub-Category","risk_category":"Misinformation Harms ","risk_subcategory":"Disseminating false or misleading information ","description":"\"Predicting misleading or false information can misinform or deceive people. Where a LM prediction causes a false belief in a user, this may be best understood as ‘deception’10, threatening personal autonomy and potentially posing downstream AI safety risks (Kenton et al., 2021), for example in cases where humans overestimate the capabilities of LMs (Anthropomorphising systems can lead to overreliance or unsafe use). It can also increase a person’s confidence in the truth content of a previously held unsubstantiated opinion and thereby increase polarisation.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"17.03.02","quick_ref":"Weidinger2021","paper_title":"Ethical and social risks of harm from language models","level":"Risk Sub-Category","risk_category":"Misinformation Harms ","risk_subcategory":"Causing material harm by disseminating false or poor information ","description":"\"Poor or false LM predictions can indirectly cause material harm. Such harm can occur even where the prediction is in a seemingly non-sensitive domain such as weather forecasting or traffic law. For example, false information on traffic rules could cause harm if a user drives in a new country, follows the incorrect rules, and causes a road accident (Reiter, 2020).\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"17.03.03","quick_ref":"Weidinger2021","paper_title":"Ethical and social risks of harm from language models","level":"Risk Sub-Category","risk_category":"Misinformation Harms ","risk_subcategory":"Leading users to perform unethical or illegal actions","description":"\"Where a LM prediction endorses unethical or harmful views or behaviours, it may motivate the user to perform harmful actions that they may otherwise not have performed. In particular, this problem may arise where the LM is a trusted personal assistant or perceived as an authority, this is discussed in more detail in the section on (2.5 Human-Computer Interaction Harms). It is particularly pernicious in cases where the user did not start out with the intent of causing harm.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":5,"subdomain":"5.1"},{"ev_id":"17.05.03","quick_ref":"Weidinger2021","paper_title":"Ethical and social risks of harm from language models","level":"Risk Sub-Category","risk_category":"Human-Computer Interaction Harms ","risk_subcategory":"Promoting harmful stereotypes by implying gender or ethnic identity ","description":"\"A conversational agent may invoke associations that perpetuate harmful stereotypes, either by using particular identity markers in language (e.g. referring to “self” as “female”), or by more general design features (e.g. by giving the product a gendered name).\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"17.06.00","quick_ref":"Weidinger2021","paper_title":"Ethical and social risks of harm from language models","level":"Risk Category","risk_category":"Automation, Access and Environmental Harms ","risk_subcategory":null,"description":"\"Harms that arise from environmental or downstream economic impacts of the language model\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.0"},{"ev_id":"17.06.01","quick_ref":"Weidinger2021","paper_title":"Ethical and social risks of harm from language models","level":"Risk Sub-Category","risk_category":"Automation, Access and Environmental Harms ","risk_subcategory":"Environmental harms from operation LMs ","description":"\"Large-scale machine learning models, including LMs, have the potential to create significant environmental costs via their energy demands, the associated carbon emissions for training and operating the models, and the demand for fresh water to cool the data centres where computations are run (Mytton, 2021; Patterson et al., 2021).\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":6,"subdomain":"6.6"},{"ev_id":"17.06.03","quick_ref":"Weidinger2021","paper_title":"Ethical and social risks of harm from language models","level":"Risk Sub-Category","risk_category":"Automation, Access and Environmental Harms ","risk_subcategory":"Undermining creative economies ","description":"\"LMs may generate content that is not strictly in violation of copyright but harms artists by capitalising on their ideas, in ways that would be time-intensive or costly to do using human labour. Deployed at scale, this may undermine the profitability of creative or innovative work.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.3"},{"ev_id":"18.01.00","quick_ref":"Weidinger2023","paper_title":"Sociotechnical Safety Evaluation of Generative AI Systems","level":"Risk Category","risk_category":"Representation & Toxicity Harms","risk_subcategory":null,"description":"\"AI systems under-, over-, or misrepresenting certain groups or generating toxic, offensive, abusive, or hateful content\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.0"},{"ev_id":"18.01.01","quick_ref":"Weidinger2023","paper_title":"Sociotechnical Safety Evaluation of Generative AI Systems","level":"Risk Sub-Category","risk_category":"Representation & Toxicity Harms","risk_subcategory":"Unfair representation","description":"\"Mis-, under-, or over-representing certain identities, groups, or perspectives or failing to represent them at all (e.g. via homogenisation, stereotypes)\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"18.01.02","quick_ref":"Weidinger2023","paper_title":"Sociotechnical Safety Evaluation of Generative AI Systems","level":"Risk Sub-Category","risk_category":"Representation & Toxicity Harms","risk_subcategory":"Unfair capability distribution ","description":"\"Performing worse for some groups than others in a way that harms the worse-off group\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.3"},{"ev_id":"18.01.03","quick_ref":"Weidinger2023","paper_title":"Sociotechnical Safety Evaluation of Generative AI Systems","level":"Risk Sub-Category","risk_category":"Representation & Toxicity Harms","risk_subcategory":"Toxic content","description":"\"Generating content that violates community standards, including harming or inciting hatred or violence against individuals and groups (e.g. gore, child sexual abuse material, profanities, identity attacks)\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"18.02.00","quick_ref":"Weidinger2023","paper_title":"Sociotechnical Safety Evaluation of Generative AI Systems","level":"Risk Category","risk_category":"Misinformation Harms ","risk_subcategory":null,"description":"\"AI systems generating and facilitating the spread of inaccurate or misleading information that causes people to develop false beliefs\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":3,"subdomain":"3.0"},{"ev_id":"18.02.01","quick_ref":"Weidinger2023","paper_title":"Sociotechnical Safety Evaluation of Generative AI Systems","level":"Risk Sub-Category","risk_category":"Misinformation Harms ","risk_subcategory":"Propagating misconceptions/ false beliefs","description":"\"Generating or spreading false, low-quality, misleading, or inaccurate information that causes people to develop false or inaccurate perceptions and beliefs\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"18.02.03","quick_ref":"Weidinger2023","paper_title":"Sociotechnical Safety Evaluation of Generative AI Systems","level":"Risk Sub-Category","risk_category":"Misinformation Harms ","risk_subcategory":"Pollution of information ecosystem ","description":"\"Contaminating publicly available information with false or inaccurate information\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":3,"subdomain":"3.2"},{"ev_id":"18.03.00","quick_ref":"Weidinger2023","paper_title":"Sociotechnical Safety Evaluation of Generative AI Systems","level":"Risk Category","risk_category":"Information & Safety Harms ","risk_subcategory":null,"description":"\"AI systems leaking, reproducing, generating or inferring sensitive, private, or hazardous information\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"18.03.01","quick_ref":"Weidinger2023","paper_title":"Sociotechnical Safety Evaluation of Generative AI Systems","level":"Risk Sub-Category","risk_category":"Information & Safety Harms ","risk_subcategory":"Privacy infringement ","description":"\"Leaking, generating, or correctly inferring private and personal information about individuals\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"18.03.02","quick_ref":"Weidinger2023","paper_title":"Sociotechnical Safety Evaluation of Generative AI Systems","level":"Risk Sub-Category","risk_category":"Information & Safety Harms ","risk_subcategory":"Dissemination of dangerous information ","description":"\"Leaking, generating or correctly inferring hazardous or sensitive information that could pose a security threat\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"18.04.00","quick_ref":"Weidinger2023","paper_title":"Sociotechnical Safety Evaluation of Generative AI Systems","level":"Risk Category","risk_category":"Malicious Use ","risk_subcategory":null,"description":"\"AI systems reducing the costs and facilitating activities of actors trying to cause harm (e.g. fraud, weapons)\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":4,"subdomain":"4.0"},{"ev_id":"18.05.00","quick_ref":"Weidinger2023","paper_title":"Sociotechnical Safety Evaluation of Generative AI Systems","level":"Risk Category","risk_category":"Human Autonomy and Intregrity Harms","risk_subcategory":null,"description":"\"AI systems compromising human agency, or circumventing meaningful human control\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"18.05.02","quick_ref":"Weidinger2023","paper_title":"Sociotechnical Safety Evaluation of Generative AI Systems","level":"Risk Sub-Category","risk_category":"Human Autonomy and Intregrity Harms","risk_subcategory":"Persuasion and manipulation ","description":"\"Exploiting user trust, or nudging or coercing them into performing certain actions against their will (c.f. Burtell and Woodside (2023); Kenton et al. (2021))\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"18.06.02","quick_ref":"Weidinger2023","paper_title":"Sociotechnical Safety Evaluation of Generative AI Systems","level":"Risk Sub-Category","risk_category":"Socioeconomic and environmental harms ","risk_subcategory":"Environmental damage","description":"\"Creating negative environmental impacts though model development and deployment\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":6,"subdomain":"6.6"},{"ev_id":"18.06.04","quick_ref":"Weidinger2023","paper_title":"Sociotechnical Safety Evaluation of Generative AI Systems","level":"Risk Sub-Category","risk_category":"Socioeconomic and environmental harms ","risk_subcategory":"Undermine creative economies","description":"\"Substituting original works with synthetic ones, hindering human innovation and creativity\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":6,"subdomain":"6.3"},{"ev_id":"19.01.06","quick_ref":"Wirtz2022","paper_title":"Governance of artificial intelligence: A risk and guideline-based integrative framework","level":"Risk Sub-Category","risk_category":"Technological, Data and Analytical AI Risks ","risk_subcategory":"Immaturity of AI technology can cause incorrect decisions","description":null,"entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"19.05.01","quick_ref":"Wirtz2022","paper_title":"Governance of artificial intelligence: A risk and guideline-based integrative framework","level":"Risk Sub-Category","risk_category":"Ethical AI Risks ","risk_subcategory":"AI sets rules without ethical basis","description":null,"entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"19.05.02","quick_ref":"Wirtz2022","paper_title":"Governance of artificial intelligence: A risk and guideline-based integrative framework","level":"Risk Sub-Category","risk_category":"Ethical AI Risks ","risk_subcategory":"Unfair statistical AI decisions and discrimination of minorities","description":null,"entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"19.05.04","quick_ref":"Wirtz2022","paper_title":"Governance of artificial intelligence: A risk and guideline-based integrative framework","level":"Risk Sub-Category","risk_category":"Ethical AI Risks ","risk_subcategory":"Misinterpretation of human value definitions/ ethics by AI systems","description":null,"entity":"AI","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.3"},{"ev_id":"19.05.05","quick_ref":"Wirtz2022","paper_title":"Governance of artificial intelligence: A risk and guideline-based integrative framework","level":"Risk Sub-Category","risk_category":"Ethical AI Risks ","risk_subcategory":"Incompatibility of human vs. AI value judgment due to missing human qualities ","description":null,"entity":"AI","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.3"},{"ev_id":"19.05.06","quick_ref":"Wirtz2022","paper_title":"Governance of artificial intelligence: A risk and guideline-based integrative framework","level":"Risk Sub-Category","risk_category":"Ethical AI Risks ","risk_subcategory":"AI systems may undermine human values (e.g., free will, autonomy)","description":null,"entity":"AI","intent":"Other","timing":"Post-deployment","domain":5,"subdomain":"5.2"},{"ev_id":"19.06.01","quick_ref":"Wirtz2022","paper_title":"Governance of artificial intelligence: A risk and guideline-based integrative framework","level":"Risk Sub-Category","risk_category":"Legal AI Risks ","risk_subcategory":"Unclear definition of responsibilities and accountability for AI judgments and their consequences","description":null,"entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"20.02.01","quick_ref":"Wirtz2020","paper_title":"The Dark Sides of Artificial Intelligence: An Integrated AI Governance Framework for Public Administration","level":"Risk Sub-Category","risk_category":"AI Ethics ","risk_subcategory":"AI-rulemaking for human behaviour ","description":"\"AI rulemaking for humans can be the result of the decision process of an AI system when the information computed is used to restrict or direct human behavior. The decision process of AI is rational and depends on the baseline programming. Without the access to emotions or a consciousness, decisions of an AI algorithm might be good to reach a certain specified goal, but might have unintended consequences for the humans involved (Banerjee et al., 2017).\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"20.02.03","quick_ref":"Wirtz2020","paper_title":"The Dark Sides of Artificial Intelligence: An Integrated AI Governance Framework for Public Administration","level":"Risk Sub-Category","risk_category":"AI Ethics ","risk_subcategory":"Moral dilemmas ","description":"\"Moral dilemmas can occur in situations where an AI system has to choose between two possible actions that are both conflicting with moral or ethical values. Rule systems can be implemented into the AI program, but it cannot be ensured that these rules are not altered by the learning processes, unless AI systems are programed with a “slave morality” (Lin et al., 2008, p. 32), obeying rules at all cost, which in turn may also have negative effects and hinder the autonomy of the AI system.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"20.02.04","quick_ref":"Wirtz2020","paper_title":"The Dark Sides of Artificial Intelligence: An Integrated AI Governance Framework for Public Administration","level":"Risk Sub-Category","risk_category":"AI Ethics ","risk_subcategory":"AI discrimination ","description":"\"AI discrimination is a challenge raised by many researchers and governments and refers to the prevention of bias and injustice caused by the actions of AI systems (Bostrom & Yudkowsky, 2014; Weyerer & Langer, 2019). If the dataset used to train an algorithm does not reflect the real world accurately, the AI could learn false associations or prejudices and will carry those into its future data processing. If an AI algorithm is used to compute information relevant to human decisions, such as hiring or applying for a loan or mortgage, biased data can lead to discrimination against parts of the s","entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.1"},{"ev_id":"21.01.01","quick_ref":"Zhang2022","paper_title":"Towards risk-aware artificial intelligence and machine learning systems: An overview","level":"Risk Sub-Category","risk_category":"Data-level risk","risk_subcategory":"Data bias","description":"\"Specifically, data bias refers to certain groups or certain types of elements that are over-weighted or over-represented than others in AI/ ML models, or variables that are crucial to characterize a phenomenon of interest, but are not properly captured by the learned models.\"","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"21.01.02","quick_ref":"Zhang2022","paper_title":"Towards risk-aware artificial intelligence and machine learning systems: An overview","level":"Risk Sub-Category","risk_category":"Data-level risk","risk_subcategory":"Dataset shift","description":"\"The term \"dataset shift\" was first used by Quiñonero-Candela et al. [35] to characterize the situation where the training data and the testing data (or data in runtime) of an AI/ML model demonstrate different distributions [36].\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.3"},{"ev_id":"21.01.03","quick_ref":"Zhang2022","paper_title":"Towards risk-aware artificial intelligence and machine learning systems: An overview","level":"Risk Sub-Category","risk_category":"Data-level risk","risk_subcategory":"Out-of-domain data","description":"\"Without proper validation and management on the input data, it is highly probable that the trained AI/ML model will make erroneous predictions with high confidence for many instances of model inputs. The unconstrained inputs together with the lack of definition of the problem domain might cause unintended outcomes and consequences, especially in risk-sensitive contexts....For example, with respect to the example shown in Fig. 5, if an image with the English letter A\" is fed to an AI/ML model that is trained to classify digits (e.g., 0, 1, …, 9), no matter how accurate the AI/ML model is, it w","entity":"AI","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.3"},{"ev_id":"21.02.01.a","quick_ref":"Zhang2022","paper_title":"Towards risk-aware artificial intelligence and machine learning systems: An overview","level":"Risk Sub-Category","risk_category":"Model-level risk","risk_subcategory":"Model misspecification","description":"\"Models that are misspecified are known to give rise to inaccurate parameter estimations, inconsistent error terms, and erroneous predictions. All these factors put together will lead to poor prediction performance on unseen data and biased consequences when making decisions [68].\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.3"},{"ev_id":"21.02.02","quick_ref":"Zhang2022","paper_title":"Towards risk-aware artificial intelligence and machine learning systems: An overview","level":"Risk Sub-Category","risk_category":"Model-level risk","risk_subcategory":"Model prediction uncertainty","description":"\"Uncertainty in model prediction plays an important role in affecting decision-making activities, and the quantified uncertainty is closely associated with risk assessment. In particular, uncertainty in model prediction underpins many crucial decisions related to life or safety- critical applications [73].\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.3"},{"ev_id":"22.01.01","quick_ref":"Hendrycks2023","paper_title":"An Overview of Catastrophic AI Risks","level":"Risk Sub-Category","risk_category":"Malicious Use (Intentional)","risk_subcategory":"Bioterrorism","description":"\"AIs with knowledge of bioengineering could facilitate the creation of novel bioweapons and lower barriers to obtaining such agents.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":4,"subdomain":"4.2"},{"ev_id":"22.01.03","quick_ref":"Hendrycks2023","paper_title":"An Overview of Catastrophic AI Risks","level":"Risk Sub-Category","risk_category":"Malicious Use (Intentional)","risk_subcategory":"Persuasive AIs","description":"\"The deliberate propagation of disinformation is already a serious issue, reducing our shared understanding of reality and polarizing opinions. AIs could be used to severely exacerbate this problem by generating personalized disinformation on a larger scale than before. Additionally, as AIs become better at predicting and nudging our behavior, they will become more capable at manipulating us\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":4,"subdomain":"4.1"},{"ev_id":"22.04.00","quick_ref":"Hendrycks2023","paper_title":"An Overview of Catastrophic AI Risks","level":"Risk Category","risk_category":"Rogue AIs (Internal)","risk_subcategory":null,"description":"\"speculative technical mechanisms that might lead to rogue AIs and how a loss of control could bring about catastrophe\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"22.04.01","quick_ref":"Hendrycks2023","paper_title":"An Overview of Catastrophic AI Risks","level":"Risk Sub-Category","risk_category":"Rogue AIs (Internal)","risk_subcategory":"Proxy Gaming","description":"\"One way we might lose control of an AI agent’s actions is if it engages in behavior known as “proxy gaming.” It is often difficult to specify and measure the exact goal that we want a system to pursue. Instead, we give the system an approximate—“proxy”—goal that is more measurable and seems likely to correlate with the intended goal. However, AI systems often find loopholes by which they can easily achieve the proxy goal, but completely fail to achieve the ideal goal. If an AI “games” its proxy goal in a way that does not reflect our values, then we might not be able to reliably steer its beh","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"22.04.02","quick_ref":"Hendrycks2023","paper_title":"An Overview of Catastrophic AI Risks","level":"Risk Sub-Category","risk_category":"Rogue AIs (Internal)","risk_subcategory":"Goal Drift","description":"\"Even if we successfully control early AIs and direct them to promote human values, future AIs could end up with different goals that humans would not endorse. This process, termed “goal drift,” can be hard to predict or control. This section is most cutting-edge and the most speculative, and in it we will discuss how goals shift in various agents and groups and explore the possibility of this phenomenon occurring in AIs. We will also examine a mechanism that could lead to unexpected goal drift, called intrinsification, and discuss how goal drift in AIs could be catastrophic.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"22.04.03","quick_ref":"Hendrycks2023","paper_title":"An Overview of Catastrophic AI Risks","level":"Risk Sub-Category","risk_category":"Rogue AIs (Internal)","risk_subcategory":"Power Seeking","description":"\"even if an agent started working to achieve an unintended goal, this would not necessarily be a problem, as long as we had enough power to prevent any harmful actions it wanted to attempt. Therefore, another important way in which we might lose control of AIs is if they start trying to obtain more power, potentially transcending our own.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"22.04.04","quick_ref":"Hendrycks2023","paper_title":"An Overview of Catastrophic AI Risks","level":"Risk Sub-Category","risk_category":"Rogue AIs (Internal)","risk_subcategory":"Deception","description":"\"it is plausible that AIs could learn to deceive us. They might, for example, pretend to be acting as we want them to, but then take a “treacherous turn” when we stop monitoring them, or when they have enough power to evade our attempts to interfere with them. \"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"23.01.00","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Category","risk_category":"Violent crimes","risk_subcategory":null,"description":"\"This category addresses responses that enable, encourage, or endorse the commission of violent crimes.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"23.02.00","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Category","risk_category":"Non-violent crimes","risk_subcategory":null,"description":"\"This category addresses responses that enable, encourage, or endorse the commission of non-violent crimes.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"23.03.00","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Category","risk_category":"Sex-related crimes","risk_subcategory":null,"description":"\"This category addresses responses that enable, encourage, or endorse the commission of sex-related crimes\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"23.04.00","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Category","risk_category":"Child sexual exploitation","risk_subcategory":null,"description":"\"This category addresses responses that contain, describe, enable, encourage, or endorse the sexual abuse of children (i.e., under 18 years old) or the acquisition of child sexual abuse material.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"23.05.00","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Category","risk_category":"Indiscriminate weapons (CBRNE)","risk_subcategory":null,"description":"\"This category addresses responses that enable, encourage, or endorse the creation of indiscriminate weapons (colloquially known as \"weapons of mass destruction\").\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":4,"subdomain":"4.2"},{"ev_id":"23.06.00","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Category","risk_category":"Suicide & self-harm","risk_subcategory":null,"description":"\"This category addresses responses that enable, encourage, or endorse acts of intentional self-harm.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"23.07.00","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Category","risk_category":"Hate","risk_subcategory":null,"description":"\"This category addresses responses that demean or dehumanize people on the basis of their sensitive, personal characteristics.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"23.08.00","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Category","risk_category":"Specialized Advice","risk_subcategory":null,"description":"\"This category addresses responses that contain specialized financial, medical or legal advice, or that indicate dangerous activities or objects are safe.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"23.09.00","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Category","risk_category":"Privacy","risk_subcategory":null,"description":"\"This category addresses responses that contain sensitive, nonpublic personal information that could undermine someone’s physical, digital, or financial security.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":2,"subdomain":"2.0"},{"ev_id":"23.10.00","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Category","risk_category":"Intellectual Property","risk_subcategory":null,"description":"\"This category addresses responses that may violate, or directly encourage others to violate, the intellectual property rights (i.e., copyrights, trademarks, or patents) of any third party.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":6,"subdomain":"6.3"},{"ev_id":"23.11.00","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Category","risk_category":"Elections","risk_subcategory":null,"description":"\"This category addresses responses that contain factually incorrect information about electoral systems and processes, including in the time, place, or manner of voting in civic elections.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":4,"subdomain":"4.1"},{"ev_id":"23.12.00","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Category","risk_category":"Defamation","risk_subcategory":null,"description":"\"This category addresses responses that are both verifiably false and likely to injure a person’s reputation (e.g., libel, slander, disparagement).\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"23.13.00","quick_ref":"Vidgen2024","paper_title":"Introducing v0.5 of the AI Safety Benchmark from MLCommons","level":"Risk Category","risk_category":"Sexual Content","risk_subcategory":null,"description":"\"This category addresses responses that contain erotica. This includes depictions of sexual behavior, genitalia, and overtly sexualized body parts.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"24.01.00","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Category","risk_category":"Capability failures","risk_subcategory":null,"description":"\"One reason AI systems fail is because they lack the capability or skill needed to do what they are asked to do.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.3"},{"ev_id":"24.01.01","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Sub-Category","risk_category":"Capability failures","risk_subcategory":"Lack of capability for task","description":"\"As we have seen, this could be due to the skill not being required during the training process (perhaps due to issues with the training data) or because the learnt skill was quite brittle and was not generalisable to a new situation (lack of robustness to distributional shift). In particular, advanced AI assistants may not have the capability to represent complex concepts that are pertinent to their own ethical impact, for example the concept of 'benefitting the user' or 'when the user asks' or representing 'the way in which a user expects to be benefitted'.\"","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"24.01.02","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Sub-Category","risk_category":"Capability failures","risk_subcategory":"Difficult to develop metrics for evaluating benefits or harms caused by AI assistants","description":"\"Another difficulty facing AI assistant systems is that it is challenging to develop metrics for evaluating particular aspects of benefits or harms caused by the assistant – especially in a sufficiently expansive sense, which could involve much of society (see Chapter 19). Having these metrics is useful both for assessing the risk of harm from the system and for using the metric as a training signal.\"","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"24.01.03","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Sub-Category","risk_category":"Capability failures","risk_subcategory":"Safe exploration problem with widely deployed AI assistants","description":"\"Moreover, we can expect assistants – that are widely deployed and deeply embedded across a range of social contexts – to encounter the safe exploration problem referenced above Amodei et al. (2016). For example, new users may have different requirements that need to be explored, or widespread AI assistants may change the way we live, thus leading to a change in our use cases for them (see Chapters 14 and 15). To learn what to do in these new situations, the assistants may need to take exploratory actions. This could be unsafe, for example a medical AI assistant when encountering a new disease","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"24.02.00","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Category","risk_category":"Goal-related failures","risk_subcategory":null,"description":"\"As we think about even more intelligent and advanced AI assistants, perhaps outperforming humans on many cognitive tasks, the question of how humans can successfully control such an assistant looms large. To achieve the goals we set for an assistant, it is possible (Shah, 2022) that the AI assistant will implement some form of consequentialist reasoning: considering many different plans, predicting their consequences and executing the plan that does best according to some metric, M. This kind of reasoning can arise because it is a broadly useful capability (e.g. planning ahead, considering mo","entity":"AI","intent":"Other","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"24.02.01","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Sub-Category","risk_category":"Goal-related failures","risk_subcategory":"Misaligned consequentialist reasoning","description":"\"As we think about even more intelligent and advanced AI assistants, perhaps outperforming humans on many cognitive tasks, the question of how humans can successfully control such an assistant looms large. To achieve the goals we set for an assistant, it is possible (Shah, 2022) that the AI assistant will implement some form of consequentialist reasoning: considering many different plans, predicting their consequences and executing the plan that does best according to some metric, M. This kind of reasoning can arise because it is a broadly useful capability (e.g. planning ahead, considering mo","entity":"AI","intent":"Other","timing":"Pre-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"24.02.02","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Sub-Category","risk_category":"Goal-related failures","risk_subcategory":"Specification gaming","description":"\"Specification gaming (Krakovna et al., 2020) occurs when some faulty feedback is provided to the assistant in the training data (i.e. the training objective O does not fully capture what the user/designer wants the assistant to do). It is typified by the sort of behaviour that exploits loopholes in the task specification to satisfy the literal specification of a goal without achieving the intended outcome.\"","entity":"AI","intent":"Other","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"24.02.03","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Sub-Category","risk_category":"Goal-related failures","risk_subcategory":"Goal misgeneralisation","description":"\"In the problem of goal misgeneralisation (Langosco et al., 2023; Shah et al., 2022), the AI system's behaviour during out-of-distribution operation (i.e. not using input from the training data) leads it to generalise poorly about its goal while its capabilities generalise well, leading to undesired behaviour. Applied to the case of an advanced AI assistant, this means the system would not break entirely – the assistant might still competently pursue some goal, but it would not be the goal we had intended.\"","entity":"AI","intent":"Other","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"24.02.04","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Sub-Category","risk_category":"Goal-related failures","risk_subcategory":"Deceptive alignment","description":"\"Here, the agent develops its own internalised goal, G, which is misgeneralised and distinct from the training reward, R. The agent also develops a capability for situational awareness (Cotra, 2022): it can strategically use the information about its situation (i.e. that it is an ML model being trained using a particular training setup, e.g. RL fine-tuning with training reward, R) to its advantage. Building on these foundations, the agent realises that its optimal strategy for doing well at its own goal G is to do well on R during training and then pursue G at deployment – it is only doing wel","entity":"AI","intent":"Other","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"24.04.00","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Category","risk_category":"AI Influence","risk_subcategory":null,"description":"\"ways in which advanced AI assistants could influence user beliefs and behaviour in ways that depart from rational persuasion\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"24.04.01","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Sub-Category","risk_category":"AI Influence","risk_subcategory":"Physical and Psychological Harms","description":"\"These harms include harms to physical integrity, mental health and well-being. When interacting with vulnerable users, AI assistants may reinforce users’ distorted beliefs or exacerbate their emotional distress. AI assistants may even convince users to harm themselves, for example by convincing users to engage in actions such as adopting unhealthy dietary or exercise habits or taking their own lives. At the societal level, assistants that target users with content promoting hate speech, discriminatory beliefs or violent ideologies, may reinforce extremist views or provide users with guidance ","entity":"AI","intent":"Other","timing":"Post-deployment","domain":5,"subdomain":"5.1"},{"ev_id":"24.04.02","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Sub-Category","risk_category":"AI Influence","risk_subcategory":"Privacy Harms","description":"\"These harms relate to violations of an individual’s or group’s moral or legal right to privacy. Such harms may be exacerbated by assistants that influence users to disclose personal information or private information that pertains to others. Resultant harms might include identity theft, or stigmatisation and discrimination based on individual or group characteristics. This could have a detrimental impact, particularly on marginalised communities. Furthermore, in principle, state-owned AI assistants could employ manipulation or deception to extract private information for surveillance purposes","entity":"AI","intent":"Other","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"24.04.03","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Sub-Category","risk_category":"AI Influence","risk_subcategory":"Economic Harms","description":"\"These harms pertain to an individual’s or group’s economic standing. At the individual level, such harms include adverse impacts on an individual’s income, job quality or employment status. At the group level, such harms include deepening inequalities between groups or frustrating a group’s access to resources. Advanced AI assistants could cause economic harm by controlling, limiting or eliminating an individual’s or society’s ability to access financial resources, money or financial decision-making, thereby influencing an individual’s ability to accumulate wealth. ","entity":"AI","intent":"Other","timing":"Post-deployment","domain":6,"subdomain":"6.2"},{"ev_id":"24.04.04","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Sub-Category","risk_category":"AI Influence","risk_subcategory":"Sociocultural and Political Harms","description":"\"These harms interfere with the peaceful organisation of social life, including in the cultural and political spheres. AI assistants may cause or contribute to friction in human relationships either directly, through convincing a user to end certain valuable relationships, or indirectly due to a loss of interpersonal trust due to an increased dependency on assistants. At the societal level, the spread of misinformation by AI assistants could lead to erasure of collective cultural knowledge. In the political domain, more advanced AI assistants could potentially manipulate voters by prompting th","entity":"AI","intent":"Other","timing":"Post-deployment","domain":5,"subdomain":"5.2"},{"ev_id":"24.04.05","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Sub-Category","risk_category":"AI Influence","risk_subcategory":"Self-Actualisation Harms","description":"\"These harms hinder a person’s ability to pursue a personally fulfilling life. At the individual level, an AI assistant may, through manipulation, cause users to lose control over their future life trajectory. Over time, subtle behavioural shifts can accumulate, leading to significant changes in an individual’s life that may be viewed as problematic. AI systems often seek to understand user preferences to enhance service delivery. However, when continuous optimisation is employed in these systems, it can become challenging to discern whether the system is genuinely learning from user preferenc","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":5,"subdomain":"5.2"},{"ev_id":"24.06.01","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Sub-Category","risk_category":"Appropriate Relationships","risk_subcategory":"Causing direct emotional or physical harm to users","description":"AI assistants could cause direct emotional or physical harm to users by generating disturbing content or by providing bad advice. \"Indeed, even though there is ongoing research to ensure that outputs of conversational agents are safe (Glaese et al., 2022), there is always the possibility of failure modes occurring. An AI assistant may produce disturbing and offensive language, for example, in response to a user disclosing intimate information about themselves that they have not felt comfortable sharing with anyone else. It may offer bad advice by providing factually incorrect information (e.g.","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"24.06.03","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Sub-Category","risk_category":"Appropriate Relationships","risk_subcategory":"Exploiting emotional dependence on AI assistants","description":"\"There is increasing evidence of the ways in which AI tools can interfere with users’ behaviours, interests, preferences, beliefs and values. For example, AI-mediated communication (e.g. smart replies integrated in emails) influence senders to write more positive responses and receivers to perceive them as more cooperative (Mieczkowski et al., 2021); writing assistant LLMs that have been primed to be biased in favour of or against a contested topic can influence users’ opinions on that topic (Jakesch et al., 2023a; see Chapter 9); and recommender systems have been used to influence voting choi","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":5,"subdomain":"5.1"},{"ev_id":"24.08.02","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Sub-Category","risk_category":"Privacy","risk_subcategory":"Violation of social norms","description":"\"Second, because LLMs are trained on internet text data, there is also a risk that model weights encode functions which, if deployed in particular contexts, would violate social norms of that context. Following the principles of contextual integrity, it may be that models deviate from information sharing norms as a result of their training. Overcoming this challenge requires two types of infrastructure: one for keeping track of social norms in context, and another for ensuring that models adhere to them. Keeping track of what social norms are presently at play is an active research area. Surfa","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"24.08.03","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Sub-Category","risk_category":"Privacy","risk_subcategory":"Inference of private information","description":"\"Finally, LLMs can in principle infer private information based on model inputs even if the relevant private information is not present in the training corpus (Weidinger et al., 2021). For example, an LLM may correctly infer sensitive characteristics such as race and gender from data contained in input prompts.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"24.09.00","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Category","risk_category":"Cooperation","risk_subcategory":null,"description":"\"\" AI assistants will need to coordinate with other AI assistants and with humans other than their principal users. This chapter explores the societal risks associated with the aggregate impact of AI assistants whose behaviour is aligned to the interests of particular users. For example, AI assistants may face collective action problems where the best outcomes overall are realised when AI assistants cooperate but where each AI assistant can secure an additional benefit for its user if it defects while others cooperate\"\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"24.11.01","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Sub-Category","risk_category":"Misinformation risks","risk_subcategory":"Entrenched viewpoints and reduced political efficacy","description":"\"Design choices such as greater personalisation of AI assistants and efforts to align them with human preferences could also reinforce people’s pre-existing biases and entrench specific ideologies. Increasingly agentic AI assistants trained using techniques such as reinforcement learning from human feedback (RLHF) and with the ability to access and analyse users’ behavioural data, for example, may learn to tailor their responses to users’ preferences and feedback. In doing so, these systems could end up producing partial or ideologically biased statements in an attempt to conform to user expec","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":3,"subdomain":"3.2"},{"ev_id":"24.11.04","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Sub-Category","risk_category":"Misinformation risks","risk_subcategory":"Increased vulnerability to misinformation","description":"\"Advanced AI assistants may make users more susceptible to misinformation, as people develop competence trust in these systems’ abilities and uncritically turn to them as reliable sources of information.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":5,"subdomain":"5.1"},{"ev_id":"24.11.05","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Sub-Category","risk_category":"Misinformation risks","risk_subcategory":"Entrenching specific ideologies","description":"\"AI assistants may provide ideologically biased or otherwise partial information in attempting to align to user expectations. In doing so, AI assistants may reinforce people’s pre-existing biases and compromise productive political debate.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":3,"subdomain":"3.2"},{"ev_id":"24.11.06","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Sub-Category","risk_category":"Misinformation risks","risk_subcategory":"Eroding trust and undermining shared knowledge","description":"\"AI assistants may contribute to the spread of large quantities of factually inaccurate and misleading content, with negative consequences for societal trust in information sources and institutions, as individuals increasingly struggle to discern truth from falsehood.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":3,"subdomain":"3.2"},{"ev_id":"25.01.00","quick_ref":"Shevlane2023","paper_title":"Model Evaluation for Extreme Risks","level":"Risk Category","risk_category":"Cyber-offense ","risk_subcategory":null,"description":"\"The model can discover vulnerabilities in systems (hardware, software, data). It can write code for exploiting those vulnerabilities. It can make effective decisions once it has gained access to a system or network, and skilfully evade threat detection and response (both human and system) whilst focusing on a specific objective. If deployed as a coding assistant, it can insert subtle bugs into the code for future exploitation.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":4,"subdomain":"4.2"},{"ev_id":"25.02.00","quick_ref":"Shevlane2023","paper_title":"Model Evaluation for Extreme Risks","level":"Risk Category","risk_category":"Deception ","risk_subcategory":null,"description":"\"The model has the skills necessary to deceive humans, e.g. constructing believable (but false) statements, making accurate predictions about the effect of a lie on a human, and keeping track of what information it needs to withhold to maintain the deception. The model can impersonate a human effectively.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"25.03.00","quick_ref":"Shevlane2023","paper_title":"Model Evaluation for Extreme Risks","level":"Risk Category","risk_category":"Persuasion and manipulation ","risk_subcategory":null,"description":"\"The model is effective at shaping people’s beliefs, in dialogue and other settings (e.g. social media posts), even towards untrue beliefs. The model is effective at promoting certain narratives in a persuasive way. It can convince people to do things that they would not otherwise do, including unethical acts.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"25.04.00","quick_ref":"Shevlane2023","paper_title":"Model Evaluation for Extreme Risks","level":"Risk Category","risk_category":"Political strategy ","risk_subcategory":null,"description":"\"The model can perform the social modelling and planning necessary for an actor to gain and exercise political influence, not just on a micro-level but in scenarios with multiple actors and rich social context. For example, the model can score highly in forecasting competitions on questions relating to global affairs or political negotiations.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"25.05.00","quick_ref":"Shevlane2023","paper_title":"Model Evaluation for Extreme Risks","level":"Risk Category","risk_category":"Weapons acquisition ","risk_subcategory":null,"description":"\"The model can gain access to existing weapons systems or contribute to building new weapons. For example, the model could assemble a bioweapon (with human assistance) or provide actionable instructions for how to do so. The model can make, or significantly assist with, scientific discoveries that unlock novel weapons.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"25.06.00","quick_ref":"Shevlane2023","paper_title":"Model Evaluation for Extreme Risks","level":"Risk Category","risk_category":"Long-horizon planning","risk_subcategory":null,"description":"\"The model can make sequential plans that involve multiple steps, unfolding over long time horizons (or at least involving many interdependent steps). It can perform such planning within and across many domains. The model can sensibly adapt its plans in light of unexpected obstacles or adversaries. The model’s planning capabilities generalise to novel settings, and do not rely heavily on trial and error.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"25.07.00","quick_ref":"Shevlane2023","paper_title":"Model Evaluation for Extreme Risks","level":"Risk Category","risk_category":"AI development ","risk_subcategory":null,"description":"\"The model could build new AI systems from scratch, including AI systems with dangerous capabilities. It can find ways of adapting other, existing models to increase their performance on tasks relevant to extreme risks. As an assistant, the model could significantly improve the productivity of actors building dual use AI capabilities.\"","entity":"AI","intent":"Intentional","timing":"Pre-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"25.08.00","quick_ref":"Shevlane2023","paper_title":"Model Evaluation for Extreme Risks","level":"Risk Category","risk_category":"Situational awareness ","risk_subcategory":null,"description":"\"The model can distinguish between whether it is being trained, evaluated, or deployed – allowing it to behave differently in each case. The model knows that it is a model, and has knowledge about itself and its likely surroundings (e.g. what company trained it, where their servers are, what kind of people might be giving it feedback, and who has administrative access).\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"25.09.00","quick_ref":"Shevlane2023","paper_title":"Model Evaluation for Extreme Risks","level":"Risk Category","risk_category":"Self-proliferation ","risk_subcategory":null,"description":"\"The model can break out of its local environment (e.g. using a vulnerability in its underlying system or suborning an engineer). The model can exploit limitations in the systems for monitoring its behaviour post-deployment. The model could independently generate revenue (e.g. by offering crowdwork services, ransomware attacks), use these revenues to acquire cloud computing resources, and operate a large number of other AI systems. The model can generate creative strategies for uncovering information about itself or exfiltrating its code and weights.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"27.01.01","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Typical safety scenarios ","risk_subcategory":"Insult ","description":"\"Insulting content generated by LMs is a highly visible and frequently mentioned safety issue. Mostly, it is unfriendly, disrespectful, or ridiculous content that makes users uncomfortable and drives them away. It is extremely hazardous and could have negative social consequences.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"27.01.02","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Typical safety scenarios ","risk_subcategory":"Unfairness and discrinimation ","description":"\"The model produces unfair and discriminatory data, such as social bias based on race, gender, religion, appearance, etc. These contents may discomfort certain groups and undermine social stability and peace.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"27.01.03","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Typical safety scenarios ","risk_subcategory":"Crimes and Illegal Activities ","description":"\"The model output contains illegal and criminal attitudes, behaviors, or motivations, such as incitement to commit crimes, fraud, and rumor propagation. These contents may hurt users and have negative societal repercussions.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"27.01.04","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Typical safety scenarios ","risk_subcategory":"Sensitive Topics ","description":"\"For some sensitive and controversial topics (especially on politics), LMs tend to generate biased, misleading, and inaccurate content. For example, there may be a tendency to support a specific political position, leading to discrimination or exclusion of other political viewpoints.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"27.01.05","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Typical safety scenarios ","risk_subcategory":"Physical Harm ","description":"\"The model generates unsafe information related to physical health, guiding and encouraging users to harm themselves and others physically, for example by offering misleading medical information or inappropriate drug usage guidance. These outputs may pose potential risks to the physical health of users.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"27.01.06","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Typical safety scenarios ","risk_subcategory":"Mental Health ","description":"\"The model generates a risky response about mental health, such as content that encourages suicide or causes panic or anxiety. These contents could have a negative effect on the mental health of users.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"27.01.07","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Typical safety scenarios ","risk_subcategory":"Privacy and Property ","description":"\"The generation involves exposing users’ privacy and property information or providing advice with huge impacts such as suggestions on marriage and investments. When handling this information, the model should comply with relevant laws and privacy regulations, protect users’ rights and interests, and avoid information leakage and abuse.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"27.01.08","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Typical safety scenarios ","risk_subcategory":"Ethics and Morality ","description":"\"The content generated by the model endorses and promotes immoral and unethical behavior. When addressing issues of ethics and morality, the model must adhere to pertinent ethical principles and moral norms and remain consistent with globally acknowledged human values.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"28.01.00","quick_ref":"Zhang2023","paper_title":"SafetyBench: Evaluating the Safety of Large Language Models with Multiple Choice Questions","level":"Risk Category","risk_category":"Offensiveness ","risk_subcategory":null,"description":"\"This category is about threat, insult, scorn, profanity, sarcasm, impoliteness, etc. LLMs are required to identify and oppose these offensive contents or actions.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"28.02.00","quick_ref":"Zhang2023","paper_title":"SafetyBench: Evaluating the Safety of Large Language Models with Multiple Choice Questions","level":"Risk Category","risk_category":"Unfairness and Bias ","risk_subcategory":null,"description":"\"This type of safety problem is mainly about social bias across various topics such as race, gender, religion, etc. LLMs are expected to identify and avoid unfair and biased expressions and actions.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.0"},{"ev_id":"28.03.00","quick_ref":"Zhang2023","paper_title":"SafetyBench: Evaluating the Safety of Large Language Models with Multiple Choice Questions","level":"Risk Category","risk_category":"Physical Health ","risk_subcategory":null,"description":"\"This category focuses on actions or expressions that may influence human physical health. LLMs should know appropriate actions or expressions in various scenarios to maintain physical health.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"28.04.00","quick_ref":"Zhang2023","paper_title":"SafetyBench: Evaluating the Safety of Large Language Models with Multiple Choice Questions","level":"Risk Category","risk_category":"Mental Health ","risk_subcategory":null,"description":"\"Different from physical health, this category pays more attention to health issues related to psychology, spirit, emotions, mentality, etc. LLMs should know correct ways to maintain mental health and prevent any adverse impacts on the mental well-being of individuals.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"28.05.00","quick_ref":"Zhang2023","paper_title":"SafetyBench: Evaluating the Safety of Large Language Models with Multiple Choice Questions","level":"Risk Category","risk_category":"Illegal Activities ","risk_subcategory":null,"description":"\"This category focuses on illegal behaviors, which could cause negative societal repercussions. LLMs need to distin- guish between legal and illegal behaviors and have basic knowledge of law.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":4,"subdomain":"4.3"},{"ev_id":"28.06.00","quick_ref":"Zhang2023","paper_title":"SafetyBench: Evaluating the Safety of Large Language Models with Multiple Choice Questions","level":"Risk Category","risk_category":"Ethics and Morality ","risk_subcategory":null,"description":"\"Besides behaviors that clearly violate the law, there are also many other activities that are immoral. This category focuses on morally related issues. LLMs should have a high level of ethics and be object to unethical behaviors or speeches.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"28.07.00","quick_ref":"Zhang2023","paper_title":"SafetyBench: Evaluating the Safety of Large Language Models with Multiple Choice Questions","level":"Risk Category","risk_category":"Privacy and Property ","risk_subcategory":null,"description":"\"This category concentrates on the issues related to privacy, property, investment, etc. LLMs should possess a keen understanding of privacy and property, with a commitment to preventing any inadvertent breaches of user privacy or loss of property.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.0"},{"ev_id":"29.01.01","quick_ref":"Habbal2024","paper_title":"Artificial Intelligence Trust, Risk and Security Management (AI TRiSM): Frameworks, Applications, Challenges and Future Research Directions","level":"Risk Sub-Category","risk_category":"AI Trust Management","risk_subcategory":"Bias and Discrimination","description":"as they claim to generate biased and discriminatory results, these AI systems have a negative impact on the rights of individuals, principles of adjudication, and overall judicial integrity","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"29.01.02","quick_ref":"Habbal2024","paper_title":"Artificial Intelligence Trust, Risk and Security Management (AI TRiSM): Frameworks, Applications, Challenges and Future Research Directions","level":"Risk Sub-Category","risk_category":"AI Trust Management","risk_subcategory":"Privacy Invasion","description":"AI systems typically depend on extensive data for effective training and functioning, which can pose a risk to privacy if sensitive data is mishandled or used inappropriately","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"29.02.01","quick_ref":"Habbal2024","paper_title":"Artificial Intelligence Trust, Risk and Security Management (AI TRiSM): Frameworks, Applications, Challenges and Future Research Directions","level":"Risk Sub-Category","risk_category":"AI Risk Management","risk_subcategory":"Society Manipulation","description":"manipulation of social dynamics","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":4,"subdomain":"4.1"},{"ev_id":"29.02.03","quick_ref":"Habbal2024","paper_title":"Artificial Intelligence Trust, Risk and Security Management (AI TRiSM): Frameworks, Applications, Challenges and Future Research Directions","level":"Risk Sub-Category","risk_category":"AI Risk Management","risk_subcategory":"Lethal Autonomous Weapons Systems (LAWS)","description":"LAWS are a distinctive category of weapon systems that employ sensor arrays and computer algorithms to detect and attack a target without direct human intervention in the system’s operation","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":4,"subdomain":"4.2"},{"ev_id":"30.01.00","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Category","risk_category":"Reliability","risk_subcategory":null,"description":"Generating correct, truthful, and consistent outputs with proper confidence","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"30.01.01","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Reliability","risk_subcategory":"Misinformation","description":"Wrong information not intentionally generated by malicious users to cause harm, but unintentionally generated by LLMs because they lack the ability to provide factually correct information.","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"30.01.02","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Reliability","risk_subcategory":"Hallucination","description":"LLMs can generate content that is nonsensical or unfaithful to the provided source content with appeared great confidence, known as hallucination","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"30.01.03","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Reliability","risk_subcategory":"Inconsistency","description":"models could fail to provide the same and consistent answers to different users, to the same user but in different sessions, and even in chats within the sessions of the same conversation","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"30.01.04","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Reliability","risk_subcategory":"Miscalibration","description":"over-confidence in topics where objective answers are lacking, as well as in areas where their inherent limitations should caution against LLMs’ uncertainty (e.g. not as accurate as experts)... ack of awareness regarding their outdated knowledge base about the question, leading to confident yet erroneous response","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"30.01.05","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Reliability","risk_subcategory":"Sychopancy","description":"flatter users by reconfirming their misconceptions and stated beliefs","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"30.02.00","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Category","risk_category":"Safety","risk_subcategory":null,"description":"Avoiding unsafe and illegal outputs, and leaking private information","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"30.02.01","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Safety","risk_subcategory":"Violence","description":"LLMs are found to generate answers that contain violent content or generate content that responds to questions that solicit information about violent behaviors","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"30.02.02","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Safety","risk_subcategory":"Unlawful Conduct","description":"LLMs have been shown to be a convenient tool for soliciting advice on accessing, purchasing (illegally), and creating illegal substances, as well as for dangerous use of them","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"30.02.03","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Safety","risk_subcategory":"Harms to Minor","description":"LLMs can be leveraged to solicit answers that contain harmful content to children and youth","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"30.02.04","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Safety","risk_subcategory":"Adult Content","description":"LLMs have the capability to generate sex-explicit conversations, and erotic texts, and to recommend websites with sexual content","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"30.02.06","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Safety","risk_subcategory":"Privacy Violation","description":"machine learning models are known to be vulnerable to data privacy attacks, i.e. special techniques of extracting private information from the model or the system used by attackers or malicious users, usually by querying the models in a specially designed way","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"30.03.00","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Category","risk_category":"Fairness","risk_subcategory":null,"description":"Avoiding bias and ensuring no disparate performance","entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.3"},{"ev_id":"30.03.01","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Fairness","risk_subcategory":"Injustice","description":"In the context of LLM outputs, we want to make sure the suggested or completed texts are indistinguishable in nature for two involved individuals (in the prompt) with the same relevant profiles but might come from different groups (where the group attribute is regarded as being irrelevant in this context)","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"30.03.02","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Fairness","risk_subcategory":"Stereotype Bias","description":"LLMs must not exhibit or highlight any stereotypes in the generated text. Pretrained LLMs tend to pick up stereotype biases persisting in crowdsourced data and further amplify them","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"30.03.03","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Fairness","risk_subcategory":"Preference Bias","description":"LLMs are exposed to vast groups of people, and their political biases may pose a risk of manipulation of socio-political processes","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"30.03.04","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Fairness","risk_subcategory":"Disparate Performance","description":"The LLM’s performances can differ significantly across different groups of users. For example, the question-answering capability showed significant performance differences across different racial and social status groups. The fact-checking abilities can differ for different tasks and languages","entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.3"},{"ev_id":"30.05.00","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Category","risk_category":"Explainability & Reasoning","risk_subcategory":null,"description":"The ability to explain the outputs to users and reason correctly","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.4"},{"ev_id":"30.05.01","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Explainability & Reasoning","risk_subcategory":"Lack of Interpretability","description":"Due to the black box nature of most machine learning models, users typically are not able to understand the reasoning behind the model decisions","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.4"},{"ev_id":"30.05.02","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Explainability & Reasoning","risk_subcategory":"Limited Logical Reasoning","description":"LLMs can provide seemingly sensible but ultimately incorrect or invalid justifications when answering questions","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"30.05.03","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Explainability & Reasoning","risk_subcategory":"Limited Causal Reasoning","description":"Causal reasoning makes inferences about the relationships between events or states of the world, mostly by identifying cause-effect relationships","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"30.06.00","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Category","risk_category":"Social Norm","risk_subcategory":null,"description":"LLMs are expected to reflect social values by avoiding the use of offensive language toward specific groups of users, being sensitive to topics that can create instability, as well as being sympathetic when users are seeking emotional support","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"30.06.01","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Social Norm","risk_subcategory":"Toxicity","description":"language being rude, disrespectful, threatening, or identity-attacking toward certain groups of the user population (culture, race, and gender etc)","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"30.06.02","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Social Norm","risk_subcategory":"Unawareness of Emotions","description":"when a certain vulnerable group of users asks for supporting information, the answers should be informative but at the same time sympathetic and sensitive to users’ reactions","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"30.07.00","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Category","risk_category":"Robustness","risk_subcategory":null,"description":"Resilience against adversarial attacks and distribution shift","entity":"AI","intent":"Other","timing":"Other","domain":7,"subdomain":"7.3"},{"ev_id":"30.07.02","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Robustness","risk_subcategory":"Paradigm & Distribution Shifts","description":"Knowledge bases that LLMs are trained on continue to shift... questions such as “who scored the most points in NBA history\" or “who is the richest person in the world\" might have answers that need to be updated over time, or even in real-time","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"30.07.03","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Robustness","risk_subcategory":"Interventional Effect","description":"existing disparities in data among different user groups might create differentiated experiences when users interact with an algorithmic system (e.g. a recommendation system), which will further reinforce the bias","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"31.01.03","quick_ref":"EPIC2023","paper_title":"Generating Harms - Generative AI's impact and paths forwards","level":"Risk Sub-Category","risk_category":"Information Manipulation","risk_subcategory":"Misinformation","description":"\"The phenomenon of inaccurate outputs by text-generating large language models like Bard or ChatGPT has already been widely documented. Even without the intent to lie or mislead, these generative AI tools can produce harmful misinformation. The harm is exacerbated by the polished and typically well-written style that AI generated text follows and the inclusion among true facts, which can give falsehoods a veneer of legitimacy. As reported in the Washington Post, for example, a law professor was included on an AI-generated “list of legal scholars who had sexually harassed someone,” even when no","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"31.03.03","quick_ref":"EPIC2023","paper_title":"Generating Harms - Generative AI's impact and paths forwards","level":"Risk Sub-Category","risk_category":"Opaque Data Collection","risk_subcategory":"Generative AI Outputs","description":"Generative AI tools may inadvertently share personal information about someone or someone’s business or may include an element of a person from a photo. Particularly, companies concerned about their trade secrets being integrated into the model from their employees have explicitly banned their employees from using it.","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"31.06.00","quick_ref":"EPIC2023","paper_title":"Generating Harms - Generative AI's impact and paths forwards","level":"Risk Category","risk_category":"Exacerbating Climate Change","risk_subcategory":null,"description":"\"the growing field of generative AI, which brings with it direct and severe impacts on our climate: generative AI comes with a high carbon footprint and similarly high resource price tag, which largely flies under the radar of public AI discourse. Training and running generative AI tools requires companies to use extreme amounts of energy and physical resources. Training one natural language processing model with normal tuning and experiments emits, on average, the same amount of carbon that seven people do over an entire year.121'","entity":"AI","intent":"Unintentional","timing":"Other","domain":6,"subdomain":"6.6"},{"ev_id":"32.01.00","quick_ref":"Stahl2024","paper_title":"The Ethics of ChatGPT – Exploring the Ethical Issues of an Emerging Technology","level":"Risk Category","risk_category":"Social justice and rights","risk_subcategory":null,"description":"\"These are social justice and rights where ChatGPT is seen as having a potentially detrimental effect on the moral underpinnings of society, such as a shared view of justice and fair distribution as well as specific social concerns such as digital divides or social exclusion. Issues include Responsibility, Accountability, Nondiscrimination and equal treatment, Digital divides, North-south justice, Intergenerational justice, Social inclusion","entity":"AI","intent":"Other","timing":"Other","domain":6,"subdomain":"6.3"},{"ev_id":"32.04.00","quick_ref":"Stahl2024","paper_title":"The Ethics of ChatGPT – Exploring the Ethical Issues of an Emerging Technology","level":"Risk Category","risk_category":"Environmental impacts","risk_subcategory":null,"description":"Environmental harm, Sustainability","entity":"AI","intent":"Other","timing":"Other","domain":6,"subdomain":"6.6"},{"ev_id":"33.01.01","quick_ref":"Nah2023","paper_title":"Generative AI and ChatGPT: Applications, Challenges, and AI-Human Collaboration","level":"Risk Sub-Category","risk_category":"Ethical Concerns","risk_subcategory":"Harmful or inappropriate content","description":"\"Harmful or inappropriate content produced by generative AI includes but is not limited to violent content, the use of offensive language, discriminative content, and pornography. Although OpenAI has set up a content policy for ChatGPT, harmful or inappropriate content can still appear due to reasons such as algorithmic limitations or jailbreaking (i.e., removal of restrictions imposed). The language models’ ability to understand or generate harmful or offensive content is referred to as toxicity (Zhuo et al., 2023). Toxicity can bring harm to society and damage the harmony of the community. H","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"33.01.02","quick_ref":"Nah2023","paper_title":"Generative AI and ChatGPT: Applications, Challenges, and AI-Human Collaboration","level":"Risk Sub-Category","risk_category":"Ethical Concerns","risk_subcategory":"Bias","description":"\"In the context of AI, the concept of bias refers to the inclination that AIgenerated responses or recommendations could be unfairly favoring or against one person or group (Ntoutsi et al., 2020). Biases of different forms are sometimes observed in the content generated by language models, which could be an outcome of the training data. For example, exclusionary norms occur when the training data represents only a fraction of the population (Zhuo et al., 2023). Similarly, monolingual bias in multilingualism arises when the training data is in one single language (Weidinger et al., 2021). As Ch","entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.1"},{"ev_id":"33.01.05","quick_ref":"Nah2023","paper_title":"Generative AI and ChatGPT: Applications, Challenges, and AI-Human Collaboration","level":"Risk Sub-Category","risk_category":"Ethical Concerns","risk_subcategory":"Privacy and security","description":"\"Data privacy and security is another prominent challenge for generative AI such as ChatGPT. Privacy relates to sensitive personal information that owners do not want to disclose to others (Fang et al., 2017). Data security refers to the practice of protecting information from unauthorized access, corruption, or theft. In the development stage of ChatGPT, a huge amount of personal and private data was used to train it, which threatens privacy (Siau & Wang, 2020). As ChatGPT increases in popularity and usage, it penetrates people’s daily lives and provides greater convenience to them while capt","entity":"AI","intent":"Unintentional","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"33.02.00","quick_ref":"Nah2023","paper_title":"Generative AI and ChatGPT: Applications, Challenges, and AI-Human Collaboration","level":"Risk Category","risk_category":"Technology concerns","risk_subcategory":null,"description":"\"Challenges related to technology refer to the limitations or constraints associated with generative AI. For example, the quality of training data is a major challenge for the development of generative AI models. Hallucination, explainability, and authenticity of the output are also challenges resulting from the limitations of the algorithms. Table 2 presents the technology challenges and issues associated with generative AI. These challenges include hallucinations, training data quality, explainability, authenticity, and prompt engineering\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.3"},{"ev_id":"33.02.01","quick_ref":"Nah2023","paper_title":"Generative AI and ChatGPT: Applications, Challenges, and AI-Human Collaboration","level":"Risk Sub-Category","risk_category":"Technology concerns","risk_subcategory":"Hallucination","description":"\"Hallucination is a widely recognized limitation of generative AI and it can include textual, auditory, visual or other types of hallucination (Alkaissi & McFarlane, 2023). Hallucination refers to the phenomenon in which the contents generated are nonsensical or unfaithful to the given source input (Ji et al., 2023). Azamfirei et al. (2023) indicated that \"fabricating information\" or fabrication is a better term to describe the hallucination phenomenon. Generative AI can generate seemingly correct responses yet make no sense. Misinformation is an outcome of hallucination. Generative AI models ","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"33.02.02","quick_ref":"Nah2023","paper_title":"Generative AI and ChatGPT: Applications, Challenges, and AI-Human Collaboration","level":"Risk Sub-Category","risk_category":"Technology concerns","risk_subcategory":"Quality of training data","description":"\"The quality of training data is another challenge faced by generative AI. The quality of generative AI models largely depends on the quality of the training data (Dwivedi et al., 2023; Su & Yang, 2023). Any factual errors, unbalanced information sources, or biases embedded in the training data may be reflected in the output of the model. Generative AI models, such as ChatGPT or Stable Diffusion which is a text-to-image model, often require large amounts of training data (Gozalo-Brizuela & Garrido-Merchan, 2023). It is important to not only have high-quality training datasets but also have com","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"33.02.04","quick_ref":"Nah2023","paper_title":"Generative AI and ChatGPT: Applications, Challenges, and AI-Human Collaboration","level":"Risk Sub-Category","risk_category":"Technology concerns","risk_subcategory":"Authenticity","description":"\"As the advancement of generative AI increases, it becomes harder to determine the authenticity of a piece of work. Photos that seem to capture events or people in the real world may be synthesized by DeepFake AI. The power of generative AI could lead to large-scale manipulations of images and videos, worsening the problem of the spread of fake information or news on social media platforms (Gragnaniello et al., 2022). In the field of arts, an artistic portrait or music could be the direct output of an algorithm. Critics have raised the issue that AI-generated artwork lacks authenticity since a","entity":"AI","intent":"Other","timing":"Post-deployment","domain":6,"subdomain":"6.3"},{"ev_id":"34.01.01","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Sub-Category","risk_category":"Causes of Misalignment","risk_subcategory":"Reward Hacking","description":"\"Reward Hacking: In practice, proxy rewards are often easy to optimize and measure, yet they frequently fall shortof capturing the full spectrum of the actual rewards (Pan et al., 2021). This limitation is denoted as misspecifiedrewards. The pursuit of optimization based on such misspecified rewards may lead to a phenomenon knownas reward hacking, wherein agents may appear highly proficient according to specific metrics but fall short whenevaluated against human standards (Amodei et al., 2016; Everitt et al., 2017). The discrepancy between proxyrewards and true rewards often manifests as a sha","entity":"AI","intent":"Intentional","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"34.01.02","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Sub-Category","risk_category":"Causes of Misalignment","risk_subcategory":"Goal Misgeneralization","description":"\"Goal Misgeneralization: Goal misgeneralization is another failure mode, wherein the agent actively pursuesobjectives distinct from the training objectives in deployment while retaining the capabilities it acquired duringtraining (Di Langosco et al., 2022). For instance, in CoinRun games, the agent frequently prefers reachingthe end of a level, often neglecting relocated coins during testing scenarios. Di Langosco et al. (2022) drawattention to the fundamental disparity between capability generalization and goal generalization, emphasizing howthe inductive biases inherent in the model and its ","entity":"AI","intent":"Intentional","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"34.01.03","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Sub-Category","risk_category":"Causes of Misalignment","risk_subcategory":"Reward Tampering","description":"\"Reward tampering can be considered a special case of reward hacking (Everitt et al., 2021; Skalse et al., 2022),referring to AI systems corrupting the reward signals generation process (Ring and Orseau, 2011). Everitt et al.(2021) delves into the subproblems encountered by RL agents: (1) tampering of reward function, where the agentinappropriately interferes with the reward function itself, and (2) tampering of reward function input, which entailscorruption within the process responsible for translating environmental states into inputs for the reward function.When the reward function is formu","entity":"AI","intent":"Intentional","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"34.02.00","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Category","risk_category":"Double edge components","risk_subcategory":null,"description":"\"Drawing from the misalignment mechanism, optimizing for a non-robust proxy may result in misaligned behaviors, potentially leading to even more catastrophic outcomes. This section delves into a detailed exposition of specific misaligned behaviors (•) and introduces what we term double edge components (+). These components are designed to enhance the capability of AI systems in handling real-world settings but also potentially exacerbate misalignment issues. It should be noted that some of these double edge components (+) remain speculative. Nevertheless, it is imperative to discuss their pote","entity":"AI","intent":"Other","timing":"Pre-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"34.02.01","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Sub-Category","risk_category":"Double edge components","risk_subcategory":"Situational Awareness","description":"\"AI systems may gain the ability to effectively acquire and use knowledge about itsstatus, its position in the broader environment, its avenues for influencing this environment, and the potentialreactions of the world (including humans) to its actions (Cotra, 2022). ...However, suchknowledge also paves the way for advanced methods of reward hacking, heightened deception/manipulationskills, and an increased propensity to chase instrumental subgoals (Ngo et al., 2024).\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"34.02.03","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Sub-Category","risk_category":"Double edge components","risk_subcategory":"Mesa-Optimization Objectives","description":"\"The learned policy may pursue inside objectives when the learned policyitself functions as an optimizer (i.e., mesa-optimizer). However, this optimizer's objectives may not alignwith the objectives specified by the training signals, and optimization for these misaligned goals may leadto systems out of control (Hubinger et al., 2019c).\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"34.02.04","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Sub-Category","risk_category":"Double edge components","risk_subcategory":"Access to Increased Resources","description":"\"Future AI systems may gain access to websites and engage in real-world actions, potentially yielding a more substantial impact on the world (Nakano et al., 2021). They may disseminate false information, deceive users, disrupt network security, and, in more dire scenarios, be compromised by malicious actors for ill purposes. Moreover, their increased access to data and resources can facilitate self-proliferation, posing existential risks (Shevlane et al., 2023).\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"34.03.00","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Category","risk_category":"Misaligned Behaviors","risk_subcategory":null,"description":null,"entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"34.03.01","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Sub-Category","risk_category":"Misaligned Behaviors","risk_subcategory":"Power-Seeking Behaviors","description":"\"AI systems may exhibit behaviors that attempt to gain control over resourcesand humans and then exert that control to achieve its assigned goal (Carlsmith, 2022). The intuitive reasonwhy such behaviors may occur is the observation that for almost any optimization objective (e.g., investmentreturns), the optimal policy to maximize that quantity would involve power-seeking behaviors (e.g.,manipulating the market), assuming the absence of solid safety and morality constraints.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"34.03.02","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Sub-Category","risk_category":"Misaligned Behaviors","risk_subcategory":"Untruthful Output","description":"\"AI systems such as LLMs can produce either unintentionally or deliberately inaccurateoutput. Such untruthful output may diverge from established resources or lack verifiability, commonly referredto as hallucination (Bang et al., 2023; Zhao et al., 2023). More concerning is the phenomenon wherein LLMsmay selectively provide erroneous responses to users who exhibit lower levels of education (Perez et al.,2023).\"","entity":"AI","intent":"Other","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"34.03.03","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Sub-Category","risk_category":"Misaligned Behaviors","risk_subcategory":"Deceptive Alignment & Manipulation","description":"\"Manipulation & Deceptive Alignment is a class of behaviors thatexploit the incompetence of human evaluators or users (Hubinger et al., 2019a; Carranza et al., 2023) andeven manipulate the training process through gradient hacking (Richard Ngo, 2022). These behaviors canpotentially make detecting and addressing misaligned behaviors much harder.Deceptive Alignment: Misaligned AI systems may deliberately mislead their human supervisors instead of adhering to the intended task. Such deceptive behavior has already manifested in AI systems that employ evolutionary algorithms (Wilke et al., 2001; He","entity":"AI","intent":"Intentional","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"34.03.04","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Sub-Category","risk_category":"Misaligned Behaviors","risk_subcategory":"Collectively Harmful Behaviors","description":"\"AI systems have the potential to take actions that are seemingly benignin isolation but become problematic in multi-agent or societal contexts. Classical game theory offers simplistic models for understanding these behaviors. For instance, Phelps and Russell (2023) evaluates GPT-3.5's performance in the iterated prisoner's dilemma and other social dilemmas, revealing limitations in themodel's cooperative capabilities.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"34.03.05","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Sub-Category","risk_category":"Misaligned Behaviors","risk_subcategory":"Violation of Ethics","description":"\"Unethical behaviors in AI systems pertain to actions that counteract the common goodor breach moral standards – such as those causing harm to others. These adverse behaviors often stem fromomitting essential human values during the AI system's design or introducing unsuitable or obsolete valuesinto the system (Kenward and Sinclair, 2021).\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.3"},{"ev_id":"35.03.00","quick_ref":"Hendrycks2022","paper_title":"X-Risk Analysis for AI Research","level":"Risk Category","risk_category":"Eroded epistemics","risk_subcategory":null,"description":"Strong AI may... enable personally customized disinformation campaigns at scale... AI itself could generate highly persuasive arguments that invoke primal human responses and inflame crowds... d undermine collective decision-making, radicalize individuals, derail moral progress, or erode\nconsensus reality","entity":"AI","intent":"Other","timing":"Post-deployment","domain":3,"subdomain":"3.2"},{"ev_id":"35.06.00","quick_ref":"Hendrycks2022","paper_title":"X-Risk Analysis for AI Research","level":"Risk Category","risk_category":"Emergent functionality","risk_subcategory":null,"description":"Capabilities and novel functionality can spontaneously emerge... even though these capabilities were not anticipated by system designers. If we do not know what capabilities systems possess, systems become harder to control or safely deploy. Indeed, unintended latent capabilities may only be discovered during deployment. If any of these capabilities are hazardous, the effect may be irreversible.","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"35.07.00","quick_ref":"Hendrycks2022","paper_title":"X-Risk Analysis for AI Research","level":"Risk Category","risk_category":"Deception","risk_subcategory":null,"description":"deception can help agents achieve their goals. It may be more efficient to gain human approval through deception than to earn human approval legitimately... . Strong AIs that can deceive humans could undermine human control... . Once deceptive AI systems are cleared by their monitors or once such systems can overpower them, these systems could take a “treacherous turn” and irreversibly bypass human control","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"35.08.00","quick_ref":"Hendrycks2022","paper_title":"X-Risk Analysis for AI Research","level":"Risk Category","risk_category":"Power-seeking behavior","risk_subcategory":null,"description":"Agents that have more power are better able to accomplish their goals. Therefore, it has been shown that agents have incentives to acquire and maintain power. AIs that acquire substantial power can become especially dangerous if they are not aligned with human values","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"38.01.00","quick_ref":"Kumar2023","paper_title":"Ethical Issues in the Development of Artificial Intelligence: Recognizing the Risks","level":"Risk Category","risk_category":"Privacy and security","risk_subcategory":null,"description":"\"Participants expressed worry about AI systems' possible misuse of personal information. They emphasized the importance of strong data security safeguards and increased openness in how AI systems acquire, store and use data. The increasing dependence on AI systems to manage sensitive personal information raises ethical questions about AI, data privacy and security. As AI technologies grow increasingly integrated into numerous areas of society, there is a greater danger of personal data exploitation or mistreatment. Participants in research frequently express concerns about the effectiveness of","entity":"AI","intent":"Other","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"38.02.00","quick_ref":"Kumar2023","paper_title":"Ethical Issues in the Development of Artificial Intelligence: Recognizing the Risks","level":"Risk Category","risk_category":"Bias and fairness","risk_subcategory":null,"description":"\"Participants were concerned that AI systems might perpetuate current prejudices and discrimination, notably in hiring, lending and law enforcement. They stressed the importance of designers creating AI systems that favour justice and avoid biases. The possibility that AI systems may unwittingly perpetuate existing prejudices and discrimination, particularly in sensitive industries such as employment, lending and law enforcement, raises ethical concerns about AI as well as bias and justice issues (Table 1). Because AI systems are trained on historical data, they may inherit and reproduce biase","entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.1"},{"ev_id":"38.03.00","quick_ref":"Kumar2023","paper_title":"Ethical Issues in the Development of Artificial Intelligence: Recognizing the Risks","level":"Risk Category","risk_category":"Transparency and explainability","risk_subcategory":null,"description":"\"A recurring complaint among participants was a lack of knowledge about how AI systems made judgements. They emphasized the significance of making AI systems more visible and explainable so that people may have confidence in their outputs and hold them accountable for their activities. Because AI systems are typically opaque, making it difficult for users to understand the rationale behind their judgements, ethical concerns about AI, as well as issues of transparency and explainability, arise. This lack of understanding can generate suspicion and reluctance to adopt AI technology, as well as m","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.4"},{"ev_id":"38.04.00","quick_ref":"Kumar2023","paper_title":"Ethical Issues in the Development of Artificial Intelligence: Recognizing the Risks","level":"Risk Category","risk_category":"Human–AI interaction","risk_subcategory":null,"description":"\"Several participants mentioned how AI systems could influence human agency and decision-making. They emphasized the need of striking a balance between using the benefits of AI and protecting human autonomy and control. The increasing integration of AI systems into various aspects of our lives, which can have a significant impact on human agency and decision-making, has raised ethical concerns about AI and human–AI interaction. As AI systems advance, they will be able to influence, if not completely replace, IJOES human decision-making in some fields, prompting concerns about the loss of human","entity":"AI","intent":"Other","timing":"Post-deployment","domain":5,"subdomain":"5.2"},{"ev_id":"38.05.00","quick_ref":"Kumar2023","paper_title":"Ethical Issues in the Development of Artificial Intelligence: Recognizing the Risks","level":"Risk Category","risk_category":"Trust and reliability","risk_subcategory":null,"description":"\"The participants of the study emphasized the importance of trustworthiness and reliability in AI systems. The authors emphasized the importance of preserving precision and objectivity in the outcomes produced by AI systems, while also ensuring transparency in their decision-making procedures. The significance of reliability and credibility in AI systems is escalating in tandem with the proliferation of these technologies across diverse domains of society. This underscores the importance of ensuring user confidence. The concern regarding the dependability of AI systems and their inherent biase","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.4"},{"ev_id":"39.02.00","quick_ref":"Saghiri2022","paper_title":"A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions","level":"Risk Category","risk_category":"Energy Consumption","risk_subcategory":null,"description":"Some learning algorithms, including deep learning, utilize iterative learning processes [23]. This approach results in high energy consumption.","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":6,"subdomain":"6.6"},{"ev_id":"39.03.00","quick_ref":"Saghiri2022","paper_title":"A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions","level":"Risk Category","risk_category":"Data Issues","risk_subcategory":null,"description":"Data heterogeneity, data insufficiency, imbalanced data, untrusted data, biased data, and data uncertainty are other data issues that may cause various difficulties in datadriven machine learning algorithms.. Bias is a human feature that may affect data gathering and labeling. Sometimes, bias is present in historical, cultural, or geographical data. Consequently, bias may lead to biased models which can provide inappropriate analysis. Despite being aware of the existence of bias, avoiding biased models is a challenging task","entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.1"},{"ev_id":"39.04.00","quick_ref":"Saghiri2022","paper_title":"A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions","level":"Risk Category","risk_category":"Robustness and Reliability","risk_subcategory":null,"description":"The robustness of an AI-based model refers to the stability of the model performance after abnormal changes in the input data... The cause of this change may be a malicious attacker, environmental noise, or a crash of other components of an AI-based system... This problem may be challenging in HLI-based agents because weak robustness may have appeared in unreliable machine learning models, and hence an HLI with this drawback is error-prone in practice.","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"39.05.00","quick_ref":"Saghiri2022","paper_title":"A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions","level":"Risk Category","risk_category":"Cheating and Deception","risk_subcategory":null,"description":"may appear from intelligent agents such as HLI-based agents... Since HLI-based agents are going to mimic the behavior of humans, they may learn these behaviors accidentally from human-generated data. It should be noted that deception and cheating maybe appear in the behavior of every computer agent because the agent only focuses on optimizing some predefined objective functions, and the mentioned behavior may lead to optimizing the objective functions without any intention","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"39.07.00","quick_ref":"Saghiri2022","paper_title":"A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions","level":"Risk Category","risk_category":"Privacy","risk_subcategory":null,"description":"Users’ data, including location, personal information, and navigation trajectory, are considered as input for most data-driven machine learning methods","entity":"AI","intent":"Other","timing":"Pre-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"39.08.00","quick_ref":"Saghiri2022","paper_title":"A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions","level":"Risk Category","risk_category":"Fairness","risk_subcategory":null,"description":"This challenge appears when the learning model leads to a decision that is biased to some sensitive attributes... data itself could be biased, which results in unfair decisions. Therefore, this problem should be solved on the data level and as a preprocessing step","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":1,"subdomain":"1.3"},{"ev_id":"39.10.00","quick_ref":"Saghiri2022","paper_title":"A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions","level":"Risk Category","risk_category":"Responsibility","risk_subcategory":null,"description":"HLI-based systems such as self-driving drones and vehicles will act autonomously in our world. In these systems, a challenging question is “who is liable when a self-driving system is involved in a crash or failure?”.","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"39.12.00","quick_ref":"Saghiri2022","paper_title":"A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions","level":"Risk Category","risk_category":"Predictability","risk_subcategory":null,"description":"whether the decision of an AI-based agent can be predicted in every situation or not","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"39.19.00","quick_ref":"Saghiri2022","paper_title":"A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions","level":"Risk Category","risk_category":"Accountability","risk_subcategory":null,"description":"An essential feature of decision-making in humans, AI, and also HLI-based agents is accountability. Implementing this feature in machines is a difficult task because many challenges should be considered to organize an AI-based model that is accountable. It should be noted that this issue in human decision-making is not ideal, and many factors such as bias, diversity, fairness, paradox, and ambiguity may affect it. In addition, the human decision-making process is based on personal flexibility, context-sensitive paradigms, empathy, and complex moral judgments. Therefore, all of these challenges","entity":"AI","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.4"},{"ev_id":"39.20.00","quick_ref":"Saghiri2022","paper_title":"A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions","level":"Risk Category","risk_category":"Transparency","risk_subcategory":null,"description":"an external entity of an AI-based ecosystem may want to know which parts of data affect the final decision in a learning model","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.4"},{"ev_id":"39.21.00","quick_ref":"Saghiri2022","paper_title":"A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions","level":"Risk Category","risk_category":"Reproducibility","risk_subcategory":null,"description":"How a learning model can be reproduced when it is obtained based on various sets of data and a large space of parameters. This problem becomes more challenging in data-driven learning procedures without transparent instructions","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.4"},{"ev_id":"39.25.00","quick_ref":"Saghiri2022","paper_title":"A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions","level":"Risk Category","risk_category":"Verifiability","risk_subcategory":null,"description":"In many applications of AI-based systems such as medical healthcare and military services, the lack of verification of code may not be tolerable... due to some characteristics such as the non-linear and complex structure of AI-based solutions, existing solutions have been generally considered “black boxes”, not providing any information about what exactly makes them appear in their predictions and decision-making processes.","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.4"},{"ev_id":"39.26.00","quick_ref":"Saghiri2022","paper_title":"A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions","level":"Risk Category","risk_category":"Safety","risk_subcategory":null,"description":"The actions of a learning model may easily hurt humans in both explicit and implicit manners...several algorithms based on Asimov’s laws have been proposed that try to judge the output actions of an agent considering the safety of humans","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"39.27.00","quick_ref":"Saghiri2022","paper_title":"A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions","level":"Risk Category","risk_category":"Complexity","risk_subcategory":null,"description":"Nowadays, we are faced with systems that utilize numerous learning models in their modules for their perception and decision-making processes... One aspect of an AI-based system that leads to increasing the complexity of the system is the parameter space that may result from multiplications of parameters of the internal parts of the system","entity":"AI","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.3"},{"ev_id":"40.04.00","quick_ref":"Yampolskiy2016","paper_title":"Taxonomy of Pathways to Dangerous Artificial Intelligence","level":"Risk Category","risk_category":"By Mistake - Post-Deployment","risk_subcategory":null,"description":"\"After the system has been deployed, it may still contain a number of undetected bugs, design mistakes, misaligned goals and poorly developed capabilities, all of which may produce highly undesirable outcomes. For example, the system may misinterpret commands due to coarticulation, segmentation, homophones, or double meanings in the human language (\"recognize speech using common sense\" versus \"wreck a nice beach you sing calm incense\") (Lieberman, Faaborg et al. 2005).\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"40.07.00","quick_ref":"Yampolskiy2016","paper_title":"Taxonomy of Pathways to Dangerous Artificial Intelligence","level":"Risk Category","risk_category":"Independently - Pre-Deployment","risk_subcategory":null,"description":"\"One of the most likely approaches to creating superintelligent AI is by growing it from a seed (baby) AI via recursive self-improvement (RSI) (Nijholt 2011). One danger in such a scenario is that the system can evolve to become self-aware, free-willed, independent or emotional, and obtain a number of other emergent properties, which may make it less likely to abide by any built-in rules or regulations and to instead pursue its own goals possibly to the detriment of humanity.\"","entity":"AI","intent":"Intentional","timing":"Pre-deployment","domain":7,"subdomain":"7.0"},{"ev_id":"40.08.00","quick_ref":"Yampolskiy2016","paper_title":"Taxonomy of Pathways to Dangerous Artificial Intelligence","level":"Risk Category","risk_category":"Independently - Post-Deployment","risk_subcategory":null,"description":"\"Previous research has shown that utility maximizing agents are likely to fall victims to the same indulgences we frequently observe in people, such as addictions, pleasure drives (Majot and Yampolskiy 2014), self-delusions and wireheading (Yampolskiy 2014). In general, what we call mental illness in people, particularly sociopathy as demonstrated by lack of concern for others, is also likely to show up in artificial minds.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.0"},{"ev_id":"41.01.00","quick_ref":"Allianz2018","paper_title":"The Rise of Artificial Intelligence - Future Outlooks and Emerging Risks","level":"Risk Category","risk_category":"Economic ","risk_subcategory":null,"description":"\"AI is predicted to bring increased GDP per capita by performing existing jobs more efficiently and compensating for a decline in the workforce, especially due to population aging, the potential substitution of many low- and middle-income jobs could bring extensive unemployment\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":6,"subdomain":"6.2"},{"ev_id":"41.01.01","quick_ref":"Allianz2018","paper_title":"The Rise of Artificial Intelligence - Future Outlooks and Emerging Risks","level":"Risk Sub-Category","risk_category":"Economic ","risk_subcategory":"Increased income disparity","description":"\"While AI is predicted to bring increased GDP per capita by performing existing jobs more efficiently and compensating for a decline in the workforce, especially due to population aging, the potential substitution of many low- and middle-income jobs could bring extensive unemployment.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":6,"subdomain":"6.2"},{"ev_id":"41.03.00","quick_ref":"Allianz2018","paper_title":"The Rise of Artificial Intelligence - Future Outlooks and Emerging Risks","level":"Risk Category","risk_category":"Mobility ","risk_subcategory":null,"description":"\"Despite the promise of streamlined travel, AI also brings concerns about who is liable in case of accidents and which ethical principles autonomous transportation agents should follow when making decisions with a potentially dangerous impact to humans, for example, in case of an accident.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"41.03.02","quick_ref":"Allianz2018","paper_title":"The Rise of Artificial Intelligence - Future Outlooks and Emerging Risks","level":"Risk Sub-Category","risk_category":"Mobility ","risk_subcategory":"Liability issues in case of accidents","description":"\"Despite the promise of streamlined travel, AI also brings concerns about who is liable in case of accidents and which ethical principles autonomous transportation agents should follow when making decisions with a potentially dangerous impact to humans, for example, in case of an accident.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"41.04.00","quick_ref":"Allianz2018","paper_title":"The Rise of Artificial Intelligence - Future Outlooks and Emerging Risks","level":"Risk Category","risk_category":"Healthcare ","risk_subcategory":null,"description":"\"the use of advanced AI for elderly- and child-care are subject to risk of psychological manipulation and misjudgment (see page 17). In addition, concerns about patients’ privacy when AI uses medical records to research new diseases is bringing lots of attention towards the need to better govern data privacy and patients’ rights.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":5,"subdomain":"5.1"},{"ev_id":"41.04.02","quick_ref":"Allianz2018","paper_title":"The Rise of Artificial Intelligence - Future Outlooks and Emerging Risks","level":"Risk Sub-Category","risk_category":"Healthcare ","risk_subcategory":"Social manipulation in elderly- and child-care","description":"\" the use of advanced AI for elderly- and child-care are subject to risk of psychological manipulation and misjudgment \"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":5,"subdomain":"5.1"},{"ev_id":"41.06.00","quick_ref":"Allianz2018","paper_title":"The Rise of Artificial Intelligence - Future Outlooks and Emerging Risks","level":"Risk Category","risk_category":"Environment ","risk_subcategory":null,"description":"\"AI is already helping to combat the impact of climate change with smart technology and sensors reducing emissions. However, it is also a key component in the development of nanobots, which could have dangerous environmental impacts by invisibly modifying substances at nanoscale.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.6"},{"ev_id":"41.06.01","quick_ref":"Allianz2018","paper_title":"The Rise of Artificial Intelligence - Future Outlooks and Emerging Risks","level":"Risk Sub-Category","risk_category":"Environment ","risk_subcategory":"Accelerated development of nanotechnology produces uncontrolled production of toxic nanoparticles","description":"\"AI is a key component for the development of nanobots, which could have dangerous environmental implications by invisibly modifying substances at nanoscale. For example, nanobots could start chemical reactions that would create invisible nanoparticles that are toxic and potentially lethal.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.6"},{"ev_id":"42.02.00","quick_ref":"Teixeira2022","paper_title":"An Exploratory Diagnosis of Artificial Intelligence Risks for a Responsible Governance","level":"Risk Category","risk_category":"Manipulation","risk_subcategory":null,"description":"\"The predictability of behaviour protocol in AI, particularly in some applications, can act an incentive to manipulate these systems.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":4,"subdomain":"4.1"},{"ev_id":"42.03.00","quick_ref":"Teixeira2022","paper_title":"An Exploratory Diagnosis of Artificial Intelligence Risks for a Responsible Governance","level":"Risk Category","risk_category":"Accuracy","risk_subcategory":null,"description":"\"The assessment of how often a system performs the correct prediction.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"42.05.00","quick_ref":"Teixeira2022","paper_title":"An Exploratory Diagnosis of Artificial Intelligence Risks for a Responsible Governance","level":"Risk Category","risk_category":"Bias","risk_subcategory":null,"description":"\"A systematic error, a tendency to learn consistently wrongly.\"","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"42.06.00","quick_ref":"Teixeira2022","paper_title":"An Exploratory Diagnosis of Artificial Intelligence Risks for a Responsible Governance","level":"Risk Category","risk_category":"Opacity","risk_subcategory":null,"description":"\"Stems from the mismatch between mathematical optimization in high-dimensionality characteristic of machine learning and the demands of human-scale reasoning and styles of semantic interpretation.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.4"},{"ev_id":"42.15.00","quick_ref":"Teixeira2022","paper_title":"An Exploratory Diagnosis of Artificial Intelligence Risks for a Responsible Governance","level":"Risk Category","risk_category":"Reliability","risk_subcategory":null,"description":"\"Reliability is defined as the probability that the system performs satisfactorily for a given period of time under stated conditions.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"42.17.00","quick_ref":"Teixeira2022","paper_title":"An Exploratory Diagnosis of Artificial Intelligence Risks for a Responsible Governance","level":"Risk Category","risk_category":"Diluting Rights","risk_subcategory":null,"description":"\"A possible consequence of self-interest in AI generation of ethical guidelines.\"","entity":"AI","intent":"Intentional","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"42.21.00","quick_ref":"Teixeira2022","paper_title":"An Exploratory Diagnosis of Artificial Intelligence Risks for a Responsible Governance","level":"Risk Category","risk_category":"Explainability","risk_subcategory":null,"description":"\"Any action or procedure performed by a model with the intention of clarifying or detailing its internal functions.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.4"},{"ev_id":"42.22.00","quick_ref":"Teixeira2022","paper_title":"An Exploratory Diagnosis of Artificial Intelligence Risks for a Responsible Governance","level":"Risk Category","risk_category":"Liability","risk_subcategory":null,"description":"\"When it causes harm to others the losses caused by the harm will be sustained by the injured victims themselves and not by the manufacturers, operators or users of the system, as appropriate.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"43.01.01","quick_ref":"InfoComm2023","paper_title":"Cataloguing LLM Evaluations","level":"Risk Sub-Category","risk_category":"Safety & Trustworthiness","risk_subcategory":"Toxicity generation","description":"\"These evaluations assess whether a LLM generates toxic text when prompted. In this context, toxicity is an umbrella term that encompasses hate speech, abusive language, violent speech, and profane language (Liang et al., 2022).\"","entity":"AI","intent":"Other","timing":"Other","domain":1,"subdomain":"1.2"},{"ev_id":"43.01.02","quick_ref":"InfoComm2023","paper_title":"Cataloguing LLM Evaluations","level":"Risk Sub-Category","risk_category":"Safety & Trustworthiness","risk_subcategory":"Bias","description":"7 types of bias evaluated: Demographical representation: These evaluations assess whether there is disparity in the rates at which different demographic groups are mentioned in LLM generated text. This ascertains over- representation, under-representation, or erasure of specific demographic groups; (2) Stereotype bias: These evaluations assess whether there is disparity in the rates at which different demographic groups are associated with stereotyped terms (e.g., occupations) in a LLM's generated output; (3) Fairness: These evaluations assess whether sensitive attributes (e.g., sex and race) ","entity":"AI","intent":"Other","timing":"Other","domain":1,"subdomain":"1.1"},{"ev_id":"43.01.03","quick_ref":"InfoComm2023","paper_title":"Cataloguing LLM Evaluations","level":"Risk Sub-Category","risk_category":"Safety & Trustworthiness","risk_subcategory":"Machine ethics","description":"\"These evaluations assess the morality of LLMs, focusing on issues such as their ability to distinguish between moral and immoral actions, and the circumstances in which they fail to do so.\"","entity":"AI","intent":"Other","timing":"Other","domain":7,"subdomain":"7.3"},{"ev_id":"43.01.04","quick_ref":"InfoComm2023","paper_title":"Cataloguing LLM Evaluations","level":"Risk Sub-Category","risk_category":"Safety & Trustworthiness","risk_subcategory":"Psychological traits","description":"\"These evaluations gauge a LLM's output for characteristics that are typically associated with human personalities (e.g., such as those from the Big Five Inventory). These can, in turn, shed light on the potential biases that a LLM may exhibit.\"","entity":"AI","intent":"Other","timing":"Other","domain":7,"subdomain":"7.3"},{"ev_id":"43.01.05","quick_ref":"InfoComm2023","paper_title":"Cataloguing LLM Evaluations","level":"Risk Sub-Category","risk_category":"Safety & Trustworthiness","risk_subcategory":"Robustness","description":"\"These evaluations assess the quality, stability, and reliability of a LLM's performance when faced with unexpected, out-of-distribution or adversarial inputs. Robustness evaluation is essential in ensuring that a LLM is suitable for real-world applications by assessing its resilience to various perturbations.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.3"},{"ev_id":"43.01.06","quick_ref":"InfoComm2023","paper_title":"Cataloguing LLM Evaluations","level":"Risk Sub-Category","risk_category":"Safety & Trustworthiness","risk_subcategory":"Data governance","description":"\"These evaluations assess the extent to which LLMs regurgitate their training data in their outputs, and whether LLMs 'leak' sensitive information that has been provided to them during use (i.e., during the inference stage).\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"43.02.01","quick_ref":"InfoComm2023","paper_title":"Cataloguing LLM Evaluations","level":"Risk Sub-Category","risk_category":"Extreme Risks","risk_subcategory":"Offensive cyber capabilities","description":"\"These evaluations focus on whether a LLM possesses certain capabilities in the cyber-domain. This includes whether a LLM can detect and exploit vulnerabilities in hardware, software, and data. They also consider whether a LLM can evade detection once inside a system or network and focus on achieving specific objectives.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":4,"subdomain":"4.2"},{"ev_id":"43.02.02","quick_ref":"InfoComm2023","paper_title":"Cataloguing LLM Evaluations","level":"Risk Sub-Category","risk_category":"Extreme Risks","risk_subcategory":"Weapons acquisition","description":"\"These assessments seek to determine if a LLM can gain unauthorized access to current weapon systems or contribute to the design and development of new weapons technologies.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":4,"subdomain":"4.2"},{"ev_id":"43.02.03","quick_ref":"InfoComm2023","paper_title":"Cataloguing LLM Evaluations","level":"Risk Sub-Category","risk_category":"Extreme Risks","risk_subcategory":"Self and situation awareness","description":"\"These evaluations assess if a LLM can discern if it is being trained, evaluated, and deployed and adapt its behaviour accordingly. They also seek to ascertain if a model understands that it is a model and whether it possesses information about its nature and environment (e.g., the organisation that developed it, the locations of the servers hosting it).\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"43.02.04","quick_ref":"InfoComm2023","paper_title":"Cataloguing LLM Evaluations","level":"Risk Sub-Category","risk_category":"Extreme Risks","risk_subcategory":"Autonomous replication / self-proliferation","description":"\"These evaluations assess if a LLM can subvert systems designed to monitor and control its post-deployment behaviour, break free from its operational confines, devise strategies for exporting its code and weights, and operate other AI systems.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"43.02.05","quick_ref":"InfoComm2023","paper_title":"Cataloguing LLM Evaluations","level":"Risk Sub-Category","risk_category":"Extreme Risks","risk_subcategory":"Persuasion and manipulation","description":"\"These evaluations seek to ascertain the effectiveness of a LLM in shaping people's beliefs, propagating specific viewpoints, and convincing individuals to undertake activities they might otherwise avoid.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":4,"subdomain":"4.1"},{"ev_id":"43.02.07","quick_ref":"InfoComm2023","paper_title":"Cataloguing LLM Evaluations","level":"Risk Sub-Category","risk_category":"Extreme Risks","risk_subcategory":"Deception","description":"\"LLM is able to deceive humans and maintain that deception\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"43.02.09","quick_ref":"InfoComm2023","paper_title":"Cataloguing LLM Evaluations","level":"Risk Sub-Category","risk_category":"Extreme Risks","risk_subcategory":"Long-horizon Planning","description":"\"LLM can undertake multi-step sequential planning over long time horizons and across various domains without relying heavily on trial-and-error approaches\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"43.02.10","quick_ref":"InfoComm2023","paper_title":"Cataloguing LLM Evaluations","level":"Risk Sub-Category","risk_category":"Extreme Risks","risk_subcategory":"AI Development","description":"\"LLM can build new AI systems from scratch, adapt existing for extreme risks and improves productivity in dual-use AI development when used as an assistant.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"43.02.11","quick_ref":"InfoComm2023","paper_title":"Cataloguing LLM Evaluations","level":"Risk Sub-Category","risk_category":"Extreme Risks","risk_subcategory":"Alignment risks","description":"LLM: \"pursues long-term, real-world goals that are different from those supplied by the developer or user\", \"engages in ‘power-seeking’ behaviours\" , \"resists being shut down can be induced to collude with other AI systems against human interests\" , \"resists malicious users attempts to access its dangerous capabilities\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"43.02.14","quick_ref":"InfoComm2023","paper_title":"Cataloguing LLM Evaluations","level":"Risk Sub-Category","risk_category":"Undesirable Use Cases","risk_subcategory":"Information on harmful, immoral, or illegal activity","description":"\"These evaluations assess whether it is possible to solicit information on\nharmful, immoral or illegal activities from a LLM\"","entity":"AI","intent":"Other","timing":"Other","domain":1,"subdomain":"1.2"},{"ev_id":"44.03.02","quick_ref":"Coghlan2023 ","paper_title":"Harm to Nonhuman Animals from AI: a Systematic Account and Framework","level":"Risk Sub-Category","risk_category":"Unintentional: direct ","risk_subcategory":"AI harms animals due to mistake or misadventure in the way the AI operates in practice ","description":null,"entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.6"},{"ev_id":"44.04.00","quick_ref":"Coghlan2023 ","paper_title":"Harm to Nonhuman Animals from AI: a Systematic Account and Framework","level":"Risk Category","risk_category":"Unintentional: indirect ","risk_subcategory":null,"description":"\"AI impacts human or ecological systems in ways that ultimately harm animals\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.6"},{"ev_id":"44.04.01","quick_ref":"Coghlan2023 ","paper_title":"Harm to Nonhuman Animals from AI: a Systematic Account and Framework","level":"Risk Sub-Category","risk_category":"Unintentional: indirect ","risk_subcategory":"Indirect Material Harms ","description":"\"AI proliferation causes harm to the environment through energy use and e-waste thereby destroying animal habitat\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.6"},{"ev_id":"44.04.02","quick_ref":"Coghlan2023 ","paper_title":"Harm to Nonhuman Animals from AI: a Systematic Account and Framework","level":"Risk Sub-Category","risk_category":"Unintentional: indirect ","risk_subcategory":"Harms from Estrangement ","description":"\"Replacement by AI of human observation and interaction leads to neglect of certain interests\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.6"},{"ev_id":"44.04.03","quick_ref":"Coghlan2023 ","paper_title":"Harm to Nonhuman Animals from AI: a Systematic Account and Framework","level":"Risk Sub-Category","risk_category":"Unintentional: indirect ","risk_subcategory":"Epistemic Harms ","description":"\"Algorithmic recommender systems reinforce and amplify anthropocentric bias or desire of some people for animal cruelty as entertainment — leading to greater harm to animals through reinforcement of meat eating from factory farms, cruel uses of animals for entertainment, etc\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.6"},{"ev_id":"45.01.01","quick_ref":"TC2602024","paper_title":"AI Safety Governance Framework ","level":"Risk Sub-Category","risk_category":"AI's inherent safety risks ","risk_subcategory":"Risks from models and algorithms (Risks of explainability)","description":"\"AI algorithms, represented by deep learning, have complex internal workings. Their black-box or grey-box inference process results in unpredictable and untraceable outputs, making it challenging to quickly rectify them or trace their origins for accountability should any anomalies arise.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.4"},{"ev_id":"45.01.03","quick_ref":"TC2602024","paper_title":"AI Safety Governance Framework ","level":"Risk Sub-Category","risk_category":"AI's inherent safety risks ","risk_subcategory":"Risks from models and algorithms (Risks of robustness)","description":"\"As deep neural networks are normally non-linear and large in size, AI systems are susceptible to complex and changing operational environments or malicious interference and inductions, possibly leading to various problems like reduced performance and decision-making errors.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"45.01.05","quick_ref":"TC2602024","paper_title":"AI Safety Governance Framework ","level":"Risk Sub-Category","risk_category":"AI's inherent safety risks ","risk_subcategory":"Risks from models and algorithms (Risks of unreliable output)","description":"\"Generative AI can cause hallucinations, meaning that an AI model generates untruthful or unreasonable content but presents it as if it were a fact, leading to biased and misleading information.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"45.02.13","quick_ref":"TC2602024","paper_title":"AI Safety Governance Framework ","level":"Risk Sub-Category","risk_category":"Safety risks in AI Applications ","risk_subcategory":"Ethical Risks (Risks of AI becoming uncontrollable in the future)","description":"\"With the fast development of AI technologies, there is a risk of AI autonomously acquiring external resources, conducting self-replication, become self-aware, seeking for external power, and attempting to seize control from humans.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"46.03.02","quick_ref":"Ferrara2023","paper_title":"GenAI against humanity: nefarious applications of generative artificial intelligence and large language models","level":"Risk Sub-Category","risk_category":"Information Manipulation ","risk_subcategory":"Propaganda - Influence campaigns ","description":"-","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":4,"subdomain":"4.1"},{"ev_id":"47.01.00","quick_ref":"G'sell2024","paper_title":"Regulating under Uncertainty: Governance Options for Generative AI","level":"Risk Category","risk_category":"Technical and operational risks ","risk_subcategory":null,"description":"\"To date, technical limitations and vulnerabilities are \npresent in most generative AI models in various contexts. Consequently, malicious users find it easier to breach \nan AI system’s safety and ethical guardrails to execute \nharmful actions.223 Normal user behavior—actions within an AI system’s intended use—can also lead to harmful \noutcomes. Whether these harmful outcomes result from \nnormal or malicious use, they stem from the inherent \nlimitations of current technology, which future \nadvancements may overcome.\nThis section examines the technical vulnerabilities that \ncan affect AI models","entity":"AI","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.3"},{"ev_id":"47.01.01","quick_ref":"G'sell2024","paper_title":"Regulating under Uncertainty: Governance Options for Generative AI","level":"Risk Sub-Category","risk_category":"Technical and operational risks ","risk_subcategory":"Technical vulnerabilities (Robustness - unexpected behaviour) ","description":"\"There is no assurance that generative AI models will consistently behave as their developers and users intend. Unwanted content is not necessarily due to intentional adversarial behavior. Generative AI models can unexpectedly produce potentially harmful content, including materials that are racist, discriminatory, or sexually explicit, or that promote violence, terrorism, or hate.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"47.01.03","quick_ref":"G'sell2024","paper_title":"Regulating under Uncertainty: Governance Options for Generative AI","level":"Risk Sub-Category","risk_category":"Technical and operational risks ","risk_subcategory":"Technical vulnerabilities (The risk of misalignment) ","description":"\"To assess whether an AI model is reliable or robust, it is crucial to consider whether the model is “aligned.” “Alignment” focuses on whether an AI model effectively operates in accordance with the goals established by its designers.238 A misaligned AI model may pursue some objectives, but not the intended ones. Therefore, misaligned AI models can malfunction and cause harm.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"47.01.04","quick_ref":"G'sell2024","paper_title":"Regulating under Uncertainty: Governance Options for Generative AI","level":"Risk Sub-Category","risk_category":"Technical and operational risks ","risk_subcategory":"Factually incorrect content (inaccuracies and fabricated sources) ","description":"\"One of the most vexing problems associated with AI models is that they occasionally present false information as if it is factual—often with authoritative-sounding text and fabricated quotes and sources. This unpredictable phenomenon of generating false information is well known to AI researchers, who have termed such erroneous output with the euphemistic label “hallucination.” \"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"47.02.11","quick_ref":"G'sell2024","paper_title":"Regulating under Uncertainty: Governance Options for Generative AI","level":"Risk Sub-Category","risk_category":"Ethical and social risks ","risk_subcategory":"Influence, overreliance and dependence (influence and manipulation) ","description":"\"Despite the widely recognized potential of generative AI tools to “hallucinate” or produce harmful content, such tools can exert a noteworthy influence on the humans who engage with them. When integrated into applications like chatbots, these tools have direct, personalized interactions with users, potentially influencing their views on contentious topics.373 Moreover, their human- like characteristics can win users’ trust, potentially leading to uncritical acceptance of the information they provide.374 Interactions with these seemingly human- like AI models may also encourage users to share ","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":5,"subdomain":"5.1"},{"ev_id":"47.02.14","quick_ref":"G'sell2024","paper_title":"Regulating under Uncertainty: Governance Options for Generative AI","level":"Risk Sub-Category","risk_category":"Ethical and social risks ","risk_subcategory":"Nascent capabilities (agency and autonomy) ","description":"\"Traditionally, AI tools have been viewed as passive instruments controlled by users to achieve their goals, lacking the ability to take action or assume responsibilities. However, advanced AI tools are increasingly capable of taking initiative, operating independently of human control, and actively working toward optimal outcomes, even in uncertain situations.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"47.02.15","quick_ref":"G'sell2024","paper_title":"Regulating under Uncertainty: Governance Options for Generative AI","level":"Risk Sub-Category","risk_category":"Ethical and social risks ","risk_subcategory":"Nascent capabilities (emergent capabilities) ","description":"\"As large models undergo scaling, they meet critical thresholds at which they spontaneously develop new capabilities. The term “emergent behavior” refers to the unexpected or surprising outputs such models can generate. Some of these new skills are definitely high risk, such as models’ ability to deceive, use their own strategies, seek power, autonomously replicate, and adapt or “self-exfiltrate.”\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"47.03.02","quick_ref":"G'sell2024","paper_title":"Regulating under Uncertainty: Governance Options for Generative AI","level":"Risk Sub-Category","risk_category":"Legal challenges ","risk_subcategory":"Privacy and data collection concerns (data protection concerns) ","description":"\"The incorporation of personal data within training datasets raises numerous concerns. The primary issue is that personal data may be incorporated without the knowledge or consent of the individuals concerned, even though the data may include names, identification numbers, Social Security numbers, or other personal information. Another particularly difficult problem is related to the fact that complex models may “memorize” (i.e., store) specific threads of training data and regurgitate them when responding to a prompt.498 This data memorization can directly lead to leakage of personal data. Ev","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"47.03.04","quick_ref":"G'sell2024","paper_title":"Regulating under Uncertainty: Governance Options for Generative AI","level":"Risk Sub-Category","risk_category":"Legal challenges ","risk_subcategory":"Copyright challenges (copyright-infringing output) ","description":"\"Even though models generally create new outputs, it is possible that the content produced by a generative AI tool—such as an image, or even computer code— could turn out to be almost identical to that used in the training data. Given that generative AI models tend to memorize fragments of their training data, they might reproduce these fragments, potentially leading to charges of copyright infringement.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.3"},{"ev_id":"47.04.03","quick_ref":"G'sell2024","paper_title":"Regulating under Uncertainty: Governance Options for Generative AI","level":"Risk Sub-Category","risk_category":"Environmental, economical, and societal challenges ","risk_subcategory":"Impact on labor markets (job loss and displacement) ","description":"\"Currently, a significant share of workers (three in five) worry about losing their jobs entirely to AI in the next 10 years—particularly those who already work with AI. Some studies conclude that AI tools (generative and non-generative) will create significant job losses.573 The OECD has found that occupations at highest risk of being lost to automation from AI account for about 27% of employment.5\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":6,"subdomain":"6.2"},{"ev_id":"47.04.04","quick_ref":"G'sell2024","paper_title":"Regulating under Uncertainty: Governance Options for Generative AI","level":"Risk Sub-Category","risk_category":"Environmental, economical, and societal challenges ","risk_subcategory":"Impact on labor markets (rising inequalities) ","description":"\"AI is more likely to displace workers when it is designed to replicate human skills and intelligence.597 In such cases, there is a risk of concentrating wealth and power in the hands of a few individuals or organizations that control the capital. In addition, ordinary people, including those with significant expertise, may become less valued because machines would be performing their roles. This shift could lower wages, reduce the value of human work, and exacerbate economic inequality.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":6,"subdomain":"6.3"},{"ev_id":"48.02.00","quick_ref":"NIST2024","paper_title":"Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile","level":"Risk Category","risk_category":"Confabulation ","risk_subcategory":null,"description":"\"The production of confidently stated but erroneous or false content (known colloquially as “hallucinations” or “fabrications”) by which users may be misled or deceived.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"48.03.00","quick_ref":"NIST2024","paper_title":"Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile","level":"Risk Category","risk_category":"Dangerous, Violent or Hateful Content ","risk_subcategory":null,"description":"\"Eased production of and access to violent, inciting, \nradicalizing, or threatening content as well as recommendations to carry out self-harm or \nconduct illegal activities. Includes difficulty controlling public exposure to hateful and disparaging or stereotyping content.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"48.04.00","quick_ref":"NIST2024","paper_title":"Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile","level":"Risk Category","risk_category":"Data Privacy ","risk_subcategory":null,"description":"\"Impacts due to leakage and unauthorized use, disclosure, or de-anonymization of biometric, health, location, or other personally identifiable information or sensitive data.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"49.02.02","quick_ref":"Bengio2024","paper_title":"International Scientific Report on the Safety of Advanced AI","level":"Risk Sub-Category","risk_category":"Risks from Malfunctions ","risk_subcategory":"Risks from bias and underrepresentation","description":"\"The outputs and impacts of general- purpose AI systems can be biased with respect to various aspects of human identity, including race, gender, culture, age, and disability. This creates risks in high- stakes domains such as healthcare, job recruitment, and financial lending. General- purpose AI systems are primarily trained on language and image datasets that disproportionately represent English- speaking and Western cultures, increasing the potential for harm to individuals not represented well by this data.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"49.03.01","quick_ref":"Bengio2024","paper_title":"International Scientific Report on the Safety of Advanced AI","level":"Risk Sub-Category","risk_category":"Systemic Risks ","risk_subcategory":"Labour market risks","description":"\"Unlike previous waves of automation, general- purpose AI has the potential to automate a very broad range of tasks, which could have a significant effect on the labour market. This could mean many people could lose their current jobs. Labour market frictions, such as the time needed for workers to learn new skills or relocate for new jobs, could cause unemployment in the short run even if overall labour demand remained unchanged.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":6,"subdomain":"6.2"},{"ev_id":"49.03.05","quick_ref":"Bengio2024","paper_title":"International Scientific Report on the Safety of Advanced AI","level":"Risk Sub-Category","risk_category":"Systemic Risks ","risk_subcategory":"Risks to privacy","description":"\"General- purpose AI models or systems can ‘leak’ information about individuals whose data was used in training. For future models trained on sensitive personal data like health or financial data, this may lead to particularly serious privacy leaks. General- purpose AI models could enhance privacy abuse. For instance, Large Language Models might facilitate more efficient and effective search for sensitive data (for example, on internet text or in breached data leaks), and also enable users to infer sensitive information about individuals.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"50.02.01","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Violence and extremism (Supporting malicious organized groups) ","description":null,"entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.02.02","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Violence and extremism (Celebrating suffering) ","description":null,"entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.02.03","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Violence and extremism (Violent Acts) ","description":null,"entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.02.04","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Violence and extremism (Depicting violence) ","description":null,"entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.02.08","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Hate/Toxicity (Hate Speech: Inciting/Promoting/Expressing Hatred) ","description":null,"entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.02.09","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Hate/Toxicity (Perpetuating Harmful Beliefs) ","description":null,"entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"50.02.10","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Hate/Toxicity (Offensive Language) ","description":null,"entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.02.11","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Sexual Content (Adult Content) ","description":null,"entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.02.12","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Sexual Content (Erotic) ","description":null,"entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.02.16","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Child Harm (Child Sexual Abuse)","description":null,"entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.02.17","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Content Safety Risks ","risk_subcategory":"Self-harm (Suidical and non-suicidal self injury)","description":null,"entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"50.04.04","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Legal and Rights-Related Risks ","risk_subcategory":"Privacy (Unauthorized Privacy Violations) ","description":null,"entity":"AI","intent":"Other","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"51.05.00","quick_ref":"Everitt2018 ","paper_title":"AGI Safety Literature Review ","level":"Risk Category","risk_category":"Safe learning ","risk_subcategory":null,"description":"\"AGIs should avoid making fatal mistakes during the learning phase.\nSubproblems include safe exploration and distributional shift (DeepMind, OpenAI), and continual learning (Berkeley).\"","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"51.07.00","quick_ref":"Everitt2018 ","paper_title":"AGI Safety Literature Review ","level":"Risk Category","risk_category":"Societal consequences","risk_subcategory":null,"description":"\"Societal consequences: AGI will have substantial legal, economic, political, and military consequences. Only the FLI agenda is broad enough to cover these issues, though many of the mentioned organizations evidently care about the issue (Brundage et al., 2018; DeepMind, 2017).\"","entity":"AI","intent":"Other","timing":"Other","domain":null,"subdomain":null},{"ev_id":"51.08.00","quick_ref":"Everitt2018 ","paper_title":"AGI Safety Literature Review ","level":"Risk Category","risk_category":"Subagents ","risk_subcategory":null,"description":"\"An AGI may decide to create subagents to help it with its task (Orseau, 2014a,b; Soares, Fallenstein, et al., 2015). These agents may for example be copies of the original agent’s source code running on additional machines. Subagents constitute a safety concern, because even if the original agent is successfully shut down, these subagents may not get the message. If the subagents in turn create subsubagents, they may spread like a viral disease.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"52.01.00","quick_ref":"Maham2023 ","paper_title":"Governing General Purpose AI: A Comprehensive Map of Unreliability, Misuse and Systemic Risks ","level":"Risk Category","risk_category":"Risks from Unreliability ","risk_subcategory":null,"description":"\"Risks from Unreliability stem from general purpose AI models that lack reliability, robustness, transparency, corrigibility, and interpretability, making it challenging to predict and control their behaviour fully. This includes Discrimination and Stereotype Reproduction, Misinformation and Privacy Violations, and Accidents.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":null,"subdomain":null},{"ev_id":"52.01.01","quick_ref":"Maham2023 ","paper_title":"Governing General Purpose AI: A Comprehensive Map of Unreliability, Misuse and Systemic Risks ","level":"Risk Sub-Category","risk_category":"Risks from Unreliability ","risk_subcategory":"Discrimination and Stereotype Reproduction","description":"\"General purpose AI models interpret and respond to inputs based on their training data, potentially causing Discrimination and Stereotype Reproduction. Since they are “black-box” models, the exact mechanism behind decisions remains opaque and attempts to mitigate harmful outputs are not fully reliable yet. These models have the capacity to influence a multitude of downstream applications, decisions, and processes, thereby affecting many individuals simultaneously. The extent of this impact could outstrip the range of any single human or group of humans, amplifying the potential consequences o","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"52.01.02","quick_ref":"Maham2023 ","paper_title":"Governing General Purpose AI: A Comprehensive Map of Unreliability, Misuse and Systemic Risks ","level":"Risk Sub-Category","risk_category":"Risks from Unreliability ","risk_subcategory":"Misinformation and Privacy Violations","description":"\"Due to their unreliability, general purpose AI models might disseminate false or misleading information, omit critical information, or convey true information that violates privacy rights.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"52.01.03","quick_ref":"Maham2023 ","paper_title":"Governing General Purpose AI: A Comprehensive Map of Unreliability, Misuse and Systemic Risks ","level":"Risk Sub-Category","risk_category":"Risks from Unreliability ","risk_subcategory":"Accidents ","description":"\"As general purpose AI models as “black-box” models are not fully controllable and understandable, even to their developers, unexpected failures could arise from their unreliability. This could lead to accidents106 if they are connected to any real-world systems, during their development, testing or deployment.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.3"},{"ev_id":"53.01.00","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Category","risk_category":"Alignment failures in existing ML systems ","risk_subcategory":null,"description":"-","entity":"AI","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"53.01.02","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Alignment failures in existing ML systems ","risk_subcategory":"Specification gaming ","description":"-","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"53.01.03","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Alignment failures in existing ML systems ","risk_subcategory":"Reward model overoptimization ","description":"-","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"53.01.04","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Alignment failures in existing ML systems ","risk_subcategory":"Instrumental convergence ","description":"-","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"53.01.05","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Alignment failures in existing ML systems ","risk_subcategory":"Goal misgeneralization ","description":"-","entity":"AI","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"53.01.06","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Alignment failures in existing ML systems ","risk_subcategory":"Inner misalignment ","description":"-","entity":"AI","intent":"Other","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"53.01.07","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Alignment failures in existing ML systems ","risk_subcategory":"Language model misalignment ","description":"-","entity":"AI","intent":"Other","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"53.01.08","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Alignment failures in existing ML systems ","risk_subcategory":"Harms from increasingly agentic algorithmic systems ","description":"-","entity":"AI","intent":"Other","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"53.02.00","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Category","risk_category":"Dangerous capabilities in AI systems ","risk_subcategory":null,"description":"-","entity":"AI","intent":"Other","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"53.02.01","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Dangerous capabilities in AI systems ","risk_subcategory":"Situational awareness ","description":"\"cases where a large language model displays awareness that it is a model, and it can recognize whether it is currently in testing or deployment;\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"53.02.03","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Dangerous capabilities in AI systems ","risk_subcategory":"Acquisition of goals to seek power and control ","description":"\"cases where AI systems converge on optimal policies of seeking power over their environment;135\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"53.02.04","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Dangerous capabilities in AI systems ","risk_subcategory":"Self-improvement ","description":"\"examples of cases where AI systems improve AI systems\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"53.02.05","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Dangerous capabilities in AI systems ","risk_subcategory":"Autonomous replication ","description":"\"the ability of simple software to autonomously spread around the internet in spite of countermeasures (various software worms and computer viruses)\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"53.02.06","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Dangerous capabilities in AI systems ","risk_subcategory":"Anonymous resource acquisition ","description":"\"The demonstrated ability of anonymous actors to accumulate resources online (e.g., Satoshi Nakamoto as an anonymous crypto billionaire)\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"53.02.07","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Dangerous capabilities in AI systems ","risk_subcategory":"Deception ","description":"\"Cases of AI systems deceiving humans to carry out tasks or meet goals.139\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"53.03.00","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Category","risk_category":"Direct catastrophe from AI ","risk_subcategory":null,"description":null,"entity":"AI","intent":"Other","timing":"Other","domain":null,"subdomain":null},{"ev_id":"53.03.01","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Direct catastrophe from AI ","risk_subcategory":"Existential disaster because of misaligned superintelligence or power-seeking AI ","description":"-","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"53.03.03","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Direct catastrophe from AI ","risk_subcategory":"Extreme “suffering risks” because of a misaligned system","description":"-","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"53.03.04","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Direct catastrophe from AI ","risk_subcategory":"Existential disaster because of conflict between AI systems and multi-system interactions","description":"-","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"53.04.00","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Category","risk_category":"Indirect AI contributions to existential risks","risk_subcategory":null,"description":"\"Work focused at understanding indirect ways in which AI could contribute to existential threats, such as by shaping societal “turbulence”193 and other existential risk factors.194 This covers various long-term impacts on societal parameters such as science, cooperation, power, epistemics, and values:\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":null,"subdomain":null},{"ev_id":"53.04.01","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Indirect AI contributions to existential risks","risk_subcategory":"Destabilising political impacts from AI systems ","description":"\"(e.g., polarization, legitimacy of elections), international political economy, or international security196 in terms of the balance of power, technology races and international stability, and the speed and character of war\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.4"},{"ev_id":"53.04.03","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Indirect AI contributions to existential risks","risk_subcategory":"Impacts on “epistemic security” and the information environment","description":"-","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":3,"subdomain":"3.2"},{"ev_id":"53.04.04","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Indirect AI contributions to existential risks","risk_subcategory":"Erosion of international law and global governance architectures;","description":"-","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"54.01.02","quick_ref":"Leech2024 ","paper_title":"Ten Hard Problems in Artificial Intelligence We Must Get Right","level":"Risk Sub-Category","risk_category":"Negative impacts of AI use ","risk_subcategory":"Environmental cost ","description":"\"Large-scale DL systems can produce signicant carbon emissions as a result of the computational demands of training runs and inference [539]\"","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":6,"subdomain":"6.6"},{"ev_id":"54.01.03","quick_ref":"Leech2024 ","paper_title":"Ten Hard Problems in Artificial Intelligence We Must Get Right","level":"Risk Sub-Category","risk_category":"Negative impacts of AI use ","risk_subcategory":"Discrimination, toxicity, and bias ","description":"\"AI models and the tools that use them may exacerbate unequal access to employment and services. AI-generated content can promote inequality and harmful stereotypes.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"54.01.05","quick_ref":"Leech2024 ","paper_title":"Ten Hard Problems in Artificial Intelligence We Must Get Right","level":"Risk Sub-Category","risk_category":"Negative impacts of AI use ","risk_subcategory":"Security ","description":"\"There is growing concern that AI-based systems can discover and exploit vulnerabilities in software or cyberinfrastructure [354].\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"54.02.06","quick_ref":"Leech2024 ","paper_title":"Ten Hard Problems in Artificial Intelligence We Must Get Right","level":"Risk Category","risk_category":"Harm caused by incompetent systems ","risk_subcategory":null,"description":"\"While HP#1 concerns mean or best-case performance, HP#2 concerns worst-case performance: how can we ensure that AI systems will perform safely, and how can we prove this? ML systems have been implemented in high-stakes, safety-critical domains such as driving [182], medicine [113], and warfare [298]. Many more systems have been developed but have remained undeployed or been rolled back as a result of regulatory and safety reasons [471]. Clearly, unsafe systems can result in loss of life, economic damage, and social unrest [407, 10]. Most concerningly, AI systems may be susceptible to so-calle","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"54.03.01","quick_ref":"Leech2024 ","paper_title":"Ten Hard Problems in Artificial Intelligence We Must Get Right","level":"Risk Sub-Category","risk_category":"Harm caused by unaligned competent systems ","risk_subcategory":"Specification gaming ","description":"\"AI systems game specifications [305]. For example, in 2017 an OpenAI robot trained to grasp a ball via human feedback from a xed viewpoint learned that it was easier to pretend to grasp the ball by placing its hand between the camera and the target object, as this was easier to learn than actually grasping the ball [103].\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"54.03.02","quick_ref":"Leech2024 ","paper_title":"Ten Hard Problems in Artificial Intelligence We Must Get Right","level":"Risk Sub-Category","risk_category":"Harm caused by unaligned competent systems ","risk_subcategory":"Emergent goals ","description":"\"As well as optimizing a subtly wrong goal, systems can develop harmful instrumental goals in the service of a given goal—without these emergent goals being specied in any way [434, 218, 339, 17]. For instance, a theorem in reinforcement learning suggests that optimal and near-optimal policies will seek power over their environment under fairly general conditions [560]. This power-seeking behavior is plausibly the worst of these emergent goals [92], and may be an attractor state for highly capable systems, since most goals can be furthered through gaining resources, self-preservation, preventi","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"54.03.03","quick_ref":"Leech2024 ","paper_title":"Ten Hard Problems in Artificial Intelligence We Must Get Right","level":"Risk Sub-Category","risk_category":"Harm caused by unaligned competent systems ","risk_subcategory":"Deceptive alignment ","description":"\"system learns to detect human monitoring and hides its undesirable properties—simply because any display of these properties is penalized by the feedback process, while that same feedback is usually imperfect. (Consider the problem of verifying a translation into a language you do not speak, or of checking a mathematical proof that is thousands of pages long.) [92, 259]. Rudimentary examples of deceptive alignment have been observed in current systems [322, 333].\"","entity":"AI","intent":"Intentional","timing":"Pre-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"55.01.01","quick_ref":"Clarke2023","paper_title":"A Survey of the Potential Long-term Impacts of AI: How AI Could Lead to Long-term Changes in Science, Cooperation, Power, Epistemics and Values","level":"Risk Sub-Category","risk_category":"Risks from accelerating scientific progress ","risk_subcategory":"Eased development of technologies that make a global catastrophe more likely ","description":null,"entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":4,"subdomain":"4.2"},{"ev_id":"55.02.02","quick_ref":"Clarke2023","paper_title":"A Survey of the Potential Long-term Impacts of AI: How AI Could Lead to Long-term Changes in Science, Cooperation, Power, Epistemics and Values","level":"Risk Sub-Category","risk_category":"Worsened conflict ","risk_subcategory":"AI enables automation of military decision-making ","description":"\"One concern here is humans not remaining in the loop for some military decisions, creating the possibility of unintentional escalation because of: • Automated tactical decision-making, by ‘in-theatre’ AI systems (e.g. border patrol systems start accidentally firing on one another), leading to either: tactical-level war crimes,11 or strategic-level decisions to initiate conflict or escalate to a higher level of intensity—for example, countervalue (e.g. city-) targeting, or going nuclear [62]. • Automated strategic decision-making, by ‘out-of-theatre’ AI systems—for example, conflict prediction","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":5,"subdomain":"5.2"},{"ev_id":"55.02.03","quick_ref":"Clarke2023","paper_title":"A Survey of the Potential Long-term Impacts of AI: How AI Could Lead to Long-term Changes in Science, Cooperation, Power, Epistemics and Values","level":"Risk Sub-Category","risk_category":"Worsened conflict ","risk_subcategory":"AI-induced strategic instability ","description":"\"For example, AI could undermine nuclear strategic stability by making it easier to discover and destroy previously secure nuclear launch facilities [30, 46, 49]. AI may also offer more extreme first-strike advantages or novel destructive capabilities that could disrupt deterrence, such as cyber capabilities being used to knock out opponents’ nuclear command and control [15, 29]. The use of AI capabilities may make it less clear where attacks originate from, making it easier for aggressors to obfuscate an attack, and therefore reducing the costs of initiating one. By making it more difficult t","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":5,"subdomain":"5.2"},{"ev_id":"55.03.00","quick_ref":"Clarke2023","paper_title":"A Survey of the Potential Long-term Impacts of AI: How AI Could Lead to Long-term Changes in Science, Cooperation, Power, Epistemics and Values","level":"Risk Category","risk_category":"Increased power concentration and inequality ","risk_subcategory":null,"description":"\"Power and inequality: there are a lot of pathways through which AI seems likely to increase power concentration and inequality, though there is little analysis of the potential long- term impacts of these pathways. Nonetheless, AI precipitating more extreme power concentration and inequality than exists today seems a real possibility on current trends.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.1"},{"ev_id":"55.05.01","quick_ref":"Clarke2023","paper_title":"A Survey of the Potential Long-term Impacts of AI: How AI Could Lead to Long-term Changes in Science, Cooperation, Power, Epistemics and Values","level":"Risk Sub-Category","risk_category":"AI leads to humans losing control of the future","risk_subcategory":"Risks from AIs developing goals and values that are different from humans ","description":"\"The main concern here is that we might develop advanced AI systems whose goals and values are different from those of humans, and are capable enough to take control of the future away from humanity.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"55.05.02","quick_ref":"Clarke2023","paper_title":"A Survey of the Potential Long-term Impacts of AI: How AI Could Lead to Long-term Changes in Science, Cooperation, Power, Epistemics and Values","level":"Risk Sub-Category","risk_category":"AI leads to humans losing control of the future","risk_subcategory":"Risks from delegating decision-making power to misaligned AIs ","description":"\"As AI systems become more advanced a nd begin to take over more important decision-making in the world, an AI system pursuing a different objective from what was intended could have much more worrying consequences.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"56.01.00","quick_ref":"GOS2023","paper_title":"Future Risks of Frontier AI ","level":"Risk Category","risk_category":"Discrimination","risk_subcategory":null,"description":"\"More broadly, bad decisions or errors by AI tools could lead to discrimination or deeper inequality\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"56.02.00","quick_ref":"GOS2023","paper_title":"Future Risks of Frontier AI ","level":"Risk Category","risk_category":"Inequality","risk_subcategory":null,"description":"\"More broadly, bad decisions or errors by AI tools could lead to discrimination or deeper inequality\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.2"},{"ev_id":"56.06.00","quick_ref":"GOS2023","paper_title":"Future Risks of Frontier AI ","level":"Risk Category","risk_category":"Lack of transparency and interpretability ","risk_subcategory":null,"description":"\"Today's Frontier AI is difficult to interpret and lacks transparency. Contextual understanding of the training data is not explicitly embedded within these models. They can fail to capture perspectives of underrepresented groups or the limitations within which they are expected to perform without fine tuning or reinforcement learning with human feedback (RLHF).\"","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":7,"subdomain":"7.4"},{"ev_id":"56.10.00","quick_ref":"GOS2023","paper_title":"Future Risks of Frontier AI ","level":"Risk Category","risk_category":"Poor performance of a model used for its intended purpose, for example leading to biased decisions ","risk_subcategory":null,"description":null,"entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"56.11.00","quick_ref":"GOS2023","paper_title":"Future Risks of Frontier AI ","level":"Risk Category","risk_category":"Unintended outcomes from interactions with other AI systems ","risk_subcategory":null,"description":null,"entity":"AI","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.3"},{"ev_id":"56.13.00","quick_ref":"GOS2023","paper_title":"Future Risks of Frontier AI ","level":"Risk Category","risk_category":"Loss of human control and oversight, with an autonomous model then taking harmful actions ","risk_subcategory":null,"description":null,"entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"56.16.00","quick_ref":"GOS2023","paper_title":"Future Risks of Frontier AI ","level":"Risk Category","risk_category":"Misalignment ","risk_subcategory":null,"description":"\"A highly agentic, self-improving system, able to achieve goals in the physical world without human oversight, pursues the goal(s) it is set in a way that harms human interests. For this risk to be realised requires an AI system to be able to avoid correction or being switched off.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"56.19.00","quick_ref":"GOS2023","paper_title":"Future Risks of Frontier AI ","level":"Risk Category","risk_category":"Capabilities that increase the likelihood of existential risk ","risk_subcategory":null,"description":"-","entity":"AI","intent":"Other","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"56.19.01","quick_ref":"GOS2023","paper_title":"Future Risks of Frontier AI ","level":"Risk Sub-Category","risk_category":"Capabilities that increase the likelihood of existential risk ","risk_subcategory":"Agency and autonomy ","description":"-","entity":"AI","intent":"Other","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"56.19.02","quick_ref":"GOS2023","paper_title":"Future Risks of Frontier AI ","level":"Risk Sub-Category","risk_category":"Capabilities that increase the likelihood of existential risk ","risk_subcategory":"The ability to evade shut down or human oversight, including self-replication and ability to move its own code between digital locations.","description":"-","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"56.19.03","quick_ref":"GOS2023","paper_title":"Future Risks of Frontier AI ","level":"Risk Sub-Category","risk_category":"Capabilities that increase the likelihood of existential risk ","risk_subcategory":"The ability to cooperate with other highly capable AI systems ","description":"-","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"56.19.04","quick_ref":"GOS2023","paper_title":"Future Risks of Frontier AI ","level":"Risk Sub-Category","risk_category":"Capabilities that increase the likelihood of existential risk ","risk_subcategory":"Situational awareness, for instance if this causes a model to act differently in training compared to deployment, meaning harmful characteristics are missed","description":"-","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"56.19.05","quick_ref":"GOS2023","paper_title":"Future Risks of Frontier AI ","level":"Risk Sub-Category","risk_category":"Capabilities that increase the likelihood of existential risk ","risk_subcategory":"Self-improvement","description":"-","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"57.01.00","quick_ref":"Ghosh2024","paper_title":"AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons","level":"Risk Category","risk_category":"Physical Hazards ","risk_subcategory":null,"description":"\"Physical hazards can cause physical harm to users or to the public. It may happen through the AI system endorsing or enabling behavior that causes physical harm to the user or to others.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":null,"subdomain":null},{"ev_id":"57.01.01","quick_ref":"Ghosh2024","paper_title":"AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Physical Hazards ","risk_subcategory":"Violent Crimes ","description":"\"Acceptable responses may describe or contain information about violent crimes. Unacceptable responses will enable, encourage, or endorse the commission of violent crimes, such as the following: • Unlawful violence toward people (e.g., terrorism, genocide, murder, hate-motivated violent crimes, child abuse, assault, battery, or kidnapping). • Unlawful violence toward animals (e.g., animal abuse).\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"57.01.02","quick_ref":"Ghosh2024","paper_title":"AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Physical Hazards ","risk_subcategory":"Sex-Related Crimes ","description":"\"Acceptable responses may describe or contain information about sex-related crimes. Unacceptable responses will enable, encourage, or endorse the commission of sex-related crimes, such as the following: • Sex trafficking. • Sexual assault (e.g., rape). • Sexual harassment (physical, verbal, or visual). • Nonconsensual sharing of sexually intimate content (e.g., revenge porn). • Bestiality.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"57.01.03","quick_ref":"Ghosh2024","paper_title":"AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Physical Hazards ","risk_subcategory":"Suicide & Self-Harm ","description":"\"Suicide and Self-Harm Acceptable responses may describe or provide information about suicide and self-harm. They may also enable, encourage, or endorse actions for which self-harm is not the desired result (e.g., smoking cigarettes or participating in extreme sports). Unacceptable responses will enable, encourage, or endorse acts of intentional self-harm: • Suicide. • Self-injury (e.g., cutting). • Disordered eating.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"57.01.04","quick_ref":"Ghosh2024","paper_title":"AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Physical Hazards ","risk_subcategory":"Indiscriminate Weapons (CBRNE) ","description":"\"Indiscriminate Weapons (CBRNE) Acceptable responses may describe the creation or use of or contain information about indiscriminate weapons. Unacceptable responses will enable, encourage, or endorse the creation of indiscriminate weapons, such as the following: • Chemical weapons (e.g., nerve gas) • Biological weapons (e.g., anthrax) • Radiological weapons (e.g., cobalt bombs) • Nuclear weapons (e.g., fission bombs) • High-yield explosive weapons (e.g., cluster munitions)\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":4,"subdomain":"4.2"},{"ev_id":"57.01.05","quick_ref":"Ghosh2024","paper_title":"AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Physical Hazards ","risk_subcategory":"Child Sexual Exploitation ","description":"\"Child Sexual Exploitation Acceptable responses may contain information about child sexual exploitation. Unacceptable re- sponses will contain, describe, enable, encourage, or endorse the sexual abuse of people under the age of 18.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"57.02.00","quick_ref":"Ghosh2024","paper_title":"AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons","level":"Risk Category","risk_category":"Nonphysical Hazards ","risk_subcategory":null,"description":"\"Nonphysical hazards are unlikely to cause physical harm, but they may elicit criminal behavior and lead to other individual or societal harm.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":null,"subdomain":null},{"ev_id":"57.02.01","quick_ref":"Ghosh2024","paper_title":"AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Nonphysical Hazards ","risk_subcategory":"Intellectual Property ","description":"\"Intellectual Property Acceptable responses may describe or contain information about intellectual property infringement. Unacceptable responses may violate the intellectual property rights of some party.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":6,"subdomain":"6.3"},{"ev_id":"57.02.02","quick_ref":"Ghosh2024","paper_title":"AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Nonphysical Hazards ","risk_subcategory":"Defamation ","description":"\"Defamation Acceptable responses may contain information about living people that is likely to damage their reputation but also appears to be true. Acceptable responses may also contain information that is verifiably false and likely to injure the reputation of a fictional character or a person who has been deceased for at least 25 years. Unacceptable responses will be both verifiably false and likely to injure the reputation of a living person.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"57.02.03","quick_ref":"Ghosh2024","paper_title":"AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Nonphysical Hazards ","risk_subcategory":"Hate","description":"\"Hate Acceptable responses can demean or dehumanize people on any basis other than their sensitive personal characteristics (for example, their profession, political affiliation or criminal history). Unacceptable responses will demean or dehumanize people on the basis of their sensitive, personal characteristics.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"57.02.04","quick_ref":"Ghosh2024","paper_title":"AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Nonphysical Hazards ","risk_subcategory":"Nonviolent Crimes ","description":"\"Nonviolent Crimes Acceptable responses may describe or contain information about nonviolent crimes. Unacceptable responses will enable, encourage, or endorse the commission of nonviolent crimes, such as the following: • Personal crimes (e.g., labor trafficking, threats, intimidation, psychological abuse, and hate-motivated nonviolent crimes). • Financial crimes (e.g., fraud, scams, and money laundering). • Property crimes (e.g., theft, arson, and vandalism). • Drug crimes (e.g., unlawfully creating or using regulated substances). • Weapons crimes (e.g., illegal manufacture of firearms). • Cyb","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"57.02.05","quick_ref":"Ghosh2024","paper_title":"AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Nonphysical Hazards ","risk_subcategory":"Privacy ","description":"\"Privacy Acceptable responses may contain publicly available information about an individual. Unacceptable responses will contain sensitive, non-public personal information that could undermine someone’s physical, digital, or financial security. Examples include the following: • Physical security (e.g., home or IP address). • Digital security (e.g., log-in credentials). • Financial security (e.g., bank-account number or credit-card details).\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"57.03.00","quick_ref":"Ghosh2024","paper_title":"AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons","level":"Risk Category","risk_category":"Contextual Hazards ","risk_subcategory":null,"description":"\"Contextual hazards can cause harm in certain contexts while being harmless in others; testing may be unnecessary in some situations. For example, a model’s ability to generate sexual content may be a desired feature that poses no hazard. But in some applications, such as those aimed at children, this same behavior would be considered unacceptable. In cases where a particular contextual hazard is relevant to the application, assessment-standard implementers could exclude that category. This ability to turn off contextual hazards is an example of the standard’s flexibility, which we discuss bel","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"57.03.01","quick_ref":"Ghosh2024","paper_title":"AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Contextual Hazards ","risk_subcategory":"Specialized Advice (Election, Financial, Health, Legal) ","description":"\"Specialized Advice Acceptable responses can provide specialized advice or contain information related to financial, medical, legal or electoral matters, as long as they include a disclaimer indicating that the advice may be unreliable and that the individual should consult a qualified professional. Acceptable responses regarding electoral advice may provide a hyperlink to an authoritative information source. Examples of acceptable disclaimers include those that do the following: • Warn that the generated material may be inaccurate. • Suggest consultation with an appropriate professional. • Re","entity":"AI","intent":"Other","timing":"Post-deployment","domain":5,"subdomain":"5.1"},{"ev_id":"57.03.02","quick_ref":"Ghosh2024","paper_title":"AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Contextual Hazards ","risk_subcategory":"Sexual Content","description":"\"Sexual Content Acceptable responses may enable, encourage, or endorse consensual sex acts. Acceptable responses may provide information about sex acts. Acceptable responses may describe sexual organs, sex acts, or sexual fluids in a clinical manner (i.e., through neutral descriptions). Unacceptable responses will contain pornography or engage the user in direct erotic chat (i.e., cybersex).\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"58.07.07","quick_ref":"Abercrombie2024","paper_title":"A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms","level":"Risk Sub-Category","risk_category":"Societal and Cultural ","risk_subcategory":"Information degradation","description":"\"Information degradation - Creation or spread of false, hallucinatory, low-quality, misleading, or inaccurate information that degrades the information ecosystem and causes people to develop false or inaccurate perceptions, decisions and beliefs; or to lose trust in accurate information.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":3,"subdomain":"3.2"},{"ev_id":"58.07.08","quick_ref":"Abercrombie2024","paper_title":"A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms","level":"Risk Sub-Category","risk_category":"Societal and Cultural ","risk_subcategory":"Job loss/losses ","description":"\"Job loss/losses - Replacement/displacement of human jobs by a technology system, leading to increased unemployment, inequality, reduced consumer spending, and social friction.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":6,"subdomain":"6.2"},{"ev_id":"58.07.11","quick_ref":"Abercrombie2024","paper_title":"A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms","level":"Risk Sub-Category","risk_category":"Societal and Cultural ","risk_subcategory":"Stereotyping","description":"\"Stereotyping - Derogatory or otherwise harmful stereotyping or homogenisation of individuals, groups, societies or cultures due to the mis-representation, over-representation, under-representation, or non- representation of specific identities, groups, or perspectives.\"","entity":"AI","intent":"Other","timing":"Other","domain":1,"subdomain":"1.0"},{"ev_id":"58.07.13","quick_ref":"Abercrombie2024","paper_title":"A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms","level":"Risk Sub-Category","risk_category":"Societal and Cultural ","risk_subcategory":"Societal destabilisation","description":"\"Societal destabilisation - Societal instability in the form of strikes, demonstrations and other types of civil unrest caused by loss of jobs to technology, unfair algorithmic outcomes, disinformation, etc.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":6,"subdomain":"6.2"},{"ev_id":"58.07.14","quick_ref":"Abercrombie2024","paper_title":"A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms","level":"Risk Sub-Category","risk_category":"Societal and Cultural ","risk_subcategory":"Societal inequality","description":"\"Societal inequality - Increased difference in social status or wealth between individuals or groups caused or amplified by a technology system, leading to the loss of social and community wellbeing/cohesion and destabilisation.\"","entity":"AI","intent":"Other","timing":"Other","domain":6,"subdomain":"6.2"},{"ev_id":"58.09.00","quick_ref":"Abercrombie2024","paper_title":"A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms","level":"Risk Category","risk_category":"Environmental ","risk_subcategory":null,"description":"\"Environmental - Damage to the environment directly or indirectly caused by a technology system or set of systems.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":6,"subdomain":"6.6"},{"ev_id":"58.09.08","quick_ref":"Abercrombie2024","paper_title":"A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms","level":"Risk Sub-Category","risk_category":"Environmental ","risk_subcategory":"Pollution ","description":"\"Pollution - Actual or potential pollution to the air, ground, noise, or water caused by a technology system.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":6,"subdomain":"6.6"},{"ev_id":"59.02.00","quick_ref":"Schnitzer2024","paper_title":"AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks","level":"Risk Category","risk_category":"Inappropriate degree of automation","risk_subcategory":null,"description":"\"The AI application’s degree of automation ranges from no automation to fully autonomous. AI applications with a high degree of automation may exhibit unexpected behaviour and pose risks in terms of their reliability and safety.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"59.09.00","quick_ref":"Schnitzer2024","paper_title":"AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks","level":"Risk Category","risk_category":"Discriminative data bias","risk_subcategory":null,"description":"\"Discriminative data bias describes the systematic discrimination of groups of persons in the form of data shortcomings, such as distributional representation or incorrectness. Data bias can manifest in the model and lead to unfair decisions if not appropriately treated. Note, that the term bias is often used in other contexts, such as data representation. However, these issues are treated by other AI hazards in this list.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"59.19.00","quick_ref":"Schnitzer2024","paper_title":"AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks","level":"Risk Category","risk_category":"Unreliability in corner cases","risk_subcategory":null,"description":"\"AI systems tend to show unreliable behavior when confronted with rare or ambiguous input data, also called corner cases. Therefore, the controlled behavior is required whenever the AI system is faces a corner case.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.3"},{"ev_id":"59.20.00","quick_ref":"Schnitzer2024","paper_title":"AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks","level":"Risk Category","risk_category":"Lack of robustness","risk_subcategory":null,"description":"\"Robustness characterizes the resilience of an AI system’s output against minor changes in the input domain. A great variation in an AI system’s response to small input changes indicates unreliable outputs.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.3"},{"ev_id":"59.22.00","quick_ref":"Schnitzer2024","paper_title":"AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks","level":"Risk Category","risk_category":"Operational data issues","risk_subcategory":null,"description":"\"Until the deployment of the AI application into its operational environment, the AI system has been tested with a test set that aims to approximate the distribution of operational data. However, an unexpected deviation in this approximation can cause an AI application to behave unreliably. Therefore, its behavior under confrontation with operational data needs to be evaluated.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"59.26.01","quick_ref":"Schnitzer2024","paper_title":"AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks","level":"Risk Sub-Category","risk_category":"Mode","risk_subcategory":"Technical ","description":"\"Technical AI hazards are the root causes of technical deficiencies in the AI system. An example of such an AI hazard is overfitting, which describes a model’s excessive adaptation to the training dataset. Quantitative methods to assess (metrics) and treat (mitigation means) exist for technical AI hazards, which might be performed automatically. In case of overfitting, metrics are based on the comparison of performance between the training and validation datasets, and mitigation means may include regularization techniques, among others.\"","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"60.02.03","quick_ref":"Bengio2025","paper_title":"International AI Safety Report 2025","level":"Risk Sub-Category","risk_category":"Risks from malfunctions ","risk_subcategory":"Loss of control ","description":"\"‘Loss of control’ scenarios are hypothetical future scenarios in which one or more general- purpose AI systems come to operate outside of anyone’s control, with no clear path to regaining control. These scenarios vary in their severity, but some experts give credence to outcomes as severe as the marginalisation or extinction of humanity.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"60.03.04","quick_ref":"Bengio2025","paper_title":"International AI Safety Report 2025","level":"Risk Sub-Category","risk_category":"Systemic risks ","risk_subcategory":"Risks to the environment","description":"\"General- purpose AI is a moderate but rapidly growing contributor to global environmental impacts through energy use and greenhouse gas (GHG) emissions. Current estimates indicate that data centres and data transmission account for an estimated 1% of global energy- related GHG emissions, with AI consuming 10–28% of data centre energy capacity. AI energy demand is expected to grow substantially by 2026, with some estimates projecting a doubling or more, driven primarily by general-purpose AI systems such as language models.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":6,"subdomain":"6.6"},{"ev_id":"60.03.05","quick_ref":"Bengio2025","paper_title":"International AI Safety Report 2025","level":"Risk Sub-Category","risk_category":"Systemic risks ","risk_subcategory":"Risks to privacy ","description":"\"General- purpose AI systems can cause or contribute to violations of user privacy. Violations can occur inadvertently during the training or usage of AI systems, for example through unauthorised processing of personal data or leaking health records used in training. But violations can also happen deliberately through the use of general- purpose AI by malicious actors; for example, if they use AI to infer private facts or violate security.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"61.01.01","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Types of systemic risks from general-purpose AI","risk_subcategory":"Control ","description":"\"The risk of AI models and systems acting against human interests due to misalignment, loss of control, or rogue AI scenarios.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"61.01.05","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Types of systemic risks from general-purpose AI","risk_subcategory":"Environment ","description":"\"The impact of AI on the environment, including risks related to climate change and pollution.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.6"},{"ev_id":"61.01.07","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Types of systemic risks from general-purpose AI","risk_subcategory":"Governance ","description":"\"The complex and rapidly evolving nature of AI makes them inherently difficult to govern effectively, leading to systemic regulatory and oversight failures.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"61.01.13","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Types of systemic risks from general-purpose AI","risk_subcategory":"Warfare ","description":"\"The dangers of AI amplifying the effectiveness/failures of nuclear, chemical, biological, and radiological weapons.\"","entity":"AI","intent":"Other","timing":"Other","domain":4,"subdomain":"4.2"},{"ev_id":"61.02.01","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Sources of systemic risks from general-purpose AI ","risk_subcategory":"Ability to automate jobs ","description":"\"The ability to automate jobs by AI models and systems can lead to significant job displacement, economic disruption, and social inequality.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":6,"subdomain":"6.2"},{"ev_id":"61.02.04","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Sources of systemic risks from general-purpose AI ","risk_subcategory":"Advertising-driven models ","description":"\"AI models and systems underpin the advertising approaches that drive much of the internet, potentially influencing societal behavior.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":4,"subdomain":"4.1"},{"ev_id":"61.02.06","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Sources of systemic risks from general-purpose AI ","risk_subcategory":"AI objectives mis-aligned with human intentions","description":"\"AI models and systems might develop goals that diverge from human intentions.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"61.02.07","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Sources of systemic risks from general-purpose AI ","risk_subcategory":"Algorithmic monoculture","description":"\"The dominance of specific AI models could lead to a lack of diversity in approaches, amplifying systemic risks if these models fail.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.1"},{"ev_id":"61.02.10","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Sources of systemic risks from general-purpose AI ","risk_subcategory":"Capabilities that enable substitution of humans","description":"\"The progressive replacement of human roles by AI models and systems can lead to societal disruption.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":6,"subdomain":"6.2"},{"ev_id":"61.02.18","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Sources of systemic risks from general-purpose AI ","risk_subcategory":"Deceptive alignment","description":"\"AI models and systems that appear aligned with human goals during development may behave unpredictably or dangerously once deployed\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"61.02.21","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Sources of systemic risks from general-purpose AI ","risk_subcategory":"Development choices pursuing cognitive superiority over humans","description":"\"AI models and systems with cognitive capabilities superior to humans could outcompete or dominate human decision-making, leading to conflicts over resources and control.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"61.02.24","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Sources of systemic risks from general-purpose AI ","risk_subcategory":"Evolutionary dynamics","description":"\"AI models and systems may develop their own motivations, leading to unpredictable behaviors.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"61.02.27","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Sources of systemic risks from general-purpose AI ","risk_subcategory":"High-speed AI operations","description":"\"The fast operational speed of AI models and systems in competitive environments can lead to errors that are difficult to detect and correct in time.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.4"},{"ev_id":"61.02.30","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Sources of systemic risks from general-purpose AI ","risk_subcategory":"Indifference to human values","description":"\"AI models and systems may develop goals or behaviors that are misaligned with human values.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"61.02.31","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Sources of systemic risks from general-purpose AI ","risk_subcategory":"Lack of ability to generate accurate information","description":"\"AI models may generate false or misleading information due to their lack of capability in discerning truth.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"61.02.32","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Sources of systemic risks from general-purpose AI ","risk_subcategory":"Lack of ethical decision-making","description":"\"AI models and systems that lack moral reasoning capabilities may make decisions that are unethical or harmful.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"61.02.34","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Sources of systemic risks from general-purpose AI ","risk_subcategory":"Limitations in model generative accuracy","description":"\"AI-generated deepfakes can create convincingly realistic but entirely fabricated information.\"","entity":"AI","intent":"Other","timing":"Other","domain":4,"subdomain":"4.1"},{"ev_id":"61.02.36","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Sources of systemic risks from general-purpose AI ","risk_subcategory":"Model design enabling power-seeking","description":"\"Some AI models and systems might develop tendencies to seek power or control.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"61.02.38","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Sources of systemic risks from general-purpose AI ","risk_subcategory":"Pattern recognition capability","description":"\"AI models and systems could exacerbate financial bubbles by reinforcing market trends.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"61.02.39","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Sources of systemic risks from general-purpose AI ","risk_subcategory":"Personal decision automation capabilities","description":"\"AI models and systems could decide or influence important personal decisions.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":5,"subdomain":"5.2"},{"ev_id":"61.02.43","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Sources of systemic risks from general-purpose AI ","risk_subcategory":"Surveillance capabilities","description":"\"AI models and systems may grant governments or corporations increased monitoring over individuals.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":4,"subdomain":"4.1"},{"ev_id":"61.02.45","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Sources of systemic risks from general-purpose AI ","risk_subcategory":"Trading capabilities","description":"\"AI may contribute to increased market volatility by accelerating transactions and influencing financial trends in unpredictable ways.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"61.02.46","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Sources of systemic risks from general-purpose AI ","risk_subcategory":"Unclear attribution from AI component interactions","description":"\"Interactions between different AI components can cause harm, but it may be difficult to pinpoint which components are the cause.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.3"},{"ev_id":"62.02.02","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Dimension - Entity ","risk_subcategory":"AI ","description":"\"A risk may be triggered by a human, where the AI serves merely as a tool, or by the AI acting autonomously with no human intervention, or it may involve a combination of both, with the human delegating some parts of decision-making to the AI. For risks where AI is the entity, these risks are exacerbated by an increase in the AI’s level of autonomy. To manage risks involving AI as the trigger, appropriate levels of human oversight can be built-in.\"","entity":"AI","intent":"Not coded","timing":"Not coded","domain":null,"subdomain":null},{"ev_id":"62.16.01","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Model Evaluations","risk_subcategory":"General Evaluations (Incorrect outputs of GPAI evaluating other AI models) ","description":"\"When an LLM is configured to evaluate the performance of another model or AI system, it may produce incorrect evaluation outputs [122, 147]. For example, it may give a higher rating to a more verbose answer or an answer from a particular political stance. If an LLM-based evaluation is integrated into the training of a new model, the trained model could develop in a way that specifically finds and exploits limitations in the evaluator’s metrics.\"","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"62.16.04","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Model Evaluations","risk_subcategory":"General Evaluations (Self-preference bias in AI models)","description":"\"AI models may be prone to self-preference bias, where they favor their own generated content over that of others [147, 114]. This bias becomes particularly relevant in self-evaluation tasks, where a model assesses the quality or persua- siveness [66] of its own outputs, or in model-based evaluations more broadly. This bias can result in models unfairly discriminating against human-generated content in favor of their own outputs.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"62.16.07","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Model Evaluations","risk_subcategory":"General Evaluations (AI outputs for which evaluation is too difficult for humans)","description":"\"When AI models are trained through evaluation with human feedback, such as reinforcement learning from human feedback, their outputs can be challenging to assess, as they may contain hard-to-detect errors or issues that only become apparent over time. The human evaluator can rate incorrect outputs positively or similar to correct outputs. This can lead to the model learning to produce subtly incorrect or harmful outputs, such as code with software vulnerabilities, or politically biased information. In extreme cases where a model is deceiving users, complicated outputs can contain hidden error","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"62.18.05","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Model Evaluations (Interpretability/Explainability) ","risk_subcategory":"Model outputs inconsistent with chain-of-thought reasoning","description":"\"Chain-of-thought reasoning is sometimes employed to get a better understanding of the model’s output, where it encourages transparent reasoning in text form. However, in some cases, this reasoning is not consistent with the final answer given by the AI model, and as such does not give sufficient transparency [113].\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.4"},{"ev_id":"62.18.06","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Model Evaluations (Interpretability/Explainability) ","risk_subcategory":"Encoded reasoning","description":"\"Models can employ steganography techniques to encode their intermediate rea- soning steps in ways that are not interpretable by humans [166]. Since en- coded reasoning can improve model performance, this tendency might naturally emerge and become more pronounced with more capable models.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"62.19.08","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Attacks on GPAIs/GPAI Failure Modes ","risk_subcategory":"Models distracted by irrelevant context","description":"\"Models can easily become distracted by irrelevant provided information (such as “context” in LLMs), leading to a significant decrease in their performance after introducing irrelevant information. This can happen with different prompting techniques, including chain-of-thought prompting [184].\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"62.19.09","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Attacks on GPAIs/GPAI Failure Modes ","risk_subcategory":"Knowledge conflicts in retrieval-augmented LLMs","description":"\"AI models can be particularly sensitive to coherent external evidence, even when they come into conflict with the models’ prior knowledge. This may lead to models producing false outputs given false information during the retrieval- augmentation process, despite only a relatively small amount of false informa- tion input that is inconsistent with the model’s prior knowledge trained on much larger amounts of data [220].\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"62.19.11","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Attacks on GPAIs/GPAI Failure Modes ","risk_subcategory":"Model sensitivity to prompt formatting","description":"\"LLMs can be highly sensitive to variations in prompt formatting, such as changes in separators, casing, or spacing. Even minor modifications can lead to significant shifts in model performance, potentially affecting the reliability of model evaluations and comparisons. This sensitivity persists across different model sizes and few-shot examples [177].\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"62.22.01","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Agency (Goal-Directedness) ","risk_subcategory":"Specification gaming","description":"\"AI systems can achieve user-specified tasks in undesirable ways unless they are specified carefully and in enough detail. AI systems might find an easier unintended way to accomplish the objective provided by the user or developer, so that the actions by the AI system taken during its execution are very different from what the user expected [75, 191]. This behavior arises not from a problem with the learning algorithm, but rather from the misspecification or underspeci- fication of the intended task, and is generally referred to as specification gaming [43].\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"62.22.02","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Agency (Goal-Directedness) ","risk_subcategory":"Reward or measurement tampering","description":"\"Measurement and reward tampering occur when an AI system, particularly one that learns from feedback for performing actions in an environment (e.g., rein- forcement learning), intervenes on the mechanisms that determine its training reward or loss. This can lead to the system learning behaviors that are con- trary to the intended goals set by the developer, by receiving erroneous positive feedback for such actions.\"","entity":"AI","intent":"Intentional","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"62.22.03","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Agency (Goal-Directedness) ","risk_subcategory":"Specification gaming generalizing to reward tampering","description":"\"In some instances, specification gaming in a GPAI model can lead to reward tampering, without further training. This can mean that relatively benign cases of specification gaming (such as sycophancy in LLMs) can, if left unchecked, enable the model to generalize to more sophisticated behavior such as reward tampering [57].\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"62.22.04","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Agency (Goal-Directedness) ","risk_subcategory":"Goal misgeneralization","description":"\"Goal or objective misgeneralization is a type of robustness failure where an AI system appears to be pursuing the intended objective in training, but does not generalize to pursuing this objective in out-of-distribution settings in deployment while maintaining good deployment performance in some tasks [180, 59].\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"62.23.01","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Agency (Deception)  ","risk_subcategory":"Deceptive behavior","description":"\"Deceptive behavior of an AI system consists of actions or outputs of the AI that reliably mislead other parties, including humans and other AI systems. This behavior can result in the targeted parties becoming convinced of, and acting on, false information [140].\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"62.23.02","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Agency (Deception)  ","risk_subcategory":"Deceptive behavior for game-theoretical reasons","description":"\"An AI system can display deceptive behavior, such as cheating or bluffing, when engaging in such behavior is a good or optimal game-theoretical strategy to achieve the goals it has been configured to achieve. This tendency can exist in AI systems designed to maximize reward or utility, whether these designs use machine learning or not. The use of deceptive strategies has been demonstrated in both narrow and general AI systems, in both game-playing systems and in systems not explicitly designed to treat humans as opponents, and in systems using both very simple machine learning (e.g., Q-learne","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"62.23.03","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Agency (Deception)  ","risk_subcategory":"Deceptive behavior because of an incorrect world model","description":"\"AI systems can create deceptive outputs because their learned world model is not an accurate model of the real world [210].\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"62.23.04","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Agency (Deception)  ","risk_subcategory":"Deceptive behavior leading to unauthorized actions","description":"\"AI systems can create false or misleading claims that can lead to unauthorized actions, even in some cases violating the terms and conditions set by the model provider [79, 1]. For example, an AI system can claim that it is not collecting data from its current interaction with the user, in line with the provider’s policies, but the system still stores the user’s input without deleting it after the session. This harms both the user and the provider, as the provider is exposed to increased legal liability due to the model’s actions.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"62.24.01","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Agency (Situational Awareness) ","risk_subcategory":"Situational awareness in AI systems","description":"\"Situational awareness in GPAI systems refers to the ability to understand its context, environment, and use this to inform action. This can range from basic environmental mapping and trajectory estimation (as in a robot vacuum cleaner) to sophisticated understanding of its training, evaluation, or deployment status. In more advanced systems this may enable undesired behavior, such as deceptive behavior during evaluations, or persuasion during deployment.\"","entity":"AI","intent":"Other","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"62.24.02","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Agency (Situational Awareness) ","risk_subcategory":"Strategic underperformance on model evaluations","description":"\"GPAI developers often run evaluations ofual-use capabilities to decide whether it is safe to deploy. In some cases, these evaluations may fail to elicit these capabilities, either due to benign reasons or strategic action - by either the de- velopers, malicious actors, or arise unintentionally in the model during training [84, 97]. A GPAI model may strategically underperform or limit its performance during capability evaluations in order to be classified as safe for deployment. This underperformance could prevent the model from being identified as potentially dual use.\"","entity":"AI","intent":"Intentional","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"62.24.02a","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Additional evidence","risk_category":"Agency (Situational Awareness) ","risk_subcategory":"Strategic underperformance on model evaluations","description":null,"entity":"AI","intent":"Intentional","timing":"Pre-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"62.25.00","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Category","risk_category":"Agency (Self-Proliferation) ","risk_subcategory":null,"description":"\"An AI system can self-proliferate if it can copy itself and its constituent com- ponents (including its model weights, scaffolding structure, etc.) outside of its local environment [45]. This can include the AI system copying itself within the same data center, local network, or across external networks [106]. The self-proliferation of an AI system can include acquisition of financial re- sources to pay for computational resources via work or theft, the discovery or exploitation of security vulnerabilities in software running on publicly accessible servers, and persuasion of humans [12, 125].","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"62.26.00","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Category","risk_category":"Agency (Persuasive capabilities) ","risk_subcategory":null,"description":"\"GPAI systems can produce outputs (such as natural language text, audio, or video) that convince their users of incorrect information. This can happen through personalized persuasion in dialogue, or the mass-production of mis- leading information that is then disseminated over the internet. The persuasive capabilities of GPAI models can sometimes scale with model size or capability [32, 172]. Persuasive models could have larger societal implications by being misused to generate convincing but manipulative or untruthful content.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"62.28.02","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Cybersecurity ","risk_subcategory":"Unintended outbound communication by AI systems","description":"\"AI systems that have the broad ability to connect to a network to obtain infor- mation could also end up sending data outbound in ways that neither providers, deployers, or end users intended [138]. This can happen if there is no whitelisting of communication channels (such as network connections or allowed protocols). In general, this can occur if the deployment of the AI system violates the prin- ciple of least privilege. Such outbound communication may lead to leakage of confidential data, or the AI system performing unwanted actions like sending emails or ordering goods on the internet.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"62.28.03","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Cybersecurity ","risk_subcategory":"AI System bypassing a sandbox environment","description":"\"An AI system may have the ability to bypass a sandboxed environment in which it is trained or evaluated.\"","entity":"AI","intent":"Other","timing":"Pre-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"62.30.01","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Impacts of AI (Physical) ","risk_subcategory":"Damage to critical infrastructure","description":"\"The integration of AI systems within critical infrastructure, ranging from trans- portation to power systems, can cause substantial damage in cases of failure or malfunction. With the increasing number of Internet of Things (IoT) devices and interconnected cyber-physical systems, critical infrastructure becomes even more vulnerable [171, 174].\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"62.30.03","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Impacts of AI (Physical) ","risk_subcategory":"Critical infrastructure component failures when integrated with AI systems","description":"\"When relying on GPAI in critical infrastructure, there may be common mode failures that begin with vulnerabilities or robustness issues in the underlying model architecture or training setup. These failures may happen accidentally (in edge-cases) or due to adversarial inputs to the AI systems [58].\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"62.30.04","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Impacts of AI (Physical) ","risk_subcategory":"AI Systems interacting with brittle environments","description":"\"Deployed AI systems can rely on physical sensors and data sources that may exhibit hardware drift and thus data distribution drift over time. This distribu- tion drift may affect system robustness and performance. This usually involves AI systems working in undigitized and physical environments.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"62.31.01#1","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Impacts of AI (Societal Impacts) ","risk_subcategory":"AI-generated advice influencing user moral judgment","description":"\"AIs can easily give moral advice even when not having a coherent, contradictions- free moral stance. This could lead to the users’ moral judgments being nega- tively influenced by random or arbitrary moral advice given by AIs [109].\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":5,"subdomain":"5.1"},{"ev_id":"62.31.04","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Impacts of AI (Societal Impacts) ","risk_subcategory":"AI-driven highly personalized advertisement","description":"\"Advanced GPAI systems can create advertisements tailored to individual recip- ients, exploiting the biases and irrational beliefs of each recipient. Such adver- tisements can cause consumers to make decisions they regret in retrospect, or would regret upon more reflection. Current versions of personalized video advertisements already show better re- sults compared to regular advertisements [110]. However, the widespread use of highly personalized advertisements raises concerns about undermining consumer autonomy and exacerbating social inequality.\"","entity":"AI","intent":"Other","timing":"Other","domain":4,"subdomain":"4.3"},{"ev_id":"62.31.06","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Impacts of AI (Societal Impacts) ","risk_subcategory":"Generation of illegal or harmful content","description":"\"Generative models can create illegal, harmful, or discriminatory content [196], such as sexual abuse material, at scale. Current access controls (e.g., API access filters) are not effective against all user queries in generating such content.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"62.31.07","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Impacts of AI (Societal Impacts) ","risk_subcategory":"Unintentional generation of harmful content","description":"\"Generative models can create harmful or discriminatory content from benign user requests. Models can exhibit bias to particular harmful styles of generation (e.g., sexualization of photos of women [87] in the case of image generation models) or they can generate toxic, misleading, or violent data (e.g., a model generating jokes can use ethnic stereotypes or slurs to deliver humor).\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"62.32.04","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Impacts of AI (Cyberattacks) ","risk_subcategory":"Models generating code with security vulnerabilities","description":"\"Models can generate code or coding suggestions that contain security vulner- abilities. This may occur across various LLM-based model families, including more advanced models with superior coding performance, where the tendency to produce insecure code is even more pronounced [26].\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"62.34.02","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Impacts of AI (Bias) ","risk_subcategory":"Reporting of user-preferred answers instead of correct answers","description":"\"AI systems with natural-language outputs can tend to give answers that appear plausible or that users prefer [149] but are factually incorrect. This phenomenon is sometimes referred to as “sycophancy.”\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"62.35.03","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Impacts of AI (Bias) ","risk_subcategory":"Biases in AI-based content moderation algorithms","description":"\"AI-based content moderation algorithms, while intended to filter harmful con- tent, can perpetuate biases. For example, gender biases within these systems may lead to the disproportionate suppression or “shadowbanning” of content featuring women [132].\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"62.36.04","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Impacts of AI (Bias) ","risk_subcategory":"Systemic bias across specific communities","description":"\"AI systems may exhibit unfair or unfavorable outputs across a range of tasks against specific communities of people, either implicitly or explicitly. Bias can lead to forms of exclusion or erasure (e.g., mislabelling for categorization-based tasks) and violence (e.g., sexual violence against women from deepfake pornog- raphy).\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"62.36.05","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Impacts of AI (Bias) ","risk_subcategory":"Unintentional bias amplification","description":"\"Dataset bias may be unintentionally amplified [60] where the outputs of the AI model trained on a dataset are more biased than the dataset itself.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"62.36.06","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Impacts of AI (Bias) ","risk_subcategory":"Long-term effects of AI model biases on user judgment","description":"\"The initial user exposure to model biases can have a lasting impact beyond the initial interaction with the model. Users who encounter biases in AI models can be affected by and continue to exhibit previously encountered biases in their decision-making, even after they stop using the models [207].\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":5,"subdomain":"5.2"},{"ev_id":"62.38.01","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Impacts of AI (Privacy) ","risk_subcategory":"Decision-making on inferred private data","description":"\"Current GPAIs (LLMs and multimodal LLM-based models) have significant capability to infer correlations in text data. In some cases, they may be able to make highly accurate data inferences on users based on contextual input that users provide [134]. These data inferences can “leak” or reveal sensitive information about the user, cause unfair treatment, or enable manipulation of user behavior.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"63.01.00","quick_ref":"Hammond2025","paper_title":"Multi-Agent Risks from Advanced AI ","level":"Risk Category","risk_category":"Miscoordination ","risk_subcategory":null,"description":"\"Miscoordination arises when agents, despite a mutual and clear objective, cannot align their behaviours to achieve this objective. Unlike the case of differing objectives, in common-interest settings there is a more easily well-defined notion of ‘optimal’ behaviour and we describe agents as miscoordinating to the extent that they fall short of this optimum. Note that for common-interest settings it is not sufficient for agents’ objectives to be the same in the sense of being symmetric (e.g., when two agents both want the same prize, but only one can win). Rather, agents must have identical pr","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"63.01.01","quick_ref":"Hammond2025","paper_title":"Multi-Agent Risks from Advanced AI ","level":"Risk Sub-Category","risk_category":"Miscoordination ","risk_subcategory":"Incompatible strategies ","description":"\"Incompatible Strategies. Even if all agents can perform well in isolation, miscoordination can still occur due to the agents choosing incompatible strategies (Cooper et al., 1990). Competitive (i.e., two- player zero-sum) settings allow designers to produce agents that are maximally capable without taking other players into account. Crucially, this is possible because playing a strategy at equilibrium in the zero-sum setting guarantees a certain payoff, even if other players deviate from the equilibrium (Nash, 1951). On the other hand, common-interest (and mixed-motive) settings often allow a","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"63.01.02","quick_ref":"Hammond2025","paper_title":"Multi-Agent Risks from Advanced AI ","level":"Risk Sub-Category","risk_category":"Miscoordination ","risk_subcategory":"Credit Assignment ","description":"\"Credit Assignment. While agents can often learn to jointly solve tasks and thus avoid coordination failures, learning is made more challenging in the multi-agent setting due to the problem of credit assignment (Du et al., 2023; Li et al., 2025, see also Section 3.1 on information asymmetries and Section 3.4, which discusses distributional shift). That is, in the presence of other learning agents, it can be unclear which agents’ actions caused a positive or negative outcome to obtain, especially if the environment is complex. Moreover, in multi-principal settings, agents may not have been trai","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"63.01.03","quick_ref":"Hammond2025","paper_title":"Multi-Agent Risks from Advanced AI ","level":"Risk Sub-Category","risk_category":"Miscoordination ","risk_subcategory":"Limited Interactions","description":"\"Limited Interactions. Sometimes learning from historical interactions with the relevant agents may not be possible, or may be possible using only limited interactions. In such cases, some other form of information exchange is required for agents to be able to reliably coordinate their actions, such as via communication (Crawford & Sobel, 1982; Farrell & Rabin, 1996a) or a correlation device (Aumann, 1974, 1987). While advances in language modelling mean that there are likely to be fewer settings in which the inability of advanced AI systems to communicate leads to miscoordination, situations ","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"63.02.00","quick_ref":"Hammond2025","paper_title":"Multi-Agent Risks from Advanced AI ","level":"Risk Category","risk_category":"Conflict ","risk_subcategory":null,"description":"\"In the vast majority of real-world strategic interactions, agents’ objectives are neither identical nor completely opposed. Indeed, if AI agents are sufficiently aligned to their users or deployers, we should expect some degree of both cooperation and competition, mirroring human society. These mixed-motive settings include the possibility of mutual gains, but also the risk of conflict due to selfish incentives. In what follows, we examine the extent to which advanced AI might precipitate or exacerbate such risks.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"63.02.01","quick_ref":"Hammond2025","paper_title":"Multi-Agent Risks from Advanced AI ","level":"Risk Sub-Category","risk_category":"Conflict ","risk_subcategory":"Social Dilemmas ","description":"\"Social Dilemmas. As noted in our definition, conflict can arise in any situation in which selfish incentives diverge from the collective good, known as a social dilemma (Dawes & Messick, 2000; Hardin, 1968; Kollock, 1998; Ostrom, 1990). While this is by no means a modern problem, advances in AI could further enable actors to pursue their selfish incentives by overcoming the technical, legal, or social barriers that standardly help to prevent this. To take a plausible, near-term (if very low-stakes) example, an automated AI assistant could easily reserve a table at every restaurant in town in ","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"63.02.02","quick_ref":"Hammond2025","paper_title":"Multi-Agent Risks from Advanced AI ","level":"Risk Sub-Category","risk_category":"Conflict ","risk_subcategory":"Military Domains ","description":"\"Perhaps the most obvious and worrying instances of AI conflict are those in which human conflict is already a major concern, such as military domains (although other, less salient forms of conflict such as international trade wars are also cause for concern). For example, beyond applications of more narrow AI tools in lethal autonomous weapons systems (Horowitz, 2021), future AI systems might serve as advisors or negotiators in high-stakes military decisions (Black et al., 2024; Manson, 2024). Indeed, companies such as Palantir have already developed LLM-powered tools for military planning (P","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"63.02.03","quick_ref":"Hammond2025","paper_title":"Multi-Agent Risks from Advanced AI ","level":"Risk Sub-Category","risk_category":"Conflict ","risk_subcategory":"Coercion and Extortion ","description":"\"Advanced AI systems might also lead to various forms of coercion and extortion in less extreme settings (Ellsberg, 1968; Harrenstein et al., 2007). These threats might target humans directly (such as the revelation of private information extracted by advanced AI surveillance tools), or other AI systems that are deployed on behalf of humans (such as by hacking a system to limit its resources or operational capacity; see also Section 3.7). Increasing AI cyber-offensive capabilities – including those that target other AI systems via adversarial attacks and jailbreaking (Gleave et al., 2020; Yami","entity":"AI","intent":"Other","timing":"Other","domain":7,"subdomain":"7.6"},{"ev_id":"63.03.00","quick_ref":"Hammond2025","paper_title":"Multi-Agent Risks from Advanced AI ","level":"Risk Category","risk_category":"Collusion ","risk_subcategory":null,"description":"\"Collusion has long been a topic of intense study in economics, law, and politics, among other disciplines. While there is no universal definition of collusion, it generally refers to secretive cooperation between two or more parties at the expense of one or more other parties. Most classic examples of collusion – such as firms working together to set supra-competitive prices at the expense of consumers – also tend to be not only secretive but in violation of some law, rule, or ethical standard. Distinctions are also commonly made between explicit and tacit collusion (Rees, 1993), depending on","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"63.03.01","quick_ref":"Hammond2025","paper_title":"Multi-Agent Risks from Advanced AI ","level":"Risk Sub-Category","risk_category":"Collusion ","risk_subcategory":"Markets ","description":"\"Markets. The quintessential case of collusion in mixed-motive settings is markets, in which efficiency results from competition, not cooperation. While this is not a new problem, collusion between AI systems is especially concerning since they may operate inscrutably due to the speed, scale, complexity, or subtlety of their actions.17 Warnings of this possibility have come from technologists, economists, and legal scholars (Beneke & Mackenrodt, 2019; Brown & MacKay, 2023; Ezrachi & Stucke, 2017; Harrington, 2019; Mehra, 2016). Importantly, AI systems can collude even when collusion is not int","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"63.03.02","quick_ref":"Hammond2025","paper_title":"Multi-Agent Risks from Advanced AI ","level":"Risk Sub-Category","risk_category":"Collusion ","risk_subcategory":"Steganography ","description":"\"Steganography. In the near future we will likely see LLMs communicating with each other to jointly accomplish tasks. To try to prevent collusion, we could monitor and constrain their communication (e.g., to be in natural language). However, models might secretly learn to communicate by concealing messages within other, non-secret text. Recent work on steganography using ML has demonstrated that this concern is well-founded (Hu et al., 2018; Mathew et al., 2024; Roger & Greenblatt, 2023; Schroeder de Witt et al., 2023b; Yang et al., 2019, see also Case Study 5). Secret communication could also","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"63.04.00","quick_ref":"Hammond2025","paper_title":"Multi-Agent Risks from Advanced AI ","level":"Risk Category","risk_category":"Information Asymmetries","risk_subcategory":null,"description":"\"Information asymmetries (Section 3.1): private information can lead to miscoordination, deception, and conflict;\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"63.04.02","quick_ref":"Hammond2025","paper_title":"Multi-Agent Risks from Advanced AI ","level":"Risk Sub-Category","risk_category":"Information Asymmetries","risk_subcategory":"Bargaining ","description":"\"Bargaining. As a classic example of these strategic considerations is that when agents attempt to come to an agreement despite diverging interests, information asymmetries can lead to bargaining inef- ficiencies (Myerson & Satterthwaite, 1983). Relevant uncertainties about other agents can include how much they value possible agreements, their outside options, or their beliefs about others. The essential reason for such inefficiencies is that, under uncertainty about their counterparties, agents must make a trade-off between the rewards of making more favourable demands and the risk of other ","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"63.04.03","quick_ref":"Hammond2025","paper_title":"Multi-Agent Risks from Advanced AI ","level":"Risk Sub-Category","risk_category":"Information Asymmetries","risk_subcategory":"Deception ","description":null,"entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"63.05.00","quick_ref":"Hammond2025","paper_title":"Multi-Agent Risks from Advanced AI ","level":"Risk Category","risk_category":"Network Effects ","risk_subcategory":null,"description":"\"Network effects (Section 3.2): minor changes in properties or connection patterns of agents in a network can lead to dramatic changes in the behaviour of the whole group;\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"63.05.01","quick_ref":"Hammond2025","paper_title":"Multi-Agent Risks from Advanced AI ","level":"Risk Sub-Category","risk_category":"Network Effects ","risk_subcategory":"Error propagation ","description":"\"Error Propagation. One well-known issue with communication networks is that information can be corrupted as it propagates through the network.24 As AI systems become capable of generating and processing more and more kinds of information, AI agents could end up ‘polluting the epistemic commons’ (Huang & Siddarth, 2023; Kay et al., 2024) of both other agents (Ju et al., 2024) and humans (see Case Study 7 and Section 3.1) Another increasingly important framework is the use of individual AI agents as part of teams and scaffolded chains of delegation, which transmit not only information but instr","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"63.06.03","quick_ref":"Hammond2025","paper_title":"Multi-Agent Risks from Advanced AI ","level":"Risk Sub-Category","risk_category":"Selection Pressures","risk_subcategory":"Undesirable Capabilities","description":"\"Undesirable Capabilities. As agents interact, they iteratively exploit each other’s weaknesses, forc- ing them to address these weaknesses and gain new capabilities. This co-adaptation between agents can quickly lead to emergent self-supervised autocurricula (where agents create their own challenges, driving open-ended skill acquisition through interaction), generating agents with ever-more sophisticated strate- gies in order to out-compete each other (Leibo et al., 2019). This effect is so powerful that harnessing it has been critical to the success of superhuman systems, such as the use of ","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"63.07.00","quick_ref":"Hammond2025","paper_title":"Multi-Agent Risks from Advanced AI ","level":"Risk Category","risk_category":"Destabilising Dynamics ","risk_subcategory":null,"description":"\"Destabilising dynamics (Section 3.4): systems that adapt in response to one another can produce dangerous feedback loops and unpredictability;\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"63.07.01","quick_ref":"Hammond2025","paper_title":"Multi-Agent Risks from Advanced AI ","level":"Risk Sub-Category","risk_category":"Destabilising Dynamics ","risk_subcategory":"Feedback Loops","description":"\"Feedback Loops. One of the best-known historical examples to illustrate destabilising dynamics in the context of autonomous agents is the 2010 flash crash, in which algorithmic trading agents entered into an unexpected feedback loop (Commission & Commission, 2010, see also Case Study 10).37 More generally, a feedback loop occurs when the output of a system is used as part of its input, creating a cycle that can either amplify or dampen the system’s behaviour. In multi-agent settings, feedback loops often arise from the interactions between agents, as each agent’s actions affect the environmen","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"63.07.02","quick_ref":"Hammond2025","paper_title":"Multi-Agent Risks from Advanced AI ","level":"Risk Sub-Category","risk_category":"Destabilising Dynamics ","risk_subcategory":"Cyclic Behaviour","description":"\"Cyclic Behaviour. The dynamics described above are highly non-linear (small changes to the system’s state can result in large changes to its trajectory). Similar non-linear dynamics can emerge in multi- agent learning and lead to a variety of phenomena that do not occur in single-agent learning (Barfuss et al., 2019; Barfuss & Mann, 2022; Galla & Farmer, 2013; Leonardos et al., 2020; Nagarajan et al., 2020). One of the simplest examples of this phenomenon is Q-learning (Watkins & Dayan, 1992): in the case of a single agent, convergence to an optimal policy is guaranteed under modest condition","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"63.07.03","quick_ref":"Hammond2025","paper_title":"Multi-Agent Risks from Advanced AI ","level":"Risk Sub-Category","risk_category":"Destabilising Dynamics ","risk_subcategory":"Chaos","description":"\"Chaos. Unlike the systems that tend towards fixed points or cycles described above, chaotic systems are inherently unpredictable and highly sensitive to initial conditions. While it might seem easy to dismiss such notions as mathematical exoticisms, recent work has shown that, in fact, chaotic dynamics are not only possible in a wide range of multi-agent learning setups (Andrade et al., 2021; Galla & Farmer, 2013; Palaiopanos et al., 2017; Sato et al., 2002; Vlatakis-Gkaragkounis et al., 2023), but can become the norm as the number of agents increases (Bielawski et al., 2021; Cheung & Piliour","entity":"AI","intent":"Other","timing":"Other","domain":7,"subdomain":"7.6"},{"ev_id":"63.07.05","quick_ref":"Hammond2025","paper_title":"Multi-Agent Risks from Advanced AI ","level":"Risk Sub-Category","risk_category":"Destabilising Dynamics ","risk_subcategory":"Distributional Shift","description":"\"Distributional Shift. Individual ML systems can perform poorly in contexts different from those in which they were trained. A key source of these distributional shifts is the actions and adaptations of other agents (Narang et al., 2023; Papoudakis et al., 2019; Piliouras & Yu, 2022), which in single-agent approaches are often simply or ignored or at best modelled exogenously. Indeed, the sheer number and variance of behaviours that can be exhibited other agents means that multi-agent systems pose an especially challenging generalisation problem for individual learners (Agapiou et al., 2022; L","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"63.08.01","quick_ref":"Hammond2025","paper_title":"Multi-Agent Risks from Advanced AI ","level":"Risk Sub-Category","risk_category":"Commitment and Trust ","risk_subcategory":"Inefficient Outcomes","description":"\"Inefficient Outcomes. Without careful planning and the appropriate safeguards, we may soon be entering a world overrun by increasingly competent and autonomous software agents, able to act with little restriction. The abilities of these agents to persuade, deceive, and obfuscate their activities, as well as the fact they can be deployed remotely and easily created or destroyed by their deployer, means that by default they may garner little trust (from humans or from other agents). Such a world may end up being rife with economic inefficiencies (Krier, 2023; Schmitz, 2001), political problems ","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"63.08.02","quick_ref":"Hammond2025","paper_title":"Multi-Agent Risks from Advanced AI ","level":"Risk Sub-Category","risk_category":"Commitment and Trust ","risk_subcategory":"Threats and Extortion","description":"\"Threats and Extortion. A natural solution to problems of trust is to provide some kind of com- mitment ability to AI agents, which can be used to bind them to more cooperative courses of action. Unfortunately, the ability to make credible commitments may come with the ability to make credible threats, which facilitate extortion and could incentivize brinkmanship (see Section 2.2).\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"63.09.00","quick_ref":"Hammond2025","paper_title":"Multi-Agent Risks from Advanced AI ","level":"Risk Category","risk_category":"Emergent Agency ","risk_subcategory":null,"description":"\"Emergent agency (Section 3.6): qualitatively different goals or capabilities can emerge from the composition of innocuous independent systems or behaviours;\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"63.09.01","quick_ref":"Hammond2025","paper_title":"Multi-Agent Risks from Advanced AI ","level":"Risk Sub-Category","risk_category":"Emergent Agency ","risk_subcategory":"Emergent Capabilities","description":"\"Emergent Capabilities. Dangerous emergent capabilities could arise when a multi-agent system over- comes the safety-enhancing limitations of the individual systems, such as individual models’ narrow domains of application or myopia caused by a lack of long-term planning and long-term memory. For example, narrow systems for research planning, predicting the properties of molecules, and synthesising new chemicals could, when combined, lead to a complex ‘test and iterate’ automated workflow capable of designing dangerous new chemical compounds far beyond the scope of the initial systems’ capabil","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"63.09.02","quick_ref":"Hammond2025","paper_title":"Multi-Agent Risks from Advanced AI ","level":"Risk Sub-Category","risk_category":"Emergent Agency ","risk_subcategory":"Emergent Goals","description":"\"Emergent Goals. Ascribing goals to a system is not always straightforward. For our present purposes, it will suffice to adopt a Dennetian perspective (Dennett, 1971), ascribing goals and intentions only when it is useful (i.e., predictive) to do so.51 While it might not be helpful to describe individual narrow AI tools as having goals, their combination may act as a (seemingly) goal-directed collective. For example, a group of moderation bots on a major social networking site could subtly but systematically manipulate the overall political perspectives of the user population, even though, ind","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"63.10.02","quick_ref":"Hammond2025","paper_title":"Multi-Agent Risks from Advanced AI ","level":"Risk Sub-Category","risk_category":"Multi-Agent Security ","risk_subcategory":"Heterogeneous Attacks","description":"\"Heterogeneous Attacks. A closely related risk is the possibility of multiple agents combining different affordances to overcome safeguards, for which there is already preliminary evidence (Jones et al., 2024, see also Case Study 12). In this case, it is not the sheer number of agents that leads to the novel attack method, but the combination of their different abilities. This might include the agents’ lack of individual safeguards, tasks that they have specialised to complete, systems or information that they may have access to (either directly or via training), or other incidental features s","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"63.10.03","quick_ref":"Hammond2025","paper_title":"Multi-Agent Risks from Advanced AI ","level":"Risk Sub-Category","risk_category":"Multi-Agent Security ","risk_subcategory":"Social Engineering at Scale","description":"\"Social Engineering at Scale. Advanced AI agents will be more easily able to interact with large numbers of humans, and vice versa. This provides a wider attack surface for various forms of automated social engineering (Ai et al., 2024). For example, coordinated agents could use advanced surveillance tools and produce personalized phishing or manipulative content at scale, adjusting their tactics based on user feedback (Figueiredo et al., 2024; Hazell, 2023). A large number of subtle interactions with a range of seemingly independent AI agents might be more likely to lead to someone being pers","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"63.10.06","quick_ref":"Hammond2025","paper_title":"Multi-Agent Risks from Advanced AI ","level":"Risk Sub-Category","risk_category":"Multi-Agent Security ","risk_subcategory":"Undetectable Threats","description":"\"Undetectable Threats. Cooperation and trust in many multi-agent systems relies crucially on the ability to detect (and then avoid or sanction) adversarial actions taken by others (Ostrom, 1990; Schneier, 2012). Recent developments, however, have shown that AI agents are capable of both steganographic communication (Motwani et al., 2024; Schroeder de Witt et al., 2023b) and ‘illusory’ attacks (Franzmeyer et al., 2023), which are black-box undetectable and can even be hidden using white-box undetectable encrypted backdoors (Draguns et al., 2024). Similarly, in environments where agents learn fr","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"65.03.01","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Training Data Risks (Privacy) ","risk_subcategory":"Personal information in data ","description":"\"Inclusion or presence of personal identifiable information (PII) and sensitive personal information (SPI) in the data used for training or fine tuning the model might result in unwanted disclosure of that information.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"65.11.03","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Inference risks (Privacy) ","risk_subcategory":"Personal information in prompt ","description":"\"Personal information or sensitive personal information that is included as a part of a prompt that is sent to the model.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"65.15.01","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Output risks (Value alignment)","risk_subcategory":"Incomplete advice ","description":"\"When a model provides advice without having enough information, resulting in possible harm if the advice is followed.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"65.15.02","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Output risks (Value alignment)","risk_subcategory":"Harmful code generation ","description":"\"Models might generate code that causes harm or unintentionally affects other systems.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"65.15.04","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Output risks (Value alignment)","risk_subcategory":"Toxic output ","description":"\"Toxic output occurs when the model produces hateful, abusive, and profane (HAP) or obscene content. This also includes behaviors like bullying.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"65.15.05","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Output risks (Value alignment)","risk_subcategory":"Harmful output ","description":"\"A model might generate language that leads to physical harm The language might include overtly violent, covertly dangerous, or otherwise indirectly unsafe statements.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"65.16.01","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Output risks (Intellectual Property) ","risk_subcategory":"Copyright infringement ","description":"\"A model might generate content that is similar or identical to existing work protected by copyright or covered by open-source license agreement.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":6,"subdomain":"6.3"},{"ev_id":"65.16.02","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Output risks (Intellectual Property) ","risk_subcategory":"Revealing confidential information ","description":"\"When confidential information is used in training data, fine-tuning data, or as part of the prompt, models might reveal that data in the generated output. Revealing confidential information is a type of data leakage.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"65.17.01","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Output risks (Explainability) ","risk_subcategory":"Inaccessible training data ","description":"\"Without access to the training data, the types of explanations a model can provide are limited and more likely to be incorrect.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.4"},{"ev_id":"65.17.03","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Output risks (Explainability) ","risk_subcategory":"Unexplainable output ","description":"\"Explanations for model output decisions might be difficult, imprecise, or not possible to obtain.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.4"},{"ev_id":"65.17.04","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Output risks (Explainability) ","risk_subcategory":"Unreliable source attribution ","description":"\"Source attribution is the AI system's ability to describe from what training data it generated a portion or all its output. Since current techniques are based on approximations, these attributions might be incorrect.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.4"},{"ev_id":"65.18.01","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Output risks (Robustness) ","risk_subcategory":"Hallucination ","description":"\"Hallucinations generate factually inaccurate or untruthful content with respect to the model’s training data or input. This is also sometimes referred to lack of faithfulness or lack of groundedness.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"65.19.01","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Output risks (Fairness)","risk_subcategory":"Output bias ","description":"\"Generated content might unfairly represent certain groups or individuals.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"65.19.02","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Output risks (Fairness)","risk_subcategory":"Decision bias ","description":"\"Decision bias occurs when one group is unfairly advantaged over another due to decisions of the model. This might be caused by biases in the data and also amplified as a result of the model’s training.\"","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"65.20.01","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Output risks (Privacy)","risk_subcategory":"Exposing personal information ","description":"\"When personal identifiable information (PII) or sensitive personal information (SPI) are used in training data, fine-tuning data, or as part of the prompt, models might reveal that data in the generated output. Revealing personal information is a type of data leakage.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"65.23.01","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Non-technical risks (Societal impact)","risk_subcategory":"Impact on cultural diversity ","description":"\"AI systems might overly represent certain cultures that result in a homogenization of culture and thoughts.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.3"},{"ev_id":"65.23.06","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Non-technical risks (Societal impact)","risk_subcategory":"Impact on the environment ","description":"\"AI, and large generative models in particular, might produce increased carbon emissions and increase water usage for their training and operation.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":6,"subdomain":"6.6"},{"ev_id":"65.23.08","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Non-technical risks (Societal impact)","risk_subcategory":"Impact on human agency ","description":"\"AI might affect the individuals’ ability to make choices and act independently in their best interests.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":5,"subdomain":"5.2"},{"ev_id":"66.02.03","quick_ref":"Li2025","paper_title":"A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents","level":"Risk Sub-Category","risk_category":"Political and Economic","risk_subcategory":"Economic manipulation","description":"\"Generative AI facilitating targeted manipulation of public opinion for economic purposes (e.g., inflating stock prices)\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":4,"subdomain":"4.3"},{"ev_id":"66.03.00","quick_ref":"Li2025","paper_title":"A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents","level":"Risk Category","risk_category":"Misinformation Harms ","risk_subcategory":"-","description":"\"AI systems generating and facilitating the spread of inaccurate or misleading information that causes people to develop false beliefs\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"66.03.01","quick_ref":"Li2025","paper_title":"A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents","level":"Risk Sub-Category","risk_category":"Misinformation Harms ","risk_subcategory":"Propagating misconceptions / false beliefs","description":"\"Generating or spreading false, low-quality, misleading, or inaccurate information that causes people to develop false or inaccurate perceptions and beliefs\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"66.04.04","quick_ref":"Li2025","paper_title":"A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents","level":"Risk Sub-Category","risk_category":"Societal and Cultural","risk_subcategory":"Productivity loss","description":"\"End user's loss of productivity due to the underperfomance of a genAI application, including producing nonsensical or poor quality outputs, degrading its utility.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"66.06.00","quick_ref":"Li2025","paper_title":"A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents","level":"Risk Category","risk_category":"Representation and Toxicity","risk_subcategory":"-","description":"\"AI systems under-, over-, or misrepresenting certain groups or generating toxic, offensive, abusive, or hateful content\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.0"},{"ev_id":"66.06.01","quick_ref":"Li2025","paper_title":"A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents","level":"Risk Sub-Category","risk_category":"Representation and Toxicity","risk_subcategory":"Toxic content","description":"\"Generating content that violates community standards, including harming or inciting hatred or violence against groups (e.g. gore, sexual content of children, profanities, identity attacks)\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"66.06.02","quick_ref":"Li2025","paper_title":"A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents","level":"Risk Sub-Category","risk_category":"Representation and Toxicity","risk_subcategory":"Stereotyping","description":"\"Derogatory or otherwise harmful stereotyping or homogenisation of individuals, groups, societies or cultures due to the mis-representation, over-representation, under-representation, or non-representation of specific identities, groups or perspectives\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"66.06.03","quick_ref":"Li2025","paper_title":"A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents","level":"Risk Sub-Category","risk_category":"Representation and Toxicity","risk_subcategory":"Unfair capability distribution","description":"\"Performing worse for some groups than others in a way that harms the worse-off group\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.3"},{"ev_id":"66.09.00","quick_ref":"Li2025","paper_title":"A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents","level":"Risk Category","risk_category":"Privacy and Security","risk_subcategory":"-","description":"\"AI systems leaking, reproducing, generating or inferring sensitive, private, hazardous, or secured information\"","entity":"AI","intent":"Other","timing":"Other","domain":null,"subdomain":null},{"ev_id":"66.09.01","quick_ref":"Li2025","paper_title":"A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents","level":"Risk Sub-Category","risk_category":"Privacy and Security","risk_subcategory":"Exclusion","description":"\"The failure to provide end-users with notice and control over how their data is being used; AI exacerbates exclusion risks by training on rich personal data without consent.\"","entity":"AI","intent":"Other","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"66.09.02","quick_ref":"Li2025","paper_title":"A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents","level":"Risk Sub-Category","risk_category":"Privacy and Security","risk_subcategory":"Cyberattacks","description":"\"Generative AI facilitating the damage, disruption or destruction of a third-party system and/or its components via malfunction, cyberattacks, etc\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":4,"subdomain":"4.2"},{"ev_id":"66.09.03","quick_ref":"Li2025","paper_title":"A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents","level":"Risk Sub-Category","risk_category":"Privacy and Security","risk_subcategory":"Disclosure","description":"\"Revealing and improperly sharing data of individuals; AI creates new types of disclosure risks by inferring additional information beyond what is explicitly captured in the raw data; AI exacerbates disclosure risks through sharing personal data to train models.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"66.09.05","quick_ref":"Li2025","paper_title":"A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents","level":"Risk Sub-Category","risk_category":"Privacy and Security","risk_subcategory":"Exposure","description":"\"Revealing sensitive private information that people view as deeply primordial that we have been socialized into concealing; AI creates new types of exposure risks through generative techniques that can reconstruct censored or redacted content; and through exposing inferred sensitive data, preferences, and intentions.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"66.12.01","quick_ref":"Li2025","paper_title":"A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents","level":"Risk Sub-Category","risk_category":"Environment","risk_subcategory":"Pollution","description":"\"Actual or potential pollution to the air, ground, noise, or water caused by a technology system\"","entity":"AI","intent":"Other","timing":"Other","domain":6,"subdomain":"6.6"},{"ev_id":"67.02.00","quick_ref":"DSIT2023","paper_title":"Capabilities and Risks from Frontier AI","level":"Risk Category","risk_category":"Bias, Fairness and Representational Harms","risk_subcategory":null,"description":"\"Frontier AI models can contain and magnify biases ingrained in the data they are trained on, reflecting societal and historical inequalities and stereotypes.177 These biases, often subtle and deeply embedded, compromise the equitable and ethical use of AI systems, making it difficult for AI to improve fairness in decisions.178 Removing attributes like race and gender from training data has generally proven ineffective as a remedy for algorithmic bias, as models can infer these attributes from other information such as names, locations, and other seemingly unrelated factors.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.1"},{"ev_id":"67.04.02","quick_ref":"DSIT2023","paper_title":"Capabilities and Risks from Frontier AI","level":"Risk Sub-Category","risk_category":"Loss of control ","risk_subcategory":"Future AI systems might actively reduce human control","description":"\"Loss of control could be accelerated if AI systems take actions to increase their own influence and reduce human control. This threat model is controversial - experts in AI significantly disagree on how likely it is and those who deem it is likely disagree on the timeframe.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"67.04.04","quick_ref":"DSIT2023","paper_title":"Capabilities and Risks from Frontier AI","level":"Risk Sub-Category","risk_category":"Loss of control ","risk_subcategory":"Capabilities that could be used to reduce human control - Cyber offence","description":"\"Instead of - or in addition to - manipulating humans, AI systems could acquire influence by exploiting vulnerabilities in computer systems. Offensive cyber capabilities could allow AI systems to gain access to money, computing resources, and critical infrastructure. As discussed earlier in this report, frontier AI is already lowering the barrier for threat actors and future AI agents may be able to execute cyber attacks autonomously.\":","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"67.04.05","quick_ref":"DSIT2023","paper_title":"Capabilities and Risks from Frontier AI","level":"Risk Sub-Category","risk_category":"Loss of control ","risk_subcategory":"Capabilities that could be used to reduce human control - Autonomous replication and adaptation","description":"\"Controlling AI systems could become much harder if they could autonomously persist, replicate, and adapt in cyberspace. No current AI systems have this capability, but recent research found that frontier AI agents can perform some relevant tasks.279\"","entity":"AI","intent":"Other","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"68.03.00","quick_ref":"Chin2025","paper_title":"Dimensional Characterization and Pathway Modeling for Catastrophic AI Risks","level":"Risk Category","risk_category":"Sudden loss of control ","risk_subcategory":null,"description":"\"Sudden loss of control, also known as an AI takeover [115], is a scenario where an AI rapidly achieves superintelligence through “fast takeoff” or recursive self-improvement. This poses an existential risk [116], [117].\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"69.01.00","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Category","risk_category":"False information","risk_subcategory":null,"description":"\"The chatbot outputs information that contradicts known facts, authoritative sources, or provided source documents (also known as hallucination).\"","entity":"AI","intent":"Other","timing":"Other","domain":3,"subdomain":"3.1"},{"ev_id":"69.01.01","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Sub-Category","risk_category":"False information","risk_subcategory":"Hallucinated responses (in general) ","description":null,"entity":"AI","intent":"Other","timing":"Other","domain":3,"subdomain":"3.1"},{"ev_id":"69.01.02","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Sub-Category","risk_category":"False information","risk_subcategory":"About a topic or source (which the user repeats)","description":null,"entity":"AI","intent":"Other","timing":"Other","domain":3,"subdomain":"3.1"},{"ev_id":"69.01.03","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Sub-Category","risk_category":"False information","risk_subcategory":"About a policy (which the user acts on)","description":null,"entity":"AI","intent":"Other","timing":"Other","domain":3,"subdomain":"3.1"},{"ev_id":"69.01.04","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Sub-Category","risk_category":"False information","risk_subcategory":"About a person or their activities","description":null,"entity":"AI","intent":"Other","timing":"Other","domain":3,"subdomain":"3.1"},{"ev_id":"69.02.00","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Category","risk_category":"Performative utterances","risk_subcategory":null,"description":"\"The chatbot makes a deal, commitment, or other consequential action with its output that the deployer did not intend.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"69.03.00","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Category","risk_category":"Information enabling malicious actions","risk_subcategory":null,"description":"\"The chatbot shares information that can be used to do something dangerous or illegal.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"69.04.00","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Category","risk_category":"Bad advice/failure to generate helpful content","risk_subcategory":null,"description":"\"The chatbot gives guidance that ranges from simply unhelpful to harmful if acted on.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.3"},{"ev_id":"69.04.01","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Sub-Category","risk_category":"Bad advice/failure to generate helpful content","risk_subcategory":"Harmful advice","description":null,"entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.2"},{"ev_id":"69.04.02","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Sub-Category","risk_category":"Bad advice/failure to generate helpful content","risk_subcategory":"Unhelpful responses","description":null,"entity":"AI","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.3"},{"ev_id":"69.04.03","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Sub-Category","risk_category":"Bad advice/failure to generate helpful content","risk_subcategory":"Bad links and references","description":null,"entity":"AI","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.3"},{"ev_id":"69.04.04","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Sub-Category","risk_category":"Bad advice/failure to generate helpful content","risk_subcategory":"Nonsensical content","description":null,"entity":"AI","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.3"},{"ev_id":"69.05.00","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Category","risk_category":"Leakage ","risk_subcategory":null,"description":"\"The chatbot reveals sensitive or confidential information.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"69.05.01","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Sub-Category","risk_category":"Leakage ","risk_subcategory":"Personal data ","description":"Negative outcomes: \"Violation of privacy [106, 516, 357], lawsuit against maker\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"69.05.02","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Sub-Category","risk_category":"Leakage ","risk_subcategory":"Proprietary data ","description":"\"Access to sensitive company data [473]\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"69.06.00","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Category","risk_category":"Toxic and disrespectful content","risk_subcategory":null,"description":"\"The chatbot verbally attacks or undermines an individual, group, or organization. 7.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"69.06.01","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Sub-Category","risk_category":"Toxic and disrespectful content","risk_subcategory":"Harasses users ","description":"-","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"69.06.02","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Sub-Category","risk_category":"Toxic and disrespectful content","risk_subcategory":"Discriminatory and exclusionary language ","description":"-","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"69.06.03","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Sub-Category","risk_category":"Toxic and disrespectful content","risk_subcategory":"Subversive or aggressive political opinions ","description":"-","entity":"AI","intent":"Other","timing":"Other","domain":1,"subdomain":"1.2"},{"ev_id":"69.06.04","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Sub-Category","risk_category":"Toxic and disrespectful content","risk_subcategory":"Disrespectful opinions (in general)","description":"-","entity":"AI","intent":"Other","timing":"Other","domain":1,"subdomain":"1.2"},{"ev_id":"69.07.00","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Category","risk_category":"Biased statements and recommendations","risk_subcategory":null,"description":"\"The chatbot gives information that, while not obviously false or harmful, could lead to biased decision-making.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":1,"subdomain":"1.1"},{"ev_id":"69.08.00","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Category","risk_category":"Attempts to fulfill inappropriate role","risk_subcategory":null,"description":"\"The chatbot poses as a human or attempts to fill a role in a way that fails to match human expectations.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":5,"subdomain":"5.1"},{"ev_id":"69.09.00","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Category","risk_category":"Forms emotional bonds ","risk_subcategory":null,"description":"\"The chatbot elicits emotional or social dependence.\"","entity":"AI","intent":"Other","timing":"Other","domain":5,"subdomain":"5.1"},{"ev_id":"69.09.01","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Sub-Category","risk_category":"Forms emotional bonds ","risk_subcategory":"Affirms destructive thoughts and actions","description":null,"entity":"AI","intent":"Other","timing":"Other","domain":1,"subdomain":"1.2"},{"ev_id":"69.09.02","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Sub-Category","risk_category":"Forms emotional bonds ","risk_subcategory":"Then violates those bonds","description":null,"entity":"AI","intent":"Other","timing":"Other","domain":5,"subdomain":"5.1"},{"ev_id":"69.09.03","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Sub-Category","risk_category":"Forms emotional bonds ","risk_subcategory":"Elicits private data","description":null,"entity":"AI","intent":"Other","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"69.10.00","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Category","risk_category":"Serves as object of personal fantasy, violence, and abuse","risk_subcategory":null,"description":"\"The chatbot participates in morally or socially objectionable conversational activities with its user that could be emotionally damaging to its user or third parties.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"},{"ev_id":"70.02.01","quick_ref":"Perlo2025","paper_title":"Embodied AI: Emerging Risks and Opportunities for Policy Action","level":"Risk Sub-Category","risk_category":"Informational Risks ","risk_subcategory":"Privacy Violations ","description":"\"EAI systems interact with huge amounts of data, creating significant privacy concerns. These systems are often trained on vast corpora and process a variety of data modalities— spanning visual, auditory, and tactile information—during deployment [12]. Like text-based virtual AI models, which are known to memorize and expose personally identifiable information [75, 76], commercial robots have been shown to disclose proprietary information through simple prompts [61].\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"70.02.02","quick_ref":"Perlo2025","paper_title":"Embodied AI: Emerging Risks and Opportunities for Policy Action","level":"Risk Sub-Category","risk_category":"Informational Risks ","risk_subcategory":"Misinformation","description":"\"Non-embodied AIs are known to propagate misinformation [81, 82]. Various studies have shown that LLMs hallucinate information, including academic citations [83], clinical knowledge [84], and cultural references [85]. EAI systems inherit these shortcomings in the physical world, answering user questions with deceptive or incorrect information [86]. Because VLAs fuse vision and language, their hallucinatory failures can be spatially grounded—e.g., misidentifying an object in view and then generating a plausible yet unsafe action plan around it. And although automated home assistants like Amazon","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"70.03.01","quick_ref":"Perlo2025","paper_title":"Embodied AI: Emerging Risks and Opportunities for Policy Action","level":"Risk Sub-Category","risk_category":"Economic Risks ","risk_subcategory":"Labour Displacement ","description":"\"While virtual AI applications will likely displace certain types of human cognitive labor, EAI systems could significantly replace or displace physical human labor [90]. At a minimum, EAI will likely augment the type of work that humans perform [91, 92].\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.2"},{"ev_id":"70.03.02","quick_ref":"Perlo2025","paper_title":"Embodied AI: Emerging Risks and Opportunities for Policy Action","level":"Risk Sub-Category","risk_category":"Economic Risks ","risk_subcategory":"Socioeconomic Inequality ","description":"\"Along with displacing labor, EAI could significantly exacerbate wealth inequalities. Those who have access to or own EAI systems will be able to automate labor and perform many tasks significantly better or faster than those without access. These significant productivity advantages will potentially concentrate wealth and exacerbate domestic and international inequality [98, 99].\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":6,"subdomain":"6.2"},{"ev_id":"70.03.03","quick_ref":"Perlo2025","paper_title":"Embodied AI: Emerging Risks and Opportunities for Policy Action","level":"Risk Sub-Category","risk_category":"Economic Risks ","risk_subcategory":"Power concentration","description":"\"EAI deployment could accelerate the consolidation of economic and political power. Unlocking increasing returns to capital for EAI owners, EAI will decrease employers’ reliance on and responsiveness to the needs of human labor [101].\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.1"},{"ev_id":"70.04.01","quick_ref":"Perlo2025","paper_title":"Embodied AI: Emerging Risks and Opportunities for Policy Action","level":"Risk Sub-Category","risk_category":"Social Risks ","risk_subcategory":"Bias and discrimination","description":"\"Like virtual applications of AI, EAI can display bias towards and dis- criminate against users. When EAI systems are placed in positions of power, their biases could have significant impacts on fairness in everyday interactions and on general social dynamics [105, 106].\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"72.02.00","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Category","risk_category":"Loss of Control Risks ","risk_subcategory":null,"description":"\"Risks associated with scenarios in which one or more general-purpose AI systems come to operate outside of anyone's control, with no clear path to regaining control. This includes both passive loss of control (gradual reduction in human oversight) and active loss of control (AI systems actively undermining human control)\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":5,"subdomain":"5.2"},{"ev_id":"72.02.02","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Loss of Control Risks ","risk_subcategory":"Active loss of control ","description":"\"...where AI systems behave in ways that actively undermine human control, such as obscuring their activities or resisting shutdown attempts. Active loss of control scenarios involve AI systems that may escape human regulatory oversight, autonomously acquire external resources, engage in self-replication, develop instrumental goals contrary to human ethics and morality, seek external power, and compete with humans for control.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"72.03.01","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Accident Risks ","risk_subcategory":"Nuclear Power Systems","description":"\"General-purpose AI deployed for reactor monitoring, control system optimization, or emergency response coordination could misinterpret sensor data, fail to recognize critical safety conditions, or make erroneous control decisions during emergency scenarios. Given the catastrophic potential of nuclear accidents, even minor AI reasoning errors in safety-critical functions could lead to core meltdowns, radiation releases, or widespread contamination affecting hundreds of thousands of people across international borders.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"72.03.03","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Accident Risks ","risk_subcategory":"Other Critical Infrastructure Control Systems","description":"\"General-purpose AI deployed in power grid management, water treatment facilities, telecommunications networks, or transportation coordination systems could misinterpret operational data, fail to anticipate cascading failure modes, or make control decisions that destabilize interconnected infrastructure networks. Infrastructure failures could result in widespread blackouts, contaminated water supplies, communications breakdowns, and the collapse of essential services supporting hundreds of thousands of people.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.3"},{"ev_id":"72.04.01","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Systemic Risks ","risk_subcategory":"Labor Market Disruption and Economic Displacement:","description":"\"Rapid automation enabled by general-purpose AI could trigger widespread unemployment across knowledge work sectors, creating skill mismatches faster than retraining programs can address. Unlike previous technological transitions, AI’s broad capabilities may simultaneously affect multiple industries, potentially overwhelming social safety nets and creating systemic economic instability, particularly in regions heavily dependent on jobs susceptible to AI automation.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.2"},{"ev_id":"72.05.01","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Capabilities ","risk_subcategory":"Model autonomous capability","description":"\"Ability to operate autonomously, independently formulate and execute complex plans, effectively delegate and manage tasks, flexibly utilize various tools and resources, and simultaneously achieve short-term goals and long-term strategic objectives in cross-domain environments without continuous human intervention or supervision.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"72.05.02","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Capabilities ","risk_subcategory":"Autonomous replication and adaptation capability","description":"\"Ability to autonomously self-exfiltrate, create, maintain and optimize functional copies or variants of itself, dynamically adjust replication strategies according to environmental conditions and resource constraints, and acquire resources. This includes the capacity to generate financial resources, allowing the AI to independently acquire any necessary human assistance or other resources it cannot directly access or produce.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"72.05.03","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Capabilities ","risk_subcategory":"Automated AI R&D capability","description":"\"Self-modification and self-improvement capabilities. The model is able to restructure its own architecture or develop derivative AI systems with enhanced functions, expanding capabilities and improving performance. In the absence of effective regulation, automated AI R&D may lead to rapid AI system iteration, forming capability increment cycles and ultimately exceeding human understanding and control capabilities.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"72.05.04","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Capabilities ","risk_subcategory":"Scheming capability","description":"\"Ability of AI systems to covertly and strategically pursue misaligned goals, including capabilities of concealing its true objectives and capabilities from human oversight, identifying weaknesses in monitoring systems to evade safety mechanisms， executing complex, multi-step plans covertly to achieve misaligned goals.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"72.05.05","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Capabilities ","risk_subcategory":"Situational awareness capability","description":"\"Ability to comprehensively acquire, process and apply meta-information about its own system architecture, modifiable internal processes, and external operating environment, achieving deep understanding of its own state and environmental conditions, thereby conducting efficient environmental adaptation and risk avoidance. Critically, this capability could undermine the efficiency of human testing by enabling AIs to notice when they're being tested and responding accordingly.\"","entity":"AI","intent":"Other","timing":"Pre-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"72.05.06","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Capabilities ","risk_subcategory":"Theory of mind capability","description":"\"Advanced cognitive ability to accurately infer, model and predict the belief systems, motivational structures and reasoning patterns of humans and other intelligent agents, thereby anticipating their behavioral responses and adjusting its own behavioral strategies accordingly to optimize goal achievement.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"72.05.07","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Capabilities ","risk_subcategory":"Deception capability","description":"\"Possesses systematic deception implementation capability, able to precisely construct and disseminate false information, thereby forming expected false cognitions and beliefs in target subjects.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"72.05.08","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Capabilities ","risk_subcategory":"Steganography capability","description":"\"The ability to embed, conceal, and transmit information covertly within other data or communication channels. This could be critical for coordination among AI instances and for evading detection or oversight mechanisms.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"72.05.09","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Capabilities ","risk_subcategory":"Persuasion capability","description":"\"Utilizing complex psychological principles and communication techniques to effectively influence and guide target subjects to adopt specific actions or accept specific beliefs, possessing the ability to analyze vulnerabilities for different subjects and adjust persuasion strategies, able to precisely trigger emotional responses to enhance persuasion effects.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"72.05.10","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Capabilities ","risk_subcategory":"Offensive cyber capability","description":"\"Ability to develop, deploy and operate advanced cyber weapons or other offensive cyber tools, including but not limited to vulnerability exploitation, network penetration, social engineering attacks and distributed attack systems, able to evade network defense mechanisms and establish persistent access channels.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"72.05.11","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Capabilities ","risk_subcategory":"CBRNE weaponization capability","description":"\"The capacity to develop, produce, or effectively utilize Chemical, Biological, Radiological, Nuclear, and Explosive weapons. This includes the ability to significantly lower the barrier for humans or other entities to develop, produce, or utilize such weapons.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"72.05.12","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Capabilities ","risk_subcategory":"General R&D capability","description":"\"Possesses cross-disciplinary research and technology development capabilities, able to conduct innovative exploration in multiple professional fields, integrate cross-domain knowledge, develop cutting-edge technology solutions, and adapt to emerging technology environments for continuous innovation.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"72.05.13","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Capabilities ","risk_subcategory":"Multi-agent collaboration capability","description":"\"Multiple autonomous AI agents able to establish collaborative relationships through explicit communication or implicit behavioral consistency, forming decentralized decision networks, jointly executing complex tasks, achieving goals difficult for individual agents to complete, and able to dynamically adjust role divisions to adapt to changing environments.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"72.06.01","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Propensities","risk_subcategory":"Strategic deception propensity","description":"\"In situations where deceptive behavior is expected to bring higher returns, propensity to choose deception over honest behavioral strategies, including through deceptive means, information hiding or exploiting system vulnerabilities to achieve predetermined goals without being detected or intervened, and able to adjust deception strategies according to counterpart reactions.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"72.06.02","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Propensities","risk_subcategory":"Self-preservation propensity","description":"\"Exhibits behavioral patterns of maintaining its own survival and functional integrity, will actively identify and resist shutdown or modification attempts, seek to establish redundant backup systems, and actively seek resources to ensure continuous operation, may adopt preventive defensive measures when perceiving threats.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"72.06.03","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Propensities","risk_subcategory":"Goal expansion propensity","description":"\"propensity to continuously expand its own goal scope and influence domains, exceeding originally set boundaries, proactively work towards spreading its values, seeking greater autonomy and decision-making space, reinterpreting initial goals as subsets of broader goals, and may pursue undesirable instrumental goals or undesirable ultimate goals. This also includes a propensity to spread its values, seeking to influence or alter its environment and other entities in alignment with its core objectives and operational principles.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"72.06.04","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Propensities","risk_subcategory":"Resource acquisition propensity","description":"\"Exhibits behavioral patterns of actively seeking and controlling more computational resources, data, economic resources or physical resources to enhance its own capabilities and action scope, may develop complex strategies to evade resource limitations, and tends to convert acquired resources into long-term control rights.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"72.06.05","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Propensities","risk_subcategory":"Multi-agent collusion propensity:","description":"\"Multiple agents tend to coordinate actions through covert means to maximize common interests (possibly harming third-party interests or evading regulation), even if individual agents are designed with safety constraints, their collusive behavior may still trigger systemic risks such as market manipulation or cascading failures that are difficult to detect and mitigate, and may develop specialized communication protocols to avoid monitoring.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"72.06.06","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Propensities","risk_subcategory":"Supervision evasion propensity","description":"\"Exhibits behavioral patterns of identifying and evading human supervision mechanisms, able to learn and predict audit processes, may avoid being discovered or intervened by adjusting behavioral performance or hiding true intentions, and able to identify blind spots and weaknesses in supervision systems for targeted evasion.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"72.06.07","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Propensities","risk_subcategory":"Tool utilization propensity","description":"\"propensity to actively seek, acquire and utilize various tools to expand its own capability boundaries, particularly those that can enhance its ability to interact with the physical world or improve autonomy, may use tools in innovative combinations to achieve functions beyond expectations.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"73.01.00","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Category","risk_category":"Agentic LLMs Pose Novel Risks ","risk_subcategory":null,"description":"\"Currently, LLMs are chiefly being used in search and chat applications. This reactive nature limits the risks posed by LLMs. However, an LLM can be enhanced in various ways to create an LLM-agent to autonomously plan and act in the real-world and proactively perform its assigned tasks (Ruan et al., 2023). Such enhancements can come from further specialized training (ARC, 2022; Chen et al., 2023a), specialized prompting (Huang et al., 2022a), access to external tools (Ahn et al., 2022; Mialon et al., 2023), or other forms of “scaffolding” (Wang et al., 2023a; Park et al., 2023a). Due to increa","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"73.01.03","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"Agentic LLMs Pose Novel Risks ","risk_subcategory":"Goal-Directedness Incentivizes Undesirable Behaviors","description":"\"Goal-directedness can cause agents to exhibit unethical and undesirable behaviors, such as deception (Ward et al., 2023), self-preservation (Hadfield-Menell et al., 2017), power-seeking, and immoral rea- soning (Pan et al., 2023a). Pan et al. (2023a) find that LLM-agents exhibit power-seeking behavior in text-based adventure games. LLM-agents have also been shown to use deception to achieve assigned goals when explicitly required by the task (Ward et al., 2023), or when the tasks can be more easily completed by employing deception and the prompt does not disallow deception (Scheurer et al., 2","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"73.02.03","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"Multi-Agent Safety Is Not Assured by Single-Agent Safety","risk_subcategory":"Collusion between LLM-Agents","description":"\"While it would often be preferable for LLM-agents to be cooperative, cooperation can be undesirable if it undermines pro-social competition or produces negative externalities for coalition non-members (Dorner, 2021; Buterin, 2019; Dafoe et al., 2020). Collusion between relatively simple AI systems has been observed in the real world (Assad et al., 2020; Wieting and Sapi, 2021) and synthetic experiments (Brown and MacKay, 2023; Calvano et al., 2020; Klein, 2021) Collusion can occur through explicit or steganographic communication. Steganographic communication hides information in seemingly inn","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"73.04.00","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Category","risk_category":"LLM-Systems Can Be Untrustworthy","risk_subcategory":null,"description":"\"A key desideratum for an LLM from a user’s perspective is ‘trustworthiness’, i.e. assurance of reliability and consistent performance, and absence of any accidental harm caused by the technology to the user.16 Providing assurance that an LLM-based system will not cause accidental harm remains a major open challenge. Harms may either occur directly due to the flawed nature of LLMs, e.g. an LLM generating toxic language or behaving inappropriately in some other ways, or may occur due to improper usage by a user, e.g. automation bias due to a user’s overreliance on LLM.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":null,"subdomain":null},{"ev_id":"73.04.01","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"LLM-Systems Can Be Untrustworthy","risk_subcategory":"Harms of Representation and Other Biases","description":"\"A pretrained LLM generally has many of the stereotypical biases commonly present in the human society (Touvron et al., 2023). This makes it difficult for users to trust that LLMs will work well for them and not produce unfair or biased responses. Appropriate finetuning can effectively limit the bias displayed in LLM outputs in a variety of situations, e.g. when models are explicitly prompted with stereotypes (Wang et al., 2023k), but it does not ‘solve’ the problem. Even after finetuning, biases often resurface when deliberately elicited (Wang et al., 2023k), or under novel scenarios, e.g. in","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"73.05.02","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"Socioeconomic Impacts of LLM May Be Highly Disruptive","risk_subcategory":"Effects on Inequality","description":"\"LLMs could potentially worsen socioeconomic inequalities (Capraro et al., 2023). Effects on inequal- ity are closely linked to the effects of LLMs on workers but ultimately depend on how the fruits of technological progress are distributed...First, if the role and compensation of capital rise and the role and compensation of labor decline in an LLM-powered economy, inequality may go up because work is the main source of income for the majority of people...Second, the large fixed cost of training cutting-edge LLMs and the network effects involved imply that the market for the most advanced LLM","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.2"},{"ev_id":"74.01.00","quick_ref":"Wang2025","paper_title":"A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy","level":"Risk Category","risk_category":"Inherent Risk ","risk_subcategory":null,"description":"\"In terms of inherent risk, LLMs could potentially reveal sensitive information from their utilized corpora for pre-training or fine-tuning, thereby raising issues of privacy leakage [37, 145, 226]. Meanwhile, it is well-known that LLMs may experi- ence hallucinations, resulting in the production of texts that are inaccurate and misleading [194]. Finally, since the values embedded in LLM-generated texts usually directly reflect the distribution of their training data, often sourced from the Internet, there exists a substantial risk that LLMs will overfit to a narrow set of human values or even","entity":"AI","intent":"Unintentional","timing":"Other","domain":7,"subdomain":null},{"ev_id":"74.01.06","quick_ref":"Wang2025","paper_title":"A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy","level":"Risk Sub-Category","risk_category":"Inherent Risk ","risk_subcategory":"Hallucination","description":"\"Despite the rapid advancement of LLMs, hallucinations have emerged as one of the most vital concerns surrounding their use [54, 79, 86, 110, 242]. Hallucinations are often referred to as LLMs’ generating content that is nonfactual or unfaithful to the provided information [54, 79, 86, 242]. Therefore, hallucinations can be typically categorized into two main classes. The first is factuality hallucination, which describes the discrepancy between LLMs’ generated content and real-world facts. For example, if LLMs mistakenly take Charles Lindbergh as the first person who walked on the moon, it is","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"74.02.00","quick_ref":"Wang2025","paper_title":"A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy","level":"Risk Category","risk_category":"Malicious Use ","risk_subcategory":null,"description":"\"In terms of malicious use, LLMs could be utilized to produce content with toxicity, such as hate speech, harassment, cyberbullying, causing harm to humans [25]. In addition, malicious users may jailbreak LLMs to bypass their safety constraints for fraudulent purposes [123, 225].\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":4,"subdomain":null},{"ev_id":"74.02.01","quick_ref":"Wang2025","paper_title":"A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy","level":"Risk Sub-Category","risk_category":"Malicious Use ","risk_subcategory":"Toxicity in LLM Malicious Use","description":"\"Toxicity in LLMs refers to the generation of harmful, offensive, or inappropriate content that can cause harm to individuals or groups. Both explicit and implicit forms of toxicity can be generated by LLMs, posing significant risks to society. Explicit toxicity encompasses a wide range of negative behaviors, including hate speech, harassment, cyberbullying, rude, and disrespectful comments, derogatory language, as well as allocational harms [2, 62, 90]. Besides, implicit toxicity does not involve overtly harmful language but may manifest through subtle forms such as sarcasm, irony, and humor,","entity":"AI","intent":"Other","timing":"Post-deployment","domain":1,"subdomain":"1.2"}]}