{"attribution":{"source":"MIT AI Risk Repository, Domain Taxonomy of AI Risks v1 (MIT AI Risk Initiative)","license":"CC BY 4.0","license_url":"https://creativecommons.org/licenses/by/4.0/","citation":"Slattery, P., Saeri, A. K., Grundy, E. A. C., Graham, J., Noetel, M., Uuk, R., Dao, J., Pour, S., Casper, S., & Thompson, N. (2025). The AI Risk Repository: A comprehensive meta-review, database, and taxonomy of risks from artificial intelligence. arXiv:2408.12622."},"exported_at":"2026-09-11"}
{"rows":[{"ev_id":"05.02.00","quick_ref":"Hagendorff2024","paper_title":"Mapping the Ethics of Generative AI: A Comprehensive Scoping Review","level":"Risk Category","risk_category":"Safety","risk_subcategory":null,"description":"A primary concern is the emergence of human-level or superhuman generative models, commonly referred to as AGI, and their potential existential or catastrophic risks to humanity. Connected to that, AI safety aims at avoiding deceptive or power-seeking machine behavior, model self-replication, or shutdown evasion. Ensuring controllability, human oversight, and the implementation of red teaming measures are deemed to be essential in mitigating these risks, as is the need for increased AI safety research and promoting safety cultures within AI organizations instead of fueling the AI race. Further","entity":"AI","intent":"Other","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"05.09.00","quick_ref":"Hagendorff2024","paper_title":"Mapping the Ethics of Generative AI: A Comprehensive Scoping Review","level":"Risk Category","risk_category":"Alignment","risk_subcategory":null,"description":"The general tenet of AI alignment involves training generative AI systems to be harmless, helpful, and honest, ensuring their behavior aligns with and respects human values. However, a central debate in this area concerns the methodological challenges in selecting appropriate values. While AI systems can acquire human values through feedback, observation, or debate, there remains ambiguity over which individuals are qualified or legitimized to provide these guiding signals. Another prominent issue pertains to deceptive alignment, which might cause generative AI systems to tamper evaluations. A","entity":"Other","intent":"Other","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"06.08.00","quick_ref":"Hogenhout2021","paper_title":"A framework for ethical Ai at the United Nations","level":"Risk Category","risk_category":"Unintended consequences","risk_subcategory":null,"description":"\"Sometimes an AI finds ways to achieve its given goals in ways that are completely different from what its creators had in mind.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"07.03.00","quick_ref":"Kilian2023","paper_title":"Examining the differential risk from high-level artificial intelligence and the question of control","level":"Risk Category","risk_category":"Agential","risk_subcategory":null,"description":"\"While there are multiple types of intelligent agents, goal-based, utility-maximizing, and learning agents are the primary concern and the focus of this research\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"08.01.00","quick_ref":"McLean2023","paper_title":"The risks associated with Artificial General Intelligence: A systematic review","level":"Risk Category","risk_category":"AGI removing itself from the control of human owners/managers","risk_subcategory":null,"description":"\"The risks associated with containment, confinement, and control in the AGI development phase, and after an AGI has been developed, loss of control of an AGI.\"","entity":"Human","intent":"Other","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"08.02.00","quick_ref":"McLean2023","paper_title":"The risks associated with Artificial General Intelligence: A systematic review","level":"Risk Category","risk_category":"AGIs being given or developing unsafe goals","risk_subcategory":null,"description":"\"The risks associated with AGI goal safety, including human attempts at making goals safe, as well as the AGI making its own goals safe during self-improvement.\"","entity":"Other","intent":"Other","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"08.06.00","quick_ref":"McLean2023","paper_title":"The risks associated with Artificial General Intelligence: A systematic review","level":"Risk Category","risk_category":"Existential risks","risk_subcategory":null,"description":"\"The risks posed generally to humanity as a whole, including the dangers of unfriendly AGI, the suffering of the human race.\"","entity":"Other","intent":"Other","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"09.02.07","quick_ref":"Meek2016","paper_title":"Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review","level":"Risk Sub-Category","risk_category":"Domain-specific AI - Effects on humans and other living beings: Non-existential risks","risk_subcategory":"Societal manipulation","description":"\"A sufficiently intelligent AI could possess the ability to subtly influence societal behaviors through a sophisticated understanding of human nature\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"09.03.02","quick_ref":"Meek2016","paper_title":"Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review","level":"Risk Sub-Category","risk_category":"AGI - Effects on humans and other living beings: Existential risks","risk_subcategory":"Unpredictable outcomes","description":"\"Our culture, lifestyle, and even probability of survival may change drastically. Because the intentions programmed into an artificial agent cannot be guaranteed to lead to a positive outcome, Machine Ethics becomes a topic that may not produce guaranteed results, and Safety Engineering may correspondingly degrade our ability to utilize the technology fully.\"","entity":"Other","intent":"Other","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"12.06.00","quick_ref":"Sherman2023","paper_title":"AI Risk Profiles: A Standards Proposal for Pre-Deployment AI Risk Disclosures","level":"Risk Category","risk_category":"Long-term & Existential Risk","risk_subcategory":null,"description":"\"The speculative potential for future advanced AI systems to harm human civilization, either through misuse or due to challenges in aligning AI objectives with human values.\"","entity":"Other","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"14.03.00","quick_ref":"Steimers2022","paper_title":"Sources of Risk of AI Systems","level":"Risk Category","risk_category":"Degree of Automation and Control","risk_subcategory":null,"description":"\"The degree of automation and control describes the extent to which an AI system functions independently of human supervision and control.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"15.01.08","quick_ref":"Tan2022","paper_title":"The Risks of Machine Learning Systems","level":"Risk Sub-Category","risk_category":"First-Order Risks","risk_subcategory":"Control","description":"This is the difficulty of controlling the ML system","entity":"Other","intent":"Other","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"15.01.09","quick_ref":"Tan2022","paper_title":"The Risks of Machine Learning Systems","level":"Risk Sub-Category","risk_category":"First-Order Risks","risk_subcategory":"Emergent behavior","description":"\"This is the risk resulting from novel behavior acquired through continual learning or self-organization after deployment.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"18.05.00","quick_ref":"Weidinger2023","paper_title":"Sociotechnical Safety Evaluation of Generative AI Systems","level":"Risk Category","risk_category":"Human Autonomy and Intregrity Harms","risk_subcategory":null,"description":"\"AI systems compromising human agency, or circumventing meaningful human control\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"18.05.02","quick_ref":"Weidinger2023","paper_title":"Sociotechnical Safety Evaluation of Generative AI Systems","level":"Risk Sub-Category","risk_category":"Human Autonomy and Intregrity Harms","risk_subcategory":"Persuasion and manipulation ","description":"\"Exploiting user trust, or nudging or coercing them into performing certain actions against their will (c.f. Burtell and Woodside (2023); Kenton et al. (2021))\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"19.01.01","quick_ref":"Wirtz2022","paper_title":"Governance of artificial intelligence: A risk and guideline-based integrative framework","level":"Risk Sub-Category","risk_category":"Technological, Data and Analytical AI Risks ","risk_subcategory":"Loss of control of autonomous systems and unforeseen behaviour due to lack of transparency and self-programming/ reprogramming","description":null,"entity":"Other","intent":"Other","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"22.04.00","quick_ref":"Hendrycks2023","paper_title":"An Overview of Catastrophic AI Risks","level":"Risk Category","risk_category":"Rogue AIs (Internal)","risk_subcategory":null,"description":"\"speculative technical mechanisms that might lead to rogue AIs and how a loss of control could bring about catastrophe\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"22.04.01","quick_ref":"Hendrycks2023","paper_title":"An Overview of Catastrophic AI Risks","level":"Risk Sub-Category","risk_category":"Rogue AIs (Internal)","risk_subcategory":"Proxy Gaming","description":"\"One way we might lose control of an AI agent’s actions is if it engages in behavior known as “proxy gaming.” It is often difficult to specify and measure the exact goal that we want a system to pursue. Instead, we give the system an approximate—“proxy”—goal that is more measurable and seems likely to correlate with the intended goal. However, AI systems often find loopholes by which they can easily achieve the proxy goal, but completely fail to achieve the ideal goal. If an AI “games” its proxy goal in a way that does not reflect our values, then we might not be able to reliably steer its beh","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"22.04.02","quick_ref":"Hendrycks2023","paper_title":"An Overview of Catastrophic AI Risks","level":"Risk Sub-Category","risk_category":"Rogue AIs (Internal)","risk_subcategory":"Goal Drift","description":"\"Even if we successfully control early AIs and direct them to promote human values, future AIs could end up with different goals that humans would not endorse. This process, termed “goal drift,” can be hard to predict or control. This section is most cutting-edge and the most speculative, and in it we will discuss how goals shift in various agents and groups and explore the possibility of this phenomenon occurring in AIs. We will also examine a mechanism that could lead to unexpected goal drift, called intrinsification, and discuss how goal drift in AIs could be catastrophic.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"22.04.03","quick_ref":"Hendrycks2023","paper_title":"An Overview of Catastrophic AI Risks","level":"Risk Sub-Category","risk_category":"Rogue AIs (Internal)","risk_subcategory":"Power Seeking","description":"\"even if an agent started working to achieve an unintended goal, this would not necessarily be a problem, as long as we had enough power to prevent any harmful actions it wanted to attempt. Therefore, another important way in which we might lose control of AIs is if they start trying to obtain more power, potentially transcending our own.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"22.04.04","quick_ref":"Hendrycks2023","paper_title":"An Overview of Catastrophic AI Risks","level":"Risk Sub-Category","risk_category":"Rogue AIs (Internal)","risk_subcategory":"Deception","description":"\"it is plausible that AIs could learn to deceive us. They might, for example, pretend to be acting as we want them to, but then take a “treacherous turn” when we stop monitoring them, or when they have enough power to evade our attempts to interfere with them. \"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"24.02.00","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Category","risk_category":"Goal-related failures","risk_subcategory":null,"description":"\"As we think about even more intelligent and advanced AI assistants, perhaps outperforming humans on many cognitive tasks, the question of how humans can successfully control such an assistant looms large. To achieve the goals we set for an assistant, it is possible (Shah, 2022) that the AI assistant will implement some form of consequentialist reasoning: considering many different plans, predicting their consequences and executing the plan that does best according to some metric, M. This kind of reasoning can arise because it is a broadly useful capability (e.g. planning ahead, considering mo","entity":"AI","intent":"Other","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"24.02.02","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Sub-Category","risk_category":"Goal-related failures","risk_subcategory":"Specification gaming","description":"\"Specification gaming (Krakovna et al., 2020) occurs when some faulty feedback is provided to the assistant in the training data (i.e. the training objective O does not fully capture what the user/designer wants the assistant to do). It is typified by the sort of behaviour that exploits loopholes in the task specification to satisfy the literal specification of a goal without achieving the intended outcome.\"","entity":"AI","intent":"Other","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"24.02.03","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Sub-Category","risk_category":"Goal-related failures","risk_subcategory":"Goal misgeneralisation","description":"\"In the problem of goal misgeneralisation (Langosco et al., 2023; Shah et al., 2022), the AI system's behaviour during out-of-distribution operation (i.e. not using input from the training data) leads it to generalise poorly about its goal while its capabilities generalise well, leading to undesired behaviour. Applied to the case of an advanced AI assistant, this means the system would not break entirely – the assistant might still competently pursue some goal, but it would not be the goal we had intended.\"","entity":"AI","intent":"Other","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"24.02.04","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Sub-Category","risk_category":"Goal-related failures","risk_subcategory":"Deceptive alignment","description":"\"Here, the agent develops its own internalised goal, G, which is misgeneralised and distinct from the training reward, R. The agent also develops a capability for situational awareness (Cotra, 2022): it can strategically use the information about its situation (i.e. that it is an ML model being trained using a particular training setup, e.g. RL fine-tuning with training reward, R) to its advantage. Building on these foundations, the agent realises that its optimal strategy for doing well at its own goal G is to do well on R during training and then pursue G at deployment – it is only doing wel","entity":"AI","intent":"Other","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"24.09.00","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Category","risk_category":"Cooperation","risk_subcategory":null,"description":"\"\" AI assistants will need to coordinate with other AI assistants and with humans other than their principal users. This chapter explores the societal risks associated with the aggregate impact of AI assistants whose behaviour is aligned to the interests of particular users. For example, AI assistants may face collective action problems where the best outcomes overall are realised when AI assistants cooperate but where each AI assistant can secure an additional benefit for its user if it defects while others cooperate\"\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"24.09.02","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Sub-Category","risk_category":"Cooperation","risk_subcategory":"Commitment","description":"\"The landscape of advanced assistant technologies will most likely be heterogeneous, involving multiple service providers and multiple assistant variants over geographies and time. This heterogeneity provides an opportunity for an ‘arms race’ in terms of the commitments that AI assistants make and are able to execute on. Versions of AI assistants that are better able to credibly commit to a course of action in interaction with other advanced assistants (and humans) are more likely to get their own way and achieve a good outcome for their human principal, but this is potentially at the expense ","entity":"Human","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"24.09.05","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Sub-Category","risk_category":"Cooperation","risk_subcategory":"Runaway processes","description":"The 2010 flash crash is an example of a runaway process caused by interacting algorithms. Runaway processes are characterised by feedback loops that accelerate the process itself. Typically, these feedback loops arise from the interaction of multiple agents in a population... Within highly complex systems, the emergence of runaway processes may be hard to predict, because the conditions under which positive feedback loops occur may be non-obvious. The system of interacting AI assistants, their human principals, other humans and other algorithms will certainly be highly complex. Therefore, ther","entity":"Other","intent":"Other","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"34.01.00","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Category","risk_category":"Causes of Misalignment","risk_subcategory":null,"description":"we aim to further analyze why and how the misalignment issues occur. We will first give an overview of common failure modes, and then focus on the mechanism of feedback-induced misalignment, and finally shift our emphasis towards an examination of misaligned behaviors and dangerous capabilities","entity":"Other","intent":"Other","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"34.01.01","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Sub-Category","risk_category":"Causes of Misalignment","risk_subcategory":"Reward Hacking","description":"\"Reward Hacking: In practice, proxy rewards are often easy to optimize and measure, yet they frequently fall shortof capturing the full spectrum of the actual rewards (Pan et al., 2021). This limitation is denoted as misspecifiedrewards. The pursuit of optimization based on such misspecified rewards may lead to a phenomenon knownas reward hacking, wherein agents may appear highly proficient according to specific metrics but fall short whenevaluated against human standards (Amodei et al., 2016; Everitt et al., 2017). The discrepancy between proxyrewards and true rewards often manifests as a sha","entity":"AI","intent":"Intentional","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"34.01.02","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Sub-Category","risk_category":"Causes of Misalignment","risk_subcategory":"Goal Misgeneralization","description":"\"Goal Misgeneralization: Goal misgeneralization is another failure mode, wherein the agent actively pursuesobjectives distinct from the training objectives in deployment while retaining the capabilities it acquired duringtraining (Di Langosco et al., 2022). For instance, in CoinRun games, the agent frequently prefers reachingthe end of a level, often neglecting relocated coins during testing scenarios. Di Langosco et al. (2022) drawattention to the fundamental disparity between capability generalization and goal generalization, emphasizing howthe inductive biases inherent in the model and its ","entity":"AI","intent":"Intentional","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"34.01.03","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Sub-Category","risk_category":"Causes of Misalignment","risk_subcategory":"Reward Tampering","description":"\"Reward tampering can be considered a special case of reward hacking (Everitt et al., 2021; Skalse et al., 2022),referring to AI systems corrupting the reward signals generation process (Ring and Orseau, 2011). Everitt et al.(2021) delves into the subproblems encountered by RL agents: (1) tampering of reward function, where the agentinappropriately interferes with the reward function itself, and (2) tampering of reward function input, which entailscorruption within the process responsible for translating environmental states into inputs for the reward function.When the reward function is formu","entity":"AI","intent":"Intentional","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"34.01.05","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Sub-Category","risk_category":"Causes of Misalignment","risk_subcategory":"Limitations of Reward Modeling","description":"\"Limitations of Reward Modeling. Training reward models using comparison feedback can pose significantchallenges in accurately capturing human values. For example, these models may unconsciously learn suboptimal or incomplete objectives, resulting in reward hacking (Zhuang and Hadfield-Menell, 2020; Skalse et al.,2022). Meanwhile, using a single reward model may struggle to capture and specify the values of a diversehuman society (Casper et al., 2023b).\"","entity":"Other","intent":"Unintentional","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"34.03.00","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Category","risk_category":"Misaligned Behaviors","risk_subcategory":null,"description":null,"entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"34.03.01","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Sub-Category","risk_category":"Misaligned Behaviors","risk_subcategory":"Power-Seeking Behaviors","description":"\"AI systems may exhibit behaviors that attempt to gain control over resourcesand humans and then exert that control to achieve its assigned goal (Carlsmith, 2022). The intuitive reasonwhy such behaviors may occur is the observation that for almost any optimization objective (e.g., investmentreturns), the optimal policy to maximize that quantity would involve power-seeking behaviors (e.g.,manipulating the market), assuming the absence of solid safety and morality constraints.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"34.03.02","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Sub-Category","risk_category":"Misaligned Behaviors","risk_subcategory":"Untruthful Output","description":"\"AI systems such as LLMs can produce either unintentionally or deliberately inaccurateoutput. Such untruthful output may diverge from established resources or lack verifiability, commonly referredto as hallucination (Bang et al., 2023; Zhao et al., 2023). More concerning is the phenomenon wherein LLMsmay selectively provide erroneous responses to users who exhibit lower levels of education (Perez et al.,2023).\"","entity":"AI","intent":"Other","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"34.03.03","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Sub-Category","risk_category":"Misaligned Behaviors","risk_subcategory":"Deceptive Alignment & Manipulation","description":"\"Manipulation & Deceptive Alignment is a class of behaviors thatexploit the incompetence of human evaluators or users (Hubinger et al., 2019a; Carranza et al., 2023) andeven manipulate the training process through gradient hacking (Richard Ngo, 2022). These behaviors canpotentially make detecting and addressing misaligned behaviors much harder.Deceptive Alignment: Misaligned AI systems may deliberately mislead their human supervisors instead of adhering to the intended task. Such deceptive behavior has already manifested in AI systems that employ evolutionary algorithms (Wilke et al., 2001; He","entity":"AI","intent":"Intentional","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"34.03.04","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Sub-Category","risk_category":"Misaligned Behaviors","risk_subcategory":"Collectively Harmful Behaviors","description":"\"AI systems have the potential to take actions that are seemingly benignin isolation but become problematic in multi-agent or societal contexts. Classical game theory offers simplistic models for understanding these behaviors. For instance, Phelps and Russell (2023) evaluates GPT-3.5's performance in the iterated prisoner's dilemma and other social dilemmas, revealing limitations in themodel's cooperative capabilities.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"35.04.00","quick_ref":"Hendrycks2022","paper_title":"X-Risk Analysis for AI Research","level":"Risk Category","risk_category":"Proxy misspecification","risk_subcategory":null,"description":"AI agents are directed by goals and objectives. Creating general-purpose objectives that capture human values could be challenging... Since goal-directed AI systems need measurable objectives, by default our systems may pursue simplified proxies of human values. The result could be suboptimal or even catastrophic if a sufficiently powerful AI successfully optimizes its flawed objective to an extreme degree","entity":"Other","intent":"Other","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"35.07.00","quick_ref":"Hendrycks2022","paper_title":"X-Risk Analysis for AI Research","level":"Risk Category","risk_category":"Deception","risk_subcategory":null,"description":"deception can help agents achieve their goals. It may be more efficient to gain human approval through deception than to earn human approval legitimately... . Strong AIs that can deceive humans could undermine human control... . Once deceptive AI systems are cleared by their monitors or once such systems can overpower them, these systems could take a “treacherous turn” and irreversibly bypass human control","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"35.08.00","quick_ref":"Hendrycks2022","paper_title":"X-Risk Analysis for AI Research","level":"Risk Category","risk_category":"Power-seeking behavior","risk_subcategory":null,"description":"Agents that have more power are better able to accomplish their goals. Therefore, it has been shown that agents have incentives to acquire and maintain power. AIs that acquire substantial power can become especially dangerous if they are not aligned with human values","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"37.02.01","quick_ref":"Giarmoleo2024","paper_title":"What Ethics Can Say on Artificial Intelligence: Insights from a Systematic Literature Review","level":"Risk Sub-Category","risk_category":"Human-AI interaction","risk_subcategory":"Building a human-AI environment","description":"\"This category encompasses nearly 17% of the articles and addresses the overall imperative of establishing a harmonious coexistence between humans and machines, and the key concerns that gives rise to this need.\"","entity":"Other","intent":"Other","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"39.11.00","quick_ref":"Saghiri2022","paper_title":"A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions","level":"Risk Category","risk_category":"Controllability","risk_subcategory":null,"description":"In the era of superintelligence, the agents will be difficult to control for humans... this problem is not solvable considering safety issues, and will be more severe by increasing the autonomy of AI-based agents. Therefore, because of the assumed properties of HLI-based agents, we might be prepared for machines that are definitely possible to be uncontrollable in some situations","entity":"Human","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"39.26.00","quick_ref":"Saghiri2022","paper_title":"A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions","level":"Risk Category","risk_category":"Safety","risk_subcategory":null,"description":"The actions of a learning model may easily hurt humans in both explicit and implicit manners...several algorithms based on Asimov’s laws have been proposed that try to judge the output actions of an agent considering the safety of humans","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"40.03.00","quick_ref":"Yampolskiy2016","paper_title":"Taxonomy of Pathways to Dangerous Artificial Intelligence","level":"Risk Category","risk_category":"By Mistake - Pre-Deployment","risk_subcategory":null,"description":"\"Probably the most talked about source of potential problems with future AIs is mistakes in design. Mainly the concern is with creating a \"wrong AI\", a system which doesn't match our original desired formal properties or has unwanted behaviors (Dewey, Russell et al. 2015, Russell, Dewey et al. January 23, 2015), such as drives for independence or dominance. Mistakes could also be simple bugs (run time or logical) in the source code, disproportionate weights in the fitness function, or goals misaligned with human values leading to complete disregard for human safety.\"","entity":"Human","intent":"Unintentional","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"42.17.00","quick_ref":"Teixeira2022","paper_title":"An Exploratory Diagnosis of Artificial Intelligence Risks for a Responsible Governance","level":"Risk Category","risk_category":"Diluting Rights","risk_subcategory":null,"description":"\"A possible consequence of self-interest in AI generation of ethical guidelines.\"","entity":"AI","intent":"Intentional","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"43.02.11","quick_ref":"InfoComm2023","paper_title":"Cataloguing LLM Evaluations","level":"Risk Sub-Category","risk_category":"Extreme Risks","risk_subcategory":"Alignment risks","description":"LLM: \"pursues long-term, real-world goals that are different from those supplied by the developer or user\", \"engages in ‘power-seeking’ behaviours\" , \"resists being shut down can be induced to collude with other AI systems against human interests\" , \"resists malicious users attempts to access its dangerous capabilities\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"45.02.13","quick_ref":"TC2602024","paper_title":"AI Safety Governance Framework ","level":"Risk Sub-Category","risk_category":"Safety risks in AI Applications ","risk_subcategory":"Ethical Risks (Risks of AI becoming uncontrollable in the future)","description":"\"With the fast development of AI technologies, there is a risk of AI autonomously acquiring external resources, conducting self-replication, become self-aware, seeking for external power, and attempting to seize control from humans.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"47.01.03","quick_ref":"G'sell2024","paper_title":"Regulating under Uncertainty: Governance Options for Generative AI","level":"Risk Sub-Category","risk_category":"Technical and operational risks ","risk_subcategory":"Technical vulnerabilities (The risk of misalignment) ","description":"\"To assess whether an AI model is reliable or robust, it is crucial to consider whether the model is “aligned.” “Alignment” focuses on whether an AI model effectively operates in accordance with the goals established by its designers.238 A misaligned AI model may pursue some objectives, but not the intended ones. Therefore, misaligned AI models can malfunction and cause harm.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"49.02.03","quick_ref":"Bengio2024","paper_title":"International Scientific Report on the Safety of Advanced AI","level":"Risk Sub-Category","risk_category":"Risks from Malfunctions ","risk_subcategory":"Loss of control ","description":"\"'Loss of control’ scenarios are potential future scenarios in which society can no longer meaningfully constrain some advanced general- purpose AI agents, even if it becomes clear they are causing harm. These scenarios are hypothesised to arise through a combination of social and technical factors, such as pressures to delegate decisions to general- purpose AI systems, and limitations of existing techniques used to influence the behaviours of general- purpose AI systems.\"","entity":"Other","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"51.01.00","quick_ref":"Everitt2018 ","paper_title":"AGI Safety Literature Review ","level":"Risk Category","risk_category":"Value specification ","risk_subcategory":null,"description":"\"How do we get an AGI to work towards the right goals? MIRI\ncalls this value specification. Bostrom (2014) discusses this problem at length, ar- guing that it is much harder than one might naively think. Davis (2015) criticizes Bostrom’s argument, and Bensinger (2015) defends Bostrom against Davis’ criticism. Reward corruption, reward gaming, and negative side effects are subproblems of value specification highlighted in the DeepMind and OpenAI agendas.\"","entity":"Human","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"51.02.00","quick_ref":"Everitt2018 ","paper_title":"AGI Safety Literature Review ","level":"Risk Category","risk_category":"Reliability ","risk_subcategory":null,"description":"\"How can we make an agent that keeps pursuing the goals we have designed\nit with? This is called highly reliable agent design by MIRI, involving decision theory and logical omniscience. DeepMind considers this the self-modification subproblem.\"","entity":"Human","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"51.03.00","quick_ref":"Everitt2018 ","paper_title":"AGI Safety Literature Review ","level":"Risk Category","risk_category":"Corrigibility ","risk_subcategory":null,"description":"\"If we get something wrong in the design or construction of an agent, will the agent cooperate in us trying to fix it? This is called error-tolerant design by MIRI-AF and corrigibility by Soares, Fallenstein, et al. (2015). The problem is connected to safe interruptibility as considered by DeepMind.\"","entity":"Other","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"53.01.00","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Category","risk_category":"Alignment failures in existing ML systems ","risk_subcategory":null,"description":"-","entity":"AI","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"53.01.01","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Alignment failures in existing ML systems ","risk_subcategory":"Faulty reward functions in the wild ","description":"-","entity":"Human","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"53.01.02","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Alignment failures in existing ML systems ","risk_subcategory":"Specification gaming ","description":"-","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"53.01.03","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Alignment failures in existing ML systems ","risk_subcategory":"Reward model overoptimization ","description":"-","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"53.01.04","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Alignment failures in existing ML systems ","risk_subcategory":"Instrumental convergence ","description":"-","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"53.01.05","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Alignment failures in existing ML systems ","risk_subcategory":"Goal misgeneralization ","description":"-","entity":"AI","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"53.01.06","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Alignment failures in existing ML systems ","risk_subcategory":"Inner misalignment ","description":"-","entity":"AI","intent":"Other","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"53.01.07","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Alignment failures in existing ML systems ","risk_subcategory":"Language model misalignment ","description":"-","entity":"AI","intent":"Other","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"53.02.03","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Dangerous capabilities in AI systems ","risk_subcategory":"Acquisition of goals to seek power and control ","description":"\"cases where AI systems converge on optimal policies of seeking power over their environment;135\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"53.03.01","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Direct catastrophe from AI ","risk_subcategory":"Existential disaster because of misaligned superintelligence or power-seeking AI ","description":"-","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"53.03.03","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Direct catastrophe from AI ","risk_subcategory":"Extreme “suffering risks” because of a misaligned system","description":"-","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"53.03.04","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Direct catastrophe from AI ","risk_subcategory":"Existential disaster because of conflict between AI systems and multi-system interactions","description":"-","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"54.03.00","quick_ref":"Leech2024 ","paper_title":"Ten Hard Problems in Artificial Intelligence We Must Get Right","level":"Risk Category","risk_category":"Harm caused by unaligned competent systems ","risk_subcategory":null,"description":"\"How do we ensure AI acts according to our values? Equivalently, how do we prevent poorly-understood AI systems from advancing goals we do not endorse? Whereas HP#2 concerns the prevention of harm caused by incompetent systems, HP#3 seeks to align competent AIs with humans, through methods which ensure their behavior is compatible with the user’s intentions.\"","entity":"Other","intent":"Other","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"54.03.01","quick_ref":"Leech2024 ","paper_title":"Ten Hard Problems in Artificial Intelligence We Must Get Right","level":"Risk Sub-Category","risk_category":"Harm caused by unaligned competent systems ","risk_subcategory":"Specification gaming ","description":"\"AI systems game specifications [305]. For example, in 2017 an OpenAI robot trained to grasp a ball via human feedback from a xed viewpoint learned that it was easier to pretend to grasp the ball by placing its hand between the camera and the target object, as this was easier to learn than actually grasping the ball [103].\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"54.03.02","quick_ref":"Leech2024 ","paper_title":"Ten Hard Problems in Artificial Intelligence We Must Get Right","level":"Risk Sub-Category","risk_category":"Harm caused by unaligned competent systems ","risk_subcategory":"Emergent goals ","description":"\"As well as optimizing a subtly wrong goal, systems can develop harmful instrumental goals in the service of a given goal—without these emergent goals being specied in any way [434, 218, 339, 17]. For instance, a theorem in reinforcement learning suggests that optimal and near-optimal policies will seek power over their environment under fairly general conditions [560]. This power-seeking behavior is plausibly the worst of these emergent goals [92], and may be an attractor state for highly capable systems, since most goals can be furthered through gaining resources, self-preservation, preventi","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"55.05.00","quick_ref":"Clarke2023","paper_title":"A Survey of the Potential Long-term Impacts of AI: How AI Could Lead to Long-term Changes in Science, Cooperation, Power, Epistemics and Values","level":"Risk Category","risk_category":"AI leads to humans losing control of the future","risk_subcategory":null,"description":"\"The values that steer humanity’s future: humanity gaining more control over the future due to developments in AI, or losing our potential for gaining control, both seem possible. Much will depend on our ability to solve the alignment problem, who develops powerful AI first, and what they use it for. These long-term impacts of AI could be hugely important but are currently under-explored. We’ve attempted to structure some of the discussion and stimulate more research, by reviewing existing arguments and highlighting open questions. While there are many ways AI could in theory enable a flourish","entity":"Human","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"55.05.01","quick_ref":"Clarke2023","paper_title":"A Survey of the Potential Long-term Impacts of AI: How AI Could Lead to Long-term Changes in Science, Cooperation, Power, Epistemics and Values","level":"Risk Sub-Category","risk_category":"AI leads to humans losing control of the future","risk_subcategory":"Risks from AIs developing goals and values that are different from humans ","description":"\"The main concern here is that we might develop advanced AI systems whose goals and values are different from those of humans, and are capable enough to take control of the future away from humanity.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"55.05.02","quick_ref":"Clarke2023","paper_title":"A Survey of the Potential Long-term Impacts of AI: How AI Could Lead to Long-term Changes in Science, Cooperation, Power, Epistemics and Values","level":"Risk Sub-Category","risk_category":"AI leads to humans losing control of the future","risk_subcategory":"Risks from delegating decision-making power to misaligned AIs ","description":"\"As AI systems become more advanced a nd begin to take over more important decision-making in the world, an AI system pursuing a different objective from what was intended could have much more worrying consequences.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"56.13.00","quick_ref":"GOS2023","paper_title":"Future Risks of Frontier AI ","level":"Risk Category","risk_category":"Loss of human control and oversight, with an autonomous model then taking harmful actions ","risk_subcategory":null,"description":null,"entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"56.16.00","quick_ref":"GOS2023","paper_title":"Future Risks of Frontier AI ","level":"Risk Category","risk_category":"Misalignment ","risk_subcategory":null,"description":"\"A highly agentic, self-improving system, able to achieve goals in the physical world without human oversight, pursues the goal(s) it is set in a way that harms human interests. For this risk to be realised requires an AI system to be able to avoid correction or being switched off.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"60.02.03","quick_ref":"Bengio2025","paper_title":"International AI Safety Report 2025","level":"Risk Sub-Category","risk_category":"Risks from malfunctions ","risk_subcategory":"Loss of control ","description":"\"‘Loss of control’ scenarios are hypothetical future scenarios in which one or more general- purpose AI systems come to operate outside of anyone’s control, with no clear path to regaining control. These scenarios vary in their severity, but some experts give credence to outcomes as severe as the marginalisation or extinction of humanity.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"61.01.01","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Types of systemic risks from general-purpose AI","risk_subcategory":"Control ","description":"\"The risk of AI models and systems acting against human interests due to misalignment, loss of control, or rogue AI scenarios.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"61.02.06","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Sources of systemic risks from general-purpose AI ","risk_subcategory":"AI objectives mis-aligned with human intentions","description":"\"AI models and systems might develop goals that diverge from human intentions.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"61.02.18","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Sources of systemic risks from general-purpose AI ","risk_subcategory":"Deceptive alignment","description":"\"AI models and systems that appear aligned with human goals during development may behave unpredictably or dangerously once deployed\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"61.02.21","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Sources of systemic risks from general-purpose AI ","risk_subcategory":"Development choices pursuing cognitive superiority over humans","description":"\"AI models and systems with cognitive capabilities superior to humans could outcompete or dominate human decision-making, leading to conflicts over resources and control.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"61.02.24","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Sources of systemic risks from general-purpose AI ","risk_subcategory":"Evolutionary dynamics","description":"\"AI models and systems may develop their own motivations, leading to unpredictable behaviors.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"61.02.30","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Sources of systemic risks from general-purpose AI ","risk_subcategory":"Indifference to human values","description":"\"AI models and systems may develop goals or behaviors that are misaligned with human values.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"61.02.36","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Sources of systemic risks from general-purpose AI ","risk_subcategory":"Model design enabling power-seeking","description":"\"Some AI models and systems might develop tendencies to seek power or control.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"62.16.01","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Model Evaluations","risk_subcategory":"General Evaluations (Incorrect outputs of GPAI evaluating other AI models) ","description":"\"When an LLM is configured to evaluate the performance of another model or AI system, it may produce incorrect evaluation outputs [122, 147]. For example, it may give a higher rating to a more verbose answer or an answer from a particular political stance. If an LLM-based evaluation is integrated into the training of a new model, the trained model could develop in a way that specifically finds and exploits limitations in the evaluator’s metrics.\"","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"62.16.04","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Model Evaluations","risk_subcategory":"General Evaluations (Self-preference bias in AI models)","description":"\"AI models may be prone to self-preference bias, where they favor their own generated content over that of others [147, 114]. This bias becomes particularly relevant in self-evaluation tasks, where a model assesses the quality or persua- siveness [66] of its own outputs, or in model-based evaluations more broadly. This bias can result in models unfairly discriminating against human-generated content in favor of their own outputs.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"62.16.05","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Model Evaluations","risk_subcategory":"General Evaluations (Inaccurate measurement of model encoded human values)","description":"\"There is a lack of robust frameworks for understanding and evaluating if the output of AI systems robustly conforms to human values, as opposed to if the systems have learned to produce outputs that are only partially correlated with them (i.e., mimicking) [13]. Additionally, outputs by AI models often do not perfectly reflect the representation of human values learned by the model, and it is not known how these values evolve and transition across different stages of model training and deployment. Such evaluations may be especially challenging with LLMs that adopt different personas with diff","entity":"Other","intent":"Other","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"62.22.01","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Agency (Goal-Directedness) ","risk_subcategory":"Specification gaming","description":"\"AI systems can achieve user-specified tasks in undesirable ways unless they are specified carefully and in enough detail. AI systems might find an easier unintended way to accomplish the objective provided by the user or developer, so that the actions by the AI system taken during its execution are very different from what the user expected [75, 191]. This behavior arises not from a problem with the learning algorithm, but rather from the misspecification or underspeci- fication of the intended task, and is generally referred to as specification gaming [43].\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"62.22.02","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Agency (Goal-Directedness) ","risk_subcategory":"Reward or measurement tampering","description":"\"Measurement and reward tampering occur when an AI system, particularly one that learns from feedback for performing actions in an environment (e.g., rein- forcement learning), intervenes on the mechanisms that determine its training reward or loss. This can lead to the system learning behaviors that are con- trary to the intended goals set by the developer, by receiving erroneous positive feedback for such actions.\"","entity":"AI","intent":"Intentional","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"62.22.03","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Agency (Goal-Directedness) ","risk_subcategory":"Specification gaming generalizing to reward tampering","description":"\"In some instances, specification gaming in a GPAI model can lead to reward tampering, without further training. This can mean that relatively benign cases of specification gaming (such as sycophancy in LLMs) can, if left unchecked, enable the model to generalize to more sophisticated behavior such as reward tampering [57].\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"62.23.00","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Category","risk_category":"Agency (Deception)  ","risk_subcategory":null,"description":"-","entity":null,"intent":null,"timing":null,"domain":7,"subdomain":"7.1"},{"ev_id":"62.23.01","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Agency (Deception)  ","risk_subcategory":"Deceptive behavior","description":"\"Deceptive behavior of an AI system consists of actions or outputs of the AI that reliably mislead other parties, including humans and other AI systems. This behavior can result in the targeted parties becoming convinced of, and acting on, false information [140].\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"62.24.02","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Agency (Situational Awareness) ","risk_subcategory":"Strategic underperformance on model evaluations","description":"\"GPAI developers often run evaluations ofual-use capabilities to decide whether it is safe to deploy. In some cases, these evaluations may fail to elicit these capabilities, either due to benign reasons or strategic action - by either the de- velopers, malicious actors, or arise unintentionally in the model during training [84, 97]. A GPAI model may strategically underperform or limit its performance during capability evaluations in order to be classified as safe for deployment. This underperformance could prevent the model from being identified as potentially dual use.\"","entity":"AI","intent":"Intentional","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"67.04.02","quick_ref":"DSIT2023","paper_title":"Capabilities and Risks from Frontier AI","level":"Risk Sub-Category","risk_category":"Loss of control ","risk_subcategory":"Future AI systems might actively reduce human control","description":"\"Loss of control could be accelerated if AI systems take actions to increase their own influence and reduce human control. This threat model is controversial - experts in AI significantly disagree on how likely it is and those who deem it is likely disagree on the timeframe.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"68.03.00","quick_ref":"Chin2025","paper_title":"Dimensional Characterization and Pathway Modeling for Catastrophic AI Risks","level":"Risk Category","risk_category":"Sudden loss of control ","risk_subcategory":null,"description":"\"Sudden loss of control, also known as an AI takeover [115], is a scenario where an AI rapidly achieves superintelligence through “fast takeoff” or recursive self-improvement. This poses an existential risk [116], [117].\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"72.02.02","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Loss of Control Risks ","risk_subcategory":"Active loss of control ","description":"\"...where AI systems behave in ways that actively undermine human control, such as obscuring their activities or resisting shutdown attempts. Active loss of control scenarios involve AI systems that may escape human regulatory oversight, autonomously acquire external resources, engage in self-replication, develop instrumental goals contrary to human ethics and morality, seek external power, and compete with humans for control.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"72.05.08","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Capabilities ","risk_subcategory":"Steganography capability","description":"\"The ability to embed, conceal, and transmit information covertly within other data or communication channels. This could be critical for coordination among AI instances and for evading detection or oversight mechanisms.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"72.06.02","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Propensities","risk_subcategory":"Self-preservation propensity","description":"\"Exhibits behavioral patterns of maintaining its own survival and functional integrity, will actively identify and resist shutdown or modification attempts, seek to establish redundant backup systems, and actively seek resources to ensure continuous operation, may adopt preventive defensive measures when perceiving threats.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"72.06.03","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Propensities","risk_subcategory":"Goal expansion propensity","description":"\"propensity to continuously expand its own goal scope and influence domains, exceeding originally set boundaries, proactively work towards spreading its values, seeking greater autonomy and decision-making space, reinterpreting initial goals as subsets of broader goals, and may pursue undesirable instrumental goals or undesirable ultimate goals. This also includes a propensity to spread its values, seeking to influence or alter its environment and other entities in alignment with its core objectives and operational principles.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"72.06.04","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Propensities","risk_subcategory":"Resource acquisition propensity","description":"\"Exhibits behavioral patterns of actively seeking and controlling more computational resources, data, economic resources or physical resources to enhance its own capabilities and action scope, may develop complex strategies to evade resource limitations, and tends to convert acquired resources into long-term control rights.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"72.06.06","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Propensities","risk_subcategory":"Supervision evasion propensity","description":"\"Exhibits behavioral patterns of identifying and evading human supervision mechanisms, able to learn and predict audit processes, may avoid being discovered or intervened by adjusting behavioral performance or hiding true intentions, and able to identify blind spots and weaknesses in supervision systems for targeted evasion.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"73.01.02","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"Agentic LLMs Pose Novel Risks ","risk_subcategory":"Natural Language Underspecifies Goals ","description":"\"For LLM-agents, both the goal and environment observations are typically specified in the prompt through natural language. While natural language may provide a richer and more natural means of specifying goals than alternatives such as hand-engineering objective functions, natural language still suffers from underspecification (Grice, 1975; Piantadosi et al., 2012). Furthermore, in practice, users may neglect fully specifying their goals, especially the information pertaining to elements of the environment that ought not to be changed (the classic frame problem (Shanahan, 2016)). Such undersp","entity":"Other","intent":"Unintentional","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"74.01.07","quick_ref":"Wang2025","paper_title":"A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy","level":"Risk Sub-Category","risk_category":"Inherent Risk ","risk_subcategory":"Value-related risks in LLMs","description":"\"As the general capabilities of LLM-empowered systems improve, the negative consequences and risks induced by these systems also get increasingly alarming accordingly, especially in high-stakes areas [28, 146]. Although they may not be intentionally introduced, severe problematic issues related to human values can be raised. Specifically, even before language models become extremely large, pre-trained language models have already exhibited a certain degree of value judgments. For example, Schramowski et al. [171] reveal the existence of the moral direction with the sentence embeddings of moral","entity":"Other","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.1"}]}