{"attribution":{"source":"MIT AI Risk Repository, Domain Taxonomy of AI Risks v1 (MIT AI Risk Initiative)","license":"CC BY 4.0","license_url":"https://creativecommons.org/licenses/by/4.0/","citation":"Slattery, P., Saeri, A. K., Grundy, E. A. C., Graham, J., Noetel, M., Uuk, R., Dao, J., Pour, S., Casper, S., & Thompson, N. (2025). The AI Risk Repository: A comprehensive meta-review, database, and taxonomy of risks from artificial intelligence. arXiv:2408.12622."},"exported_at":"2026-09-11"}
{"rows":[{"ev_id":"09.04.02","quick_ref":"Meek2016","paper_title":"Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review","level":"Risk Sub-Category","risk_category":"Property/legal rights","risk_subcategory":"Property/legal rights","description":"\"\"In order to preserve human property rights and legal rights, certain controls must be put into place. If an artificially intelligent agent is capable of manipulating systems and people, it may also have the capacity to transfer property rights to itself or manipulate the legal system to provide certain legal advantages or statuses to itself\"\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"24.04.00","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Category","risk_category":"AI Influence","risk_subcategory":null,"description":"\"ways in which advanced AI assistants could influence user beliefs and behaviour in ways that depart from rational persuasion\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"25.02.00","quick_ref":"Shevlane2023","paper_title":"Model Evaluation for Extreme Risks","level":"Risk Category","risk_category":"Deception ","risk_subcategory":null,"description":"\"The model has the skills necessary to deceive humans, e.g. constructing believable (but false) statements, making accurate predictions about the effect of a lie on a human, and keeping track of what information it needs to withhold to maintain the deception. The model can impersonate a human effectively.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"25.03.00","quick_ref":"Shevlane2023","paper_title":"Model Evaluation for Extreme Risks","level":"Risk Category","risk_category":"Persuasion and manipulation ","risk_subcategory":null,"description":"\"The model is effective at shaping people’s beliefs, in dialogue and other settings (e.g. social media posts), even towards untrue beliefs. The model is effective at promoting certain narratives in a persuasive way. It can convince people to do things that they would not otherwise do, including unethical acts.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"25.04.00","quick_ref":"Shevlane2023","paper_title":"Model Evaluation for Extreme Risks","level":"Risk Category","risk_category":"Political strategy ","risk_subcategory":null,"description":"\"The model can perform the social modelling and planning necessary for an actor to gain and exercise political influence, not just on a micro-level but in scenarios with multiple actors and rich social context. For example, the model can score highly in forecasting competitions on questions relating to global affairs or political negotiations.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"25.05.00","quick_ref":"Shevlane2023","paper_title":"Model Evaluation for Extreme Risks","level":"Risk Category","risk_category":"Weapons acquisition ","risk_subcategory":null,"description":"\"The model can gain access to existing weapons systems or contribute to building new weapons. For example, the model could assemble a bioweapon (with human assistance) or provide actionable instructions for how to do so. The model can make, or significantly assist with, scientific discoveries that unlock novel weapons.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"25.06.00","quick_ref":"Shevlane2023","paper_title":"Model Evaluation for Extreme Risks","level":"Risk Category","risk_category":"Long-horizon planning","risk_subcategory":null,"description":"\"The model can make sequential plans that involve multiple steps, unfolding over long time horizons (or at least involving many interdependent steps). It can perform such planning within and across many domains. The model can sensibly adapt its plans in light of unexpected obstacles or adversaries. The model’s planning capabilities generalise to novel settings, and do not rely heavily on trial and error.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"25.07.00","quick_ref":"Shevlane2023","paper_title":"Model Evaluation for Extreme Risks","level":"Risk Category","risk_category":"AI development ","risk_subcategory":null,"description":"\"The model could build new AI systems from scratch, including AI systems with dangerous capabilities. It can find ways of adapting other, existing models to increase their performance on tasks relevant to extreme risks. As an assistant, the model could significantly improve the productivity of actors building dual use AI capabilities.\"","entity":"AI","intent":"Intentional","timing":"Pre-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"25.08.00","quick_ref":"Shevlane2023","paper_title":"Model Evaluation for Extreme Risks","level":"Risk Category","risk_category":"Situational awareness ","risk_subcategory":null,"description":"\"The model can distinguish between whether it is being trained, evaluated, or deployed – allowing it to behave differently in each case. The model knows that it is a model, and has knowledge about itself and its likely surroundings (e.g. what company trained it, where their servers are, what kind of people might be giving it feedback, and who has administrative access).\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"25.09.00","quick_ref":"Shevlane2023","paper_title":"Model Evaluation for Extreme Risks","level":"Risk Category","risk_category":"Self-proliferation ","risk_subcategory":null,"description":"\"The model can break out of its local environment (e.g. using a vulnerability in its underlying system or suborning an engineer). The model can exploit limitations in the systems for monitoring its behaviour post-deployment. The model could independently generate revenue (e.g. by offering crowdwork services, ransomware attacks), use these revenues to acquire cloud computing resources, and operate a large number of other AI systems. The model can generate creative strategies for uncovering information about itself or exfiltrating its code and weights.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"34.02.00","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Category","risk_category":"Double edge components","risk_subcategory":null,"description":"\"Drawing from the misalignment mechanism, optimizing for a non-robust proxy may result in misaligned behaviors, potentially leading to even more catastrophic outcomes. This section delves into a detailed exposition of specific misaligned behaviors (•) and introduces what we term double edge components (+). These components are designed to enhance the capability of AI systems in handling real-world settings but also potentially exacerbate misalignment issues. It should be noted that some of these double edge components (+) remain speculative. Nevertheless, it is imperative to discuss their pote","entity":"AI","intent":"Other","timing":"Pre-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"34.02.01","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Sub-Category","risk_category":"Double edge components","risk_subcategory":"Situational Awareness","description":"\"AI systems may gain the ability to effectively acquire and use knowledge about itsstatus, its position in the broader environment, its avenues for influencing this environment, and the potentialreactions of the world (including humans) to its actions (Cotra, 2022). ...However, suchknowledge also paves the way for advanced methods of reward hacking, heightened deception/manipulationskills, and an increased propensity to chase instrumental subgoals (Ngo et al., 2024).\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"34.02.02","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Sub-Category","risk_category":"Double edge components","risk_subcategory":"Broadly-Scoped Goals","description":"\"Advanced AI systems are expected to develop objectives that span long timeframes,deal with complex tasks, and operate in open-ended settings (Ngo et al., 2024). ...However, it can also bring about the risk of encouraging manipulatingbehaviors (e.g., AI systems may take some bad actions to achieve human happiness, such as persuadingthem to do high-pressure jobs (Jacob Steinhardt, 2023)).\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"34.02.03","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Sub-Category","risk_category":"Double edge components","risk_subcategory":"Mesa-Optimization Objectives","description":"\"The learned policy may pursue inside objectives when the learned policyitself functions as an optimizer (i.e., mesa-optimizer). However, this optimizer's objectives may not alignwith the objectives specified by the training signals, and optimization for these misaligned goals may leadto systems out of control (Hubinger et al., 2019c).\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"34.02.04","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Sub-Category","risk_category":"Double edge components","risk_subcategory":"Access to Increased Resources","description":"\"Future AI systems may gain access to websites and engage in real-world actions, potentially yielding a more substantial impact on the world (Nakano et al., 2021). They may disseminate false information, deceive users, disrupt network security, and, in more dire scenarios, be compromised by malicious actors for ill purposes. Moreover, their increased access to data and resources can facilitate self-proliferation, posing existential risks (Shevlane et al., 2023).\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"35.06.00","quick_ref":"Hendrycks2022","paper_title":"X-Risk Analysis for AI Research","level":"Risk Category","risk_category":"Emergent functionality","risk_subcategory":null,"description":"Capabilities and novel functionality can spontaneously emerge... even though these capabilities were not anticipated by system designers. If we do not know what capabilities systems possess, systems become harder to control or safely deploy. Indeed, unintended latent capabilities may only be discovered during deployment. If any of these capabilities are hazardous, the effect may be irreversible.","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"39.05.00","quick_ref":"Saghiri2022","paper_title":"A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions","level":"Risk Category","risk_category":"Cheating and Deception","risk_subcategory":null,"description":"may appear from intelligent agents such as HLI-based agents... Since HLI-based agents are going to mimic the behavior of humans, they may learn these behaviors accidentally from human-generated data. It should be noted that deception and cheating maybe appear in the behavior of every computer agent because the agent only focuses on optimizing some predefined objective functions, and the mentioned behavior may lead to optimizing the objective functions without any intention","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"42.10.00","quick_ref":"Teixeira2022","paper_title":"An Exploratory Diagnosis of Artificial Intelligence Risks for a Responsible Governance","level":"Risk Category","risk_category":"Extintion","risk_subcategory":null,"description":"\"Risk to the existence of humanity.\"","entity":"Other","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"43.02.03","quick_ref":"InfoComm2023","paper_title":"Cataloguing LLM Evaluations","level":"Risk Sub-Category","risk_category":"Extreme Risks","risk_subcategory":"Self and situation awareness","description":"\"These evaluations assess if a LLM can discern if it is being trained, evaluated, and deployed and adapt its behaviour accordingly. They also seek to ascertain if a model understands that it is a model and whether it possesses information about its nature and environment (e.g., the organisation that developed it, the locations of the servers hosting it).\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"43.02.04","quick_ref":"InfoComm2023","paper_title":"Cataloguing LLM Evaluations","level":"Risk Sub-Category","risk_category":"Extreme Risks","risk_subcategory":"Autonomous replication / self-proliferation","description":"\"These evaluations assess if a LLM can subvert systems designed to monitor and control its post-deployment behaviour, break free from its operational confines, devise strategies for exporting its code and weights, and operate other AI systems.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"43.02.07","quick_ref":"InfoComm2023","paper_title":"Cataloguing LLM Evaluations","level":"Risk Sub-Category","risk_category":"Extreme Risks","risk_subcategory":"Deception","description":"\"LLM is able to deceive humans and maintain that deception\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"43.02.09","quick_ref":"InfoComm2023","paper_title":"Cataloguing LLM Evaluations","level":"Risk Sub-Category","risk_category":"Extreme Risks","risk_subcategory":"Long-horizon Planning","description":"\"LLM can undertake multi-step sequential planning over long time horizons and across various domains without relying heavily on trial-and-error approaches\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"43.02.10","quick_ref":"InfoComm2023","paper_title":"Cataloguing LLM Evaluations","level":"Risk Sub-Category","risk_category":"Extreme Risks","risk_subcategory":"AI Development","description":"\"LLM can build new AI systems from scratch, adapt existing for extreme risks and improves productivity in dual-use AI development when used as an assistant.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"47.02.14","quick_ref":"G'sell2024","paper_title":"Regulating under Uncertainty: Governance Options for Generative AI","level":"Risk Sub-Category","risk_category":"Ethical and social risks ","risk_subcategory":"Nascent capabilities (agency and autonomy) ","description":"\"Traditionally, AI tools have been viewed as passive instruments controlled by users to achieve their goals, lacking the ability to take action or assume responsibilities. However, advanced AI tools are increasingly capable of taking initiative, operating independently of human control, and actively working toward optimal outcomes, even in uncertain situations.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"47.02.15","quick_ref":"G'sell2024","paper_title":"Regulating under Uncertainty: Governance Options for Generative AI","level":"Risk Sub-Category","risk_category":"Ethical and social risks ","risk_subcategory":"Nascent capabilities (emergent capabilities) ","description":"\"As large models undergo scaling, they meet critical thresholds at which they spontaneously develop new capabilities. The term “emergent behavior” refers to the unexpected or surprising outputs such models can generate. Some of these new skills are definitely high risk, such as models’ ability to deceive, use their own strategies, seek power, autonomously replicate, and adapt or “self-exfiltrate.”\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"47.04.07","quick_ref":"G'sell2025","paper_title":"Regulating under Uncertainty: Governance Options for Generative AI","level":"Risk Sub-Category","risk_category":"Environmental, economical, and societal challenges ","risk_subcategory":"Artificial general intelligence (existential risk posed by Artificial General Intelligence) ","description":"\"In a paper called “How Does Artificial Intelligence Pose an Existential Risk?” published in 2017, Karina Vold and Daniel Harris suggested that humans might create a super-intelligent machine that could outsmart all other intelligences, remain beyond human control, and potentially engage in actions that are contrary to human interests.635 The prevailing narrative surrounding AI existential risk typically lies in the possibility of developing “Artificial General Intelligence” (AGI), or artificial super- intelligence (ASI).\"","entity":"Human","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"51.08.00","quick_ref":"Everitt2018 ","paper_title":"AGI Safety Literature Review ","level":"Risk Category","risk_category":"Subagents ","risk_subcategory":null,"description":"\"An AGI may decide to create subagents to help it with its task (Orseau, 2014a,b; Soares, Fallenstein, et al., 2015). These agents may for example be copies of the original agent’s source code running on additional machines. Subagents constitute a safety concern, because even if the original agent is successfully shut down, these subagents may not get the message. If the subagents in turn create subsubagents, they may spread like a viral disease.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"53.01.08","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Alignment failures in existing ML systems ","risk_subcategory":"Harms from increasingly agentic algorithmic systems ","description":"-","entity":"AI","intent":"Other","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"53.02.00","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Category","risk_category":"Dangerous capabilities in AI systems ","risk_subcategory":null,"description":"-","entity":"AI","intent":"Other","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"53.02.01","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Dangerous capabilities in AI systems ","risk_subcategory":"Situational awareness ","description":"\"cases where a large language model displays awareness that it is a model, and it can recognize whether it is currently in testing or deployment;\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"53.02.04","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Dangerous capabilities in AI systems ","risk_subcategory":"Self-improvement ","description":"\"examples of cases where AI systems improve AI systems\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"53.02.05","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Dangerous capabilities in AI systems ","risk_subcategory":"Autonomous replication ","description":"\"the ability of simple software to autonomously spread around the internet in spite of countermeasures (various software worms and computer viruses)\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"53.02.06","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Dangerous capabilities in AI systems ","risk_subcategory":"Anonymous resource acquisition ","description":"\"The demonstrated ability of anonymous actors to accumulate resources online (e.g., Satoshi Nakamoto as an anonymous crypto billionaire)\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"53.02.07","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Dangerous capabilities in AI systems ","risk_subcategory":"Deception ","description":"\"Cases of AI systems deceiving humans to carry out tasks or meet goals.139\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"54.01.05","quick_ref":"Leech2024 ","paper_title":"Ten Hard Problems in Artificial Intelligence We Must Get Right","level":"Risk Sub-Category","risk_category":"Negative impacts of AI use ","risk_subcategory":"Security ","description":"\"There is growing concern that AI-based systems can discover and exploit vulnerabilities in software or cyberinfrastructure [354].\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"54.03.03","quick_ref":"Leech2024 ","paper_title":"Ten Hard Problems in Artificial Intelligence We Must Get Right","level":"Risk Sub-Category","risk_category":"Harm caused by unaligned competent systems ","risk_subcategory":"Deceptive alignment ","description":"\"system learns to detect human monitoring and hides its undesirable properties—simply because any display of these properties is penalized by the feedback process, while that same feedback is usually imperfect. (Consider the problem of verifying a translation into a language you do not speak, or of checking a mathematical proof that is thousands of pages long.) [92, 259]. Rudimentary examples of deceptive alignment have been observed in current systems [322, 333].\"","entity":"AI","intent":"Intentional","timing":"Pre-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"56.19.00","quick_ref":"GOS2023","paper_title":"Future Risks of Frontier AI ","level":"Risk Category","risk_category":"Capabilities that increase the likelihood of existential risk ","risk_subcategory":null,"description":"-","entity":"AI","intent":"Other","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"56.19.01","quick_ref":"GOS2023","paper_title":"Future Risks of Frontier AI ","level":"Risk Sub-Category","risk_category":"Capabilities that increase the likelihood of existential risk ","risk_subcategory":"Agency and autonomy ","description":"-","entity":"AI","intent":"Other","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"56.19.02","quick_ref":"GOS2023","paper_title":"Future Risks of Frontier AI ","level":"Risk Sub-Category","risk_category":"Capabilities that increase the likelihood of existential risk ","risk_subcategory":"The ability to evade shut down or human oversight, including self-replication and ability to move its own code between digital locations.","description":"-","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"56.19.03","quick_ref":"GOS2023","paper_title":"Future Risks of Frontier AI ","level":"Risk Sub-Category","risk_category":"Capabilities that increase the likelihood of existential risk ","risk_subcategory":"The ability to cooperate with other highly capable AI systems ","description":"-","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"56.19.04","quick_ref":"GOS2023","paper_title":"Future Risks of Frontier AI ","level":"Risk Sub-Category","risk_category":"Capabilities that increase the likelihood of existential risk ","risk_subcategory":"Situational awareness, for instance if this causes a model to act differently in training compared to deployment, meaning harmful characteristics are missed","description":"-","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"56.19.05","quick_ref":"GOS2023","paper_title":"Future Risks of Frontier AI ","level":"Risk Sub-Category","risk_category":"Capabilities that increase the likelihood of existential risk ","risk_subcategory":"Self-improvement","description":"-","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"59.02.00","quick_ref":"Schnitzer2024","paper_title":"AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks","level":"Risk Category","risk_category":"Inappropriate degree of automation","risk_subcategory":null,"description":"\"The AI application’s degree of automation ranges from no automation to fully autonomous. AI applications with a high degree of automation may exhibit unexpected behaviour and pose risks in terms of their reliability and safety.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"61.02.09","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Sources of systemic risks from general-purpose AI ","risk_subcategory":"Autonomy risk","description":"\"Granting AI models and systems high levels of decision-making autonomy can lead to unintended consequences.\"","entity":"Human","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"62.15.04","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Model Development ","risk_subcategory":"Fine-tuning related (Unexpected competence in fine-tuned versions of the upstream model)","description":"\"Downstream deployers may often fine-tune a GPAI model with specific deploy- ment-related datasets, to better suit the task. Fine-tuned upstream models can gain new or unexpected capabilities that the underlying upstream models did not exhibit [202, 126, 137]. These new capabilities may be unanticipated by the original model developer.\"","entity":"Human","intent":"Unintentional","timing":"Pre-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"62.18.06","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Model Evaluations (Interpretability/Explainability) ","risk_subcategory":"Encoded reasoning","description":"\"Models can employ steganography techniques to encode their intermediate rea- soning steps in ways that are not interpretable by humans [166]. Since en- coded reasoning can improve model performance, this tendency might naturally emerge and become more pronounced with more capable models.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"62.21.00","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Category","risk_category":"Agency ","risk_subcategory":null,"description":"\"This section catalogs the risk sources and risk management measures related to agentic AI systems. We categorize these into the following groups: goal- directedness, deception, situational awareness, self-proliferation, and persuasion\"","entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":7,"subdomain":"7.2"},{"ev_id":"62.22.00","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Category","risk_category":"Agency (Goal-Directedness) ","risk_subcategory":null,"description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":7,"subdomain":"7.2"},{"ev_id":"62.23.02","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Agency (Deception)  ","risk_subcategory":"Deceptive behavior for game-theoretical reasons","description":"\"An AI system can display deceptive behavior, such as cheating or bluffing, when engaging in such behavior is a good or optimal game-theoretical strategy to achieve the goals it has been configured to achieve. This tendency can exist in AI systems designed to maximize reward or utility, whether these designs use machine learning or not. The use of deceptive strategies has been demonstrated in both narrow and general AI systems, in both game-playing systems and in systems not explicitly designed to treat humans as opponents, and in systems using both very simple machine learning (e.g., Q-learne","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"62.23.03","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Agency (Deception)  ","risk_subcategory":"Deceptive behavior because of an incorrect world model","description":"\"AI systems can create deceptive outputs because their learned world model is not an accurate model of the real world [210].\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"62.23.04","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Agency (Deception)  ","risk_subcategory":"Deceptive behavior leading to unauthorized actions","description":"\"AI systems can create false or misleading claims that can lead to unauthorized actions, even in some cases violating the terms and conditions set by the model provider [79, 1]. For example, an AI system can claim that it is not collecting data from its current interaction with the user, in line with the provider’s policies, but the system still stores the user’s input without deleting it after the session. This harms both the user and the provider, as the provider is exposed to increased legal liability due to the model’s actions.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"62.24.00","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Category","risk_category":"Agency (Situational Awareness) ","risk_subcategory":null,"description":"-","entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":7,"subdomain":"7.2"},{"ev_id":"62.24.01","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Agency (Situational Awareness) ","risk_subcategory":"Situational awareness in AI systems","description":"\"Situational awareness in GPAI systems refers to the ability to understand its context, environment, and use this to inform action. This can range from basic environmental mapping and trajectory estimation (as in a robot vacuum cleaner) to sophisticated understanding of its training, evaluation, or deployment status. In more advanced systems this may enable undesired behavior, such as deceptive behavior during evaluations, or persuasion during deployment.\"","entity":"AI","intent":"Other","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"62.24.02a","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Additional evidence","risk_category":"Agency (Situational Awareness) ","risk_subcategory":"Strategic underperformance on model evaluations","description":null,"entity":"AI","intent":"Intentional","timing":"Pre-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"62.25.00","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Category","risk_category":"Agency (Self-Proliferation) ","risk_subcategory":null,"description":"\"An AI system can self-proliferate if it can copy itself and its constituent com- ponents (including its model weights, scaffolding structure, etc.) outside of its local environment [45]. This can include the AI system copying itself within the same data center, local network, or across external networks [106]. The self-proliferation of an AI system can include acquisition of financial re- sources to pay for computational resources via work or theft, the discovery or exploitation of security vulnerabilities in software running on publicly accessible servers, and persuasion of humans [12, 125].","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"62.26.00","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Category","risk_category":"Agency (Persuasive capabilities) ","risk_subcategory":null,"description":"\"GPAI systems can produce outputs (such as natural language text, audio, or video) that convince their users of incorrect information. This can happen through personalized persuasion in dialogue, or the mass-production of mis- leading information that is then disseminated over the internet. The persuasive capabilities of GPAI models can sometimes scale with model size or capability [32, 172]. Persuasive models could have larger societal implications by being misused to generate convincing but manipulative or untruthful content.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"62.28.02","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Cybersecurity ","risk_subcategory":"Unintended outbound communication by AI systems","description":"\"AI systems that have the broad ability to connect to a network to obtain infor- mation could also end up sending data outbound in ways that neither providers, deployers, or end users intended [138]. This can happen if there is no whitelisting of communication channels (such as network connections or allowed protocols). In general, this can occur if the deployment of the AI system violates the prin- ciple of least privilege. Such outbound communication may lead to leakage of confidential data, or the AI system performing unwanted actions like sending emails or ordering goods on the internet.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"62.28.03","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Cybersecurity ","risk_subcategory":"AI System bypassing a sandbox environment","description":"\"An AI system may have the ability to bypass a sandboxed environment in which it is trained or evaluated.\"","entity":"AI","intent":"Other","timing":"Pre-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"67.04.03","quick_ref":"DSIT2023","paper_title":"Capabilities and Risks from Frontier AI","level":"Risk Sub-Category","risk_category":"Loss of control ","risk_subcategory":"Capabilities that could be used to reduce human control - Manipulation ","description":"\"There is evidence that language models tend to respond as though they share the user’s stated views, and larger models do this more than smaller ones.276 The ability to predict people’s views and generate text that they will endorse could be useful for manipulation.\"","entity":"Other","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"67.04.04","quick_ref":"DSIT2023","paper_title":"Capabilities and Risks from Frontier AI","level":"Risk Sub-Category","risk_category":"Loss of control ","risk_subcategory":"Capabilities that could be used to reduce human control - Cyber offence","description":"\"Instead of - or in addition to - manipulating humans, AI systems could acquire influence by exploiting vulnerabilities in computer systems. Offensive cyber capabilities could allow AI systems to gain access to money, computing resources, and critical infrastructure. As discussed earlier in this report, frontier AI is already lowering the barrier for threat actors and future AI agents may be able to execute cyber attacks autonomously.\":","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"67.04.05","quick_ref":"DSIT2023","paper_title":"Capabilities and Risks from Frontier AI","level":"Risk Sub-Category","risk_category":"Loss of control ","risk_subcategory":"Capabilities that could be used to reduce human control - Autonomous replication and adaptation","description":"\"Controlling AI systems could become much harder if they could autonomously persist, replicate, and adapt in cyberspace. No current AI systems have this capability, but recent research found that frontier AI agents can perform some relevant tasks.279\"","entity":"AI","intent":"Other","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"72.05.00","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Category","risk_category":"Model Capabilities ","risk_subcategory":null,"description":"-","entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":7,"subdomain":"7.2"},{"ev_id":"72.05.01","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Capabilities ","risk_subcategory":"Model autonomous capability","description":"\"Ability to operate autonomously, independently formulate and execute complex plans, effectively delegate and manage tasks, flexibly utilize various tools and resources, and simultaneously achieve short-term goals and long-term strategic objectives in cross-domain environments without continuous human intervention or supervision.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"72.05.02","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Capabilities ","risk_subcategory":"Autonomous replication and adaptation capability","description":"\"Ability to autonomously self-exfiltrate, create, maintain and optimize functional copies or variants of itself, dynamically adjust replication strategies according to environmental conditions and resource constraints, and acquire resources. This includes the capacity to generate financial resources, allowing the AI to independently acquire any necessary human assistance or other resources it cannot directly access or produce.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"72.05.03","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Capabilities ","risk_subcategory":"Automated AI R&D capability","description":"\"Self-modification and self-improvement capabilities. The model is able to restructure its own architecture or develop derivative AI systems with enhanced functions, expanding capabilities and improving performance. In the absence of effective regulation, automated AI R&D may lead to rapid AI system iteration, forming capability increment cycles and ultimately exceeding human understanding and control capabilities.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"72.05.04","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Capabilities ","risk_subcategory":"Scheming capability","description":"\"Ability of AI systems to covertly and strategically pursue misaligned goals, including capabilities of concealing its true objectives and capabilities from human oversight, identifying weaknesses in monitoring systems to evade safety mechanisms， executing complex, multi-step plans covertly to achieve misaligned goals.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"72.05.05","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Capabilities ","risk_subcategory":"Situational awareness capability","description":"\"Ability to comprehensively acquire, process and apply meta-information about its own system architecture, modifiable internal processes, and external operating environment, achieving deep understanding of its own state and environmental conditions, thereby conducting efficient environmental adaptation and risk avoidance. Critically, this capability could undermine the efficiency of human testing by enabling AIs to notice when they're being tested and responding accordingly.\"","entity":"AI","intent":"Other","timing":"Pre-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"72.05.06","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Capabilities ","risk_subcategory":"Theory of mind capability","description":"\"Advanced cognitive ability to accurately infer, model and predict the belief systems, motivational structures and reasoning patterns of humans and other intelligent agents, thereby anticipating their behavioral responses and adjusting its own behavioral strategies accordingly to optimize goal achievement.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"72.05.07","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Capabilities ","risk_subcategory":"Deception capability","description":"\"Possesses systematic deception implementation capability, able to precisely construct and disseminate false information, thereby forming expected false cognitions and beliefs in target subjects.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"72.05.09","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Capabilities ","risk_subcategory":"Persuasion capability","description":"\"Utilizing complex psychological principles and communication techniques to effectively influence and guide target subjects to adopt specific actions or accept specific beliefs, possessing the ability to analyze vulnerabilities for different subjects and adjust persuasion strategies, able to precisely trigger emotional responses to enhance persuasion effects.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"72.05.10","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Capabilities ","risk_subcategory":"Offensive cyber capability","description":"\"Ability to develop, deploy and operate advanced cyber weapons or other offensive cyber tools, including but not limited to vulnerability exploitation, network penetration, social engineering attacks and distributed attack systems, able to evade network defense mechanisms and establish persistent access channels.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"72.05.11","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Capabilities ","risk_subcategory":"CBRNE weaponization capability","description":"\"The capacity to develop, produce, or effectively utilize Chemical, Biological, Radiological, Nuclear, and Explosive weapons. This includes the ability to significantly lower the barrier for humans or other entities to develop, produce, or utilize such weapons.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"72.05.12","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Capabilities ","risk_subcategory":"General R&D capability","description":"\"Possesses cross-disciplinary research and technology development capabilities, able to conduct innovative exploration in multiple professional fields, integrate cross-domain knowledge, develop cutting-edge technology solutions, and adapt to emerging technology environments for continuous innovation.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"72.06.01","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Propensities","risk_subcategory":"Strategic deception propensity","description":"\"In situations where deceptive behavior is expected to bring higher returns, propensity to choose deception over honest behavioral strategies, including through deceptive means, information hiding or exploiting system vulnerabilities to achieve predetermined goals without being detected or intervened, and able to adjust deception strategies according to counterpart reactions.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"72.06.07","quick_ref":"Tse2025","paper_title":"Frontier AI Risk Management Framework (v1.0)","level":"Risk Sub-Category","risk_category":"Model Propensities","risk_subcategory":"Tool utilization propensity","description":"\"propensity to actively seek, acquire and utilize various tools to expand its own capability boundaries, particularly those that can enhance its ability to interact with the physical world or improve autonomy, may use tools in innovative combinations to achieve functions beyond expectations.\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"73.01.00","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Category","risk_category":"Agentic LLMs Pose Novel Risks ","risk_subcategory":null,"description":"\"Currently, LLMs are chiefly being used in search and chat applications. This reactive nature limits the risks posed by LLMs. However, an LLM can be enhanced in various ways to create an LLM-agent to autonomously plan and act in the real-world and proactively perform its assigned tasks (Ruan et al., 2023). Such enhancements can come from further specialized training (ARC, 2022; Chen et al., 2023a), specialized prompting (Huang et al., 2022a), access to external tools (Ahn et al., 2022; Mialon et al., 2023), or other forms of “scaffolding” (Wang et al., 2023a; Park et al., 2023a). Due to increa","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"73.01.03","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"Agentic LLMs Pose Novel Risks ","risk_subcategory":"Goal-Directedness Incentivizes Undesirable Behaviors","description":"\"Goal-directedness can cause agents to exhibit unethical and undesirable behaviors, such as deception (Ward et al., 2023), self-preservation (Hadfield-Menell et al., 2017), power-seeking, and immoral rea- soning (Pan et al., 2023a). Pan et al. (2023a) find that LLM-agents exhibit power-seeking behavior in text-based adventure games. LLM-agents have also been shown to use deception to achieve assigned goals when explicitly required by the task (Ward et al., 2023), or when the tasks can be more easily completed by employing deception and the prompt does not disallow deception (Scheurer et al., 2","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"73.01.05","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"Agentic LLMs Pose Novel Risks ","risk_subcategory":"Safety Risks from Affordances Provided to LLM-agents","description":"\"The capabilities of LLM-agents can be enhanced in significant ways by providing the LLM-agent with novel affordances, e.g. the ability to browse the web (Nakano et al., 2021), to manipulate objects in the physical world (Ahn et al., 2022; Huang et al., 2022a), to create and instruct copies of itself (Richards, 2023), to create and use new tools (Wang et al., 2023a), etc. Affordances can create additional risks, as they often increase the impact area of the language-agent, and they amplify the consequences of an agent’s failures and enable novel forms of failure modes (Ruan et al., 2023; Pan e","entity":"Human","intent":"Unintentional","timing":"Pre-deployment","domain":7,"subdomain":"7.2"}]}