{"attribution":{"source":"MIT AI Risk Repository, Domain Taxonomy of AI Risks v1 (MIT AI Risk Initiative)","license":"CC BY 4.0","license_url":"https://creativecommons.org/licenses/by/4.0/","citation":"Slattery, P., Saeri, A. K., Grundy, E. A. C., Graham, J., Noetel, M., Uuk, R., Dao, J., Pour, S., Casper, S., & Thompson, N. (2025). The AI Risk Repository: A comprehensive meta-review, database, and taxonomy of risks from artificial intelligence. arXiv:2408.12622."},"exported_at":"2026-09-11"}
{"rows":[{"ev_id":"73.01.00","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Category","risk_category":"Agentic LLMs Pose Novel Risks ","risk_subcategory":null,"description":"\"Currently, LLMs are chiefly being used in search and chat applications. This reactive nature limits the risks posed by LLMs. However, an LLM can be enhanced in various ways to create an LLM-agent to autonomously plan and act in the real-world and proactively perform its assigned tasks (Ruan et al., 2023). Such enhancements can come from further specialized training (ARC, 2022; Chen et al., 2023a), specialized prompting (Huang et al., 2022a), access to external tools (Ahn et al., 2022; Mialon et al., 2023), or other forms of “scaffolding” (Wang et al., 2023a; Park et al., 2023a). Due to increa","entity":"AI","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"73.01.02","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"Agentic LLMs Pose Novel Risks ","risk_subcategory":"Natural Language Underspecifies Goals ","description":"\"For LLM-agents, both the goal and environment observations are typically specified in the prompt through natural language. While natural language may provide a richer and more natural means of specifying goals than alternatives such as hand-engineering objective functions, natural language still suffers from underspecification (Grice, 1975; Piantadosi et al., 2012). Furthermore, in practice, users may neglect fully specifying their goals, especially the information pertaining to elements of the environment that ought not to be changed (the classic frame problem (Shanahan, 2016)). Such undersp","entity":"Other","intent":"Unintentional","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"73.01.03","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"Agentic LLMs Pose Novel Risks ","risk_subcategory":"Goal-Directedness Incentivizes Undesirable Behaviors","description":"\"Goal-directedness can cause agents to exhibit unethical and undesirable behaviors, such as deception (Ward et al., 2023), self-preservation (Hadfield-Menell et al., 2017), power-seeking, and immoral rea- soning (Pan et al., 2023a). Pan et al. (2023a) find that LLM-agents exhibit power-seeking behavior in text-based adventure games. LLM-agents have also been shown to use deception to achieve assigned goals when explicitly required by the task (Ward et al., 2023), or when the tasks can be more easily completed by employing deception and the prompt does not disallow deception (Scheurer et al., 2","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"73.01.05","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"Agentic LLMs Pose Novel Risks ","risk_subcategory":"Safety Risks from Affordances Provided to LLM-agents","description":"\"The capabilities of LLM-agents can be enhanced in significant ways by providing the LLM-agent with novel affordances, e.g. the ability to browse the web (Nakano et al., 2021), to manipulate objects in the physical world (Ahn et al., 2022; Huang et al., 2022a), to create and instruct copies of itself (Richards, 2023), to create and use new tools (Wang et al., 2023a), etc. Affordances can create additional risks, as they often increase the impact area of the language-agent, and they amplify the consequences of an agent’s failures and enable novel forms of failure modes (Ruan et al., 2023; Pan e","entity":"Human","intent":"Unintentional","timing":"Pre-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"73.02.00","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Category","risk_category":"Multi-Agent Safety Is Not Assured by Single-Agent Safety","risk_subcategory":null,"description":"\"A foremost lesson of game theory is that optimal decision-making within a single-agent setting (i.e. selfishly optimizing for an agent’s own utility) can produce sub-optimal outcomes in the presence of other strategic agents. Failing to account for the strategic nature of other agents can cause an agent to adopt strategies under which potentially everyone, including the agent itself, ends up worse off (Schelling, 1981; Harsanyi, 1995; Roughgarden, 2005; Nisan, 2007). Examples include collective action problems (or ‘social dilemmas’) such as arms races or the depletion of common resources, as ","entity":"Other","intent":"Other","timing":"Other","domain":7,"subdomain":"7.6"},{"ev_id":"73.02.01","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"Multi-Agent Safety Is Not Assured by Single-Agent Safety","risk_subcategory":"Foundationality May Cause Correlated Failures","description":"\"Another important characteristic of LLM development is foundationality — due to the expense of large- scale pretraining, many deployed instances share similar or identical learned components. Foundation- ality may both be a blessing and a curse. On the one hand, it may be possible to exploit the similarity in the design of LLM-agents to facilitate cooperation (Critch et al., 2022; Conitzer and Oesterheld, 2023; Oesterheld et al., 2023). On the other hand, foundationality may leave LLM-agents vulnerable to correlated failures both in terms of safety and capabilities due to increased output hom","entity":"Other","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"73.02.02","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"Multi-Agent Safety Is Not Assured by Single-Agent Safety","risk_subcategory":"Groups of LLM-Agents May Show Emergent Functionality","description":"\"Multi-agent learning, either through explicit finetuning or implicit in-context learning, may enable LLM-agents to influence each other during their interactions (Foerster et al., 2018). Under some environmental settings, this can create feedback loops that result in novel and emergent behaviors that would not manifest in the absence of multi-agent interactions (Hammond et al., 2024, Section 3.6).  Emergent functionality is a safety risk in two ways. Firstly, it may itself be dangerous (Shevlane et al., 2023). Secondly, it makes assurance harder as such emergent behaviors are difficult to pre","entity":"Other","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"73.02.03","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"Multi-Agent Safety Is Not Assured by Single-Agent Safety","risk_subcategory":"Collusion between LLM-Agents","description":"\"While it would often be preferable for LLM-agents to be cooperative, cooperation can be undesirable if it undermines pro-social competition or produces negative externalities for coalition non-members (Dorner, 2021; Buterin, 2019; Dafoe et al., 2020). Collusion between relatively simple AI systems has been observed in the real world (Assad et al., 2020; Wieting and Sapi, 2021) and synthetic experiments (Brown and MacKay, 2023; Calvano et al., 2020; Klein, 2021) Collusion can occur through explicit or steganographic communication. Steganographic communication hides information in seemingly inn","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"73.03.00","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Category","risk_category":"Dual-Use Capabilities Enable Malicious Use and Misuse of LLMs","risk_subcategory":null,"description":"\"Like all technologies, LLMs have the possibility for misuse by malicious actors. Malicious use of dual- use capabilities of AI is a recurring concern within literature (Brundage et al., 2018; Hendrycks et al., 2023; Mozes et al., 2023)\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":4,"subdomain":"4.0"},{"ev_id":"73.03.01","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"Dual-Use Capabilities Enable Malicious Use and Misuse of LLMs","risk_subcategory":"Misinformation and Manipulation","description":"\"Recent studies have demonstrated that LLMs can be exploited to craft deceptive narratives with levels of persuasiveness similar to human-generated content (Pan et al., 2023b; Spitale et al., 2023), to fabri- cate fake news (Zellers et al., 2019; Zhou et al., 2023f), and to devise automated influence operations aimed at manipulating the perspectives of targeted audiences (Goldstein et al., 2023). LLMs have also been found to be used in malicious social botnets (Yang and Menczer, 2023), powering automated accounts used to disseminate coordinated messages. More broadly, the use of LLMs for the d","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":4,"subdomain":"4.3"},{"ev_id":"73.03.02","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"Dual-Use Capabilities Enable Malicious Use and Misuse of LLMs","risk_subcategory":"Cybersecurity","description":"\"LLMs may exacerbate cybersecurity risks in various ways (Newman, 2024). Firstly, LLMs may significantly amplify the effectiveness of deceptive operations aimed at tricking people into disclosing sensitive information or granting adversary access to critical resources. For example, LLMs might prove highly effective at crafting personalized phishing emails or messages at scale that may be harder for an average user to recognize as phishing attempts (Karanjai, 2022; Hazell, 2023). In addition to being directly harmful to the targeted individual, such ‘social engineering’ attacks are often the ba","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":4,"subdomain":"4.3"},{"ev_id":"73.03.02a","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Additional evidence","risk_category":"Dual-Use Capabilities Enable Malicious Use and Misuse of LLMs","risk_subcategory":"Cybersecurity","description":null,"entity":null,"intent":null,"timing":null,"domain":null,"subdomain":null},{"ev_id":"73.03.02b","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Additional evidence","risk_category":"Dual-Use Capabilities Enable Malicious Use and Misuse of LLMs","risk_subcategory":"Cybersecurity","description":null,"entity":null,"intent":null,"timing":null,"domain":null,"subdomain":null},{"ev_id":"73.03.02c","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Additional evidence","risk_category":"Dual-Use Capabilities Enable Malicious Use and Misuse of LLMs","risk_subcategory":"Cybersecurity","description":null,"entity":null,"intent":null,"timing":null,"domain":null,"subdomain":null},{"ev_id":"73.03.03","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"Dual-Use Capabilities Enable Malicious Use and Misuse of LLMs","risk_subcategory":"Surveillance and Censorship","description":"\"Content moderation has emerged as one of the key use-cases of LLMs (Weng et al., 2023), indicating the potential of LLMs for surveillance and censorship as well (Edwards, 2023). Surveillance and censorship are one of the primary tools employed by governments with dictatorial tendencies to suppress opposing political and social voices. These censorship measures, however, are often quite crude and can be escaped with little ingenuity...However, LLMs could enable significantly more sophisticated surveillance and censorship operations at scale (Feldstein, 2019). Multimodal-LLMs or LLMs combined w","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":4,"subdomain":"4.1"},{"ev_id":"73.03.04","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"Dual-Use Capabilities Enable Malicious Use and Misuse of LLMs","risk_subcategory":"Warfare and Physical Harm","description":"\"The use of AI in warfare is highly alarming and may pose dangers to human safety (Hendrycks et al., 2023). Autonomous drone warfare is being aggressively pursued as a tactic in the current war in Ukraine (Meaker, 2023), and may already have been used on human targets (Hambling, 2023). The use of AI- based facial recognition has been documented in the targeting of Palestinians in Gaza (International, 2023). LLMs have already been productized in limited ways for the purposes of warfare planning (Tarantola, 2023). Furthermore, active research is being carried out to develop multimodal-LLMs that ","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":4,"subdomain":"4.2"},{"ev_id":"73.03.05","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"Dual-Use Capabilities Enable Malicious Use and Misuse of LLMs","risk_subcategory":"Hazardous Biological and Chemical Technologies","description":"\"AI systems such as LLMs, chemical LLMs (Skinnider et al., 2021; Moret et al., 2023), and other LLM- based biological design tools might soon facilitate the production of bioweapons, chemical weapons, and other hazardous technologies. In particular, LLMs might enable actors with less expertise to more easily synthesize dangerous pathogens, while customized chemical and biological design tools might be more concerning in terms of expanding the capabilities of sophisticated actors (e.g. states) (Sandbrink, 2023). Gopal et al. (2023) and Soice et al. (2023) demonstrated that people with little ba","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":4,"subdomain":"4.2"},{"ev_id":"73.03.05a","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Additional evidence","risk_category":"Dual-Use Capabilities Enable Malicious Use and Misuse of LLMs","risk_subcategory":"Hazardous Biological and Chemical Technologies","description":null,"entity":null,"intent":null,"timing":null,"domain":null,"subdomain":null},{"ev_id":"73.03.06","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"Dual-Use Capabilities Enable Malicious Use and Misuse of LLMs","risk_subcategory":"Domain-Specific Misuses","description":"\"Improvements in LLMs may exert greater pressure to apply LLMs to various domains, such as health and education (Eloundou et al., 2023). Crude efforts to use LLMs in such domains, however, may incur harm and should be discouraged strongly. In particular, it is important to guard against different ways in which LLMs may be misused within any domain. One famous episode of misuse within the health sector is a mental health non-profit experimenting LLM-based therapy on its users without their informed consent (Xiang, 2023a). Within the education sector, LLMs may be misused in various ways that mig","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":4,"subdomain":"4.3"},{"ev_id":"73.04.00","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Category","risk_category":"LLM-Systems Can Be Untrustworthy","risk_subcategory":null,"description":"\"A key desideratum for an LLM from a user’s perspective is ‘trustworthiness’, i.e. assurance of reliability and consistent performance, and absence of any accidental harm caused by the technology to the user.16 Providing assurance that an LLM-based system will not cause accidental harm remains a major open challenge. Harms may either occur directly due to the flawed nature of LLMs, e.g. an LLM generating toxic language or behaving inappropriately in some other ways, or may occur due to improper usage by a user, e.g. automation bias due to a user’s overreliance on LLM.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":null,"subdomain":null},{"ev_id":"73.04.01","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"LLM-Systems Can Be Untrustworthy","risk_subcategory":"Harms of Representation and Other Biases","description":"\"A pretrained LLM generally has many of the stereotypical biases commonly present in the human society (Touvron et al., 2023). This makes it difficult for users to trust that LLMs will work well for them and not produce unfair or biased responses. Appropriate finetuning can effectively limit the bias displayed in LLM outputs in a variety of situations, e.g. when models are explicitly prompted with stereotypes (Wang et al., 2023k), but it does not ‘solve’ the problem. Even after finetuning, biases often resurface when deliberately elicited (Wang et al., 2023k), or under novel scenarios, e.g. in","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":1,"subdomain":"1.1"},{"ev_id":"73.04.01a","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Additional evidence","risk_category":"LLM-Systems Can Be Untrustworthy","risk_subcategory":"Harms of Representation and Other Biases","description":null,"entity":null,"intent":null,"timing":null,"domain":null,"subdomain":null},{"ev_id":"73.04.02","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"LLM-Systems Can Be Untrustworthy","risk_subcategory":"Inconsistent Performance across and within Domains","description":"\"Estimating true capabilities of an LLM is a difficult task (c.f. Section 3.3), especially for naive users unfamiliar with the brittle nature of machine learning technologies. Exaggeration of model capabilities by the developers (Lambert, 2023; Blair-Stanek et al., 2023), and issues such as task-contamination (Roberts et al., 2023b), underrepresentation of tasks or domains (Wu et al., 2023a; McCoy et al., 2023), and prompt-sensitivity (Anthropic, 2023d) may cause a user to misestimate the true capabilities of a model. This lack of reliability can undermine user trust or cause harm if a user ba","entity":"Human","intent":"Unintentional","timing":"Post-deployment","domain":5,"subdomain":"5.1"},{"ev_id":"73.04.02a","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Additional evidence","risk_category":"LLM-Systems Can Be Untrustworthy","risk_subcategory":"Inconsistent Performance across and within Domains","description":null,"entity":null,"intent":null,"timing":null,"domain":null,"subdomain":null},{"ev_id":"73.04.03","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"LLM-Systems Can Be Untrustworthy","risk_subcategory":"Overreliance","description":"\"If a user begins to excessively trust an LLM, this may cause them to develop an overreliance on the LLM. Overreliance can result in automation bias (Kupfer et al., 2023), and can cause errors of omission (user choosing not to verify the validity of a response) and errors of commission (user believing and acting on the basis of the LLM’s response, even if it contradicts their own knowledge) (Skitka et al., 1999). It can be particularly dangerous in domains where the user may lack relevant expertise to robustly scrutinize the LLM responses. This is particularly a source of risk for LLMs because","entity":"Human","intent":"Unintentional","timing":"Post-deployment","domain":5,"subdomain":"5.1"},{"ev_id":"73.04.03a","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Additional evidence","risk_category":"LLM-Systems Can Be Untrustworthy","risk_subcategory":"Overreliance","description":null,"entity":null,"intent":null,"timing":null,"domain":null,"subdomain":null},{"ev_id":"73.05.00","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Category","risk_category":"Socioeconomic Impacts of LLM May Be Highly Disruptive","risk_subcategory":null,"description":"\"The rapid evolution of LLMs brings significant socioeconomic opportunities and challenges, impacting the workforce, income inequality, education, and global economic development. Many of these challenges are systemic in nature, constituting what economists refer to as general equilibrium effects. These challenges do not arise directly from LLMs causing harm to users but rather from their indirect effects on the socioeconomic equilibrium.\"","entity":"Other","intent":"Unintentional","timing":"Post-deployment","domain":null,"subdomain":null},{"ev_id":"73.05.00a","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Additional evidence","risk_category":"Socioeconomic Impacts of LLM May Be Highly Disruptive","risk_subcategory":null,"description":null,"entity":null,"intent":null,"timing":null,"domain":null,"subdomain":null},{"ev_id":"73.05.01","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"Socioeconomic Impacts of LLM May Be Highly Disruptive","risk_subcategory":"Effects on the Workforce","description":"\"Rapid advances in LLMs pose three distinct sets of challenges for workers’ incomes (Korinek and Stiglitz, 2019; Susskind, 2023). First, they are likely to accelerate the rate of job turnover and disruption —– affecting more workers, including more highly skilled workers, and making the adjustment process for society more difficult than what we were used to from prior technological advances...Second, although technological progress means that society may produce more wealth overall, there is a risk that the general-purpose nature of LLMs may lead to progress that is biased against labor, meani","entity":"Other","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.2"},{"ev_id":"73.05.01a","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Additional evidence","risk_category":"Socioeconomic Impacts of LLM May Be Highly Disruptive","risk_subcategory":"Effects on the Workforce","description":null,"entity":null,"intent":null,"timing":null,"domain":null,"subdomain":null},{"ev_id":"73.05.01b","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Additional evidence","risk_category":"Socioeconomic Impacts of LLM May Be Highly Disruptive","risk_subcategory":"Effects on the Workforce","description":null,"entity":null,"intent":null,"timing":null,"domain":null,"subdomain":null},{"ev_id":"73.05.02","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"Socioeconomic Impacts of LLM May Be Highly Disruptive","risk_subcategory":"Effects on Inequality","description":"\"LLMs could potentially worsen socioeconomic inequalities (Capraro et al., 2023). Effects on inequal- ity are closely linked to the effects of LLMs on workers but ultimately depend on how the fruits of technological progress are distributed...First, if the role and compensation of capital rise and the role and compensation of labor decline in an LLM-powered economy, inequality may go up because work is the main source of income for the majority of people...Second, the large fixed cost of training cutting-edge LLMs and the network effects involved imply that the market for the most advanced LLM","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.2"},{"ev_id":"73.05.03","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"Socioeconomic Impacts of LLM May Be Highly Disruptive","risk_subcategory":"Global Economic Development","description":"\"Many of the themes and challenges that we discussed above come together when analyzing the socioeconomic effects on developing countries. The workforce of developing countries may suffer from a retrenchment of outsourcing as many simple cognitive tasks that used to be performed in developing countries — for example, in call centers –— can be automated with LLMs. This may adversely affect the economies of the poor countries (Georgieva, 2024).\"","entity":"Other","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.2"},{"ev_id":"73.05.03a","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Additional evidence","risk_category":"Socioeconomic Impacts of LLM May Be Highly Disruptive","risk_subcategory":"Global Economic Development","description":null,"entity":null,"intent":null,"timing":null,"domain":null,"subdomain":null},{"ev_id":"73.06.00","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Category","risk_category":"Corporate power may impeded effective governance ","risk_subcategory":null,"description":"\"The increasing power and influence of large corporations may make effective governance difficult. There exists a power asymmetry between corporate entities profiting from LLMs and other social groups (e.g. civil society). State-of-the-art LLMs are developed by or in partnership with, some of the world’s largest private tech companies...This poses a risk of governance protocols related to LLMs becoming excessively favorable to tech companies, potentially leading to regulatory capture at the cost of the interests of other societal groups, particularly marginalized communities who have historica","entity":"Other","intent":"Unintentional","timing":"Other","domain":6,"subdomain":"6.1"},{"ev_id":"73.07.00","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Category","risk_category":"Jailbreaks and Prompt Injections Threaten Security of LLMs","risk_subcategory":null,"description":"\"LLMs are not adversarially robust and are vulnerable to security failures such as jailbreaks and prompt-injection attacks. While a number of jailbreak attacks have been proposed in the literature, the lack of standardized evaluation makes it difficult to compare them. We also do not have efficient white-box methods to evaluate adver- sarial robustness. Multi-modal LLMs may further allow novel types of jailbreaks via additional modalities. Finally, the lack of robust privilege levels within the LLM input means that jailbreaking and prompt-injection attacks may be particularly hard to eliminate","entity":"Other","intent":"Other","timing":"Other","domain":2,"subdomain":"2.2"},{"ev_id":"73.07.01","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"Jailbreaks and Prompt Injections Threaten Security of LLMs","risk_subcategory":"Exploiting Limited Generalization of Safety Finetuning","description":"\"Safety tuning is performed over a much narrower distribution compared to the pretraining distribution. This leaves the model vulnerable to attacks that exploit gaps in the generalization of the safety training, e.g. using encoded text (Wei et al., 2023c) or low-resource languages (Deng et al., 2023a; Yong et al., 2023) (see also Section 3.2).\"","entity":"Other","intent":"Unintentional","timing":"Other","domain":2,"subdomain":"2.2"},{"ev_id":"73.07.02","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"Jailbreaks and Prompt Injections Threaten Security of LLMs","risk_subcategory":"“Model Psychology” Attacks","description":"\"LLMs are vulnerable to “psychological” tricks (Li et al., 2023e; Shen et al., 2023), which can be exploited by attackers. Examples include instructing the model to behave like a specific persona (Shah et al., 2023; Andreas, 2022), or employing various “social engineering” tricks crafted by humans (Wei et al., 2023c) or other LLMs (Perez et al., 2022b; Casper et al., 2023c).\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"73.07.03","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"Jailbreaks and Prompt Injections Threaten Security of LLMs","risk_subcategory":"Adversarial Optimization:","description":"\"Jailbreak attacks can be discovered by performing manual or auto- mated adversarial optimization against a proxy objective that is noisily correlated with the success of a jailbreak. These are mostly gradient-based attacks (Zou et al., 2023b; Shin et al., 2020) as described in the previous two challenges, but gradient-free methods also exist (Prasad et al., 2022; Deng et al., 2022; Lapid et al., 2023).\"","entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":null,"subdomain":null},{"ev_id":"73.07.04","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"Jailbreaks and Prompt Injections Threaten Security of LLMs","risk_subcategory":"Attacking LLMs via Additional Modalities a","description":"\"LLMs can now process modalities other than text, e.g. images or video frames (OpenAI, 2023c; Gemini Team, 2023). Several studies show that gradient-based attacks on multimodal models are easy and effective (Carlini et al., 2023a; Bailey et al., 2023; Qi et al., 2023b). These attacks manipulate images that are input to the model (via an appropriate encoding). GPT-4Vision (OpenAI, 2023c) is vulnerable to jailbreaks and exfiltration attacks through much simpler means as well, e.g. writing jailbreaking text in the image (Willison, 2023a; Gong et al., 2023). For indirect prompt injection, the atta","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.2"},{"ev_id":"73.08.00","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Category","risk_category":"Vulnerability to Poisoning and Backdoors","risk_subcategory":null,"description":"\"The previous section explored jailbreaks and other forms of adversarial prompts as ways to elicit harmful capabilities acquired during pretraining. These methods make no assumptions about the training data. On the other hand, poisoning attacks (Biggio et al., 2012) perturb training data to introduce specific vulnerabilities, called backdoors, that can then be exploited at inference time by the adversary. This is a challenging problem in current large language models because they are trained on data gathered from untrusted sources (e.g. internet), which can easily be poisoned by an adversary (","entity":"Human","intent":"Intentional","timing":"Pre-deployment","domain":2,"subdomain":"2.2"}]}