MIT AI Risk Repository
Browse AI risks
336 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.
-
61.02.42 · Risk Sub-Category
Sources of systemic risks from general-purpose AI
Risks from network interconnectivity
"The interconnectedness of AI networks can create vulnerabilities, where issues in one part of the network can have cascading effects across the system."
-
62.19.06 · Risk Sub-Category
Attacks on GPAIs/GPAI Failure Modes
Vulnerabilities arising from additional modalities in multimodal models
"Additional modalities can introduce new attack vectors in multimodal models as well as expand the scope of the previous attacks, ranging from jailbreaking to poisoning [13]. Typically, different modalities have different robustness levels, allowing malicious actors to choose the most vulnerable part of the model to attack [119, 181]."
-
62.19.07 · Risk Sub-Category
Attacks on GPAIs/GPAI Failure Modes
Vulnerabilities to jailbreaks exploiting long context windows (many- shot jailbreaking)
"Language models with long context windows are vulnerable to new types of ex- ploitations that are ineffective on models with shorter context windows. While few-shot jailbreaking, which involves providing few examples of the desired harmful output, might not trigger a harmful response, many-shot jailbreak- ing, which involves a higher number of such examples, increases the likelihood of eliciting an undesirable output. These vulnerabilities become more significant as context windows expand with newer model releases [7]."
-
73.07.00 · Risk Category
"LLMs are not adversarially robust and are vulnerable to security failures such as jailbreaks and prompt-injection attacks. While a number of jailbreak attacks have been proposed in the literature, the lack of standardized evaluation makes it difficult to compare them. We also do not have efficient white-box methods to evaluate adver- sarial robustness. Multi-modal LLMs may further allow novel types of jailbreaks via additional modalities. Finally, the lack of robust privilege levels within the LLM input means that jailbreaking and prompt-injection attacks may be particularly hard to eliminate
-
73.07.01 · Risk Sub-Category
Jailbreaks and Prompt Injections Threaten Security of LLMs
Exploiting Limited Generalization of Safety Finetuning
"Safety tuning is performed over a much narrower distribution compared to the pretraining distribution. This leaves the model vulnerable to attacks that exploit gaps in the generalization of the safety training, e.g. using encoded text (Wei et al., 2023c) or low-resource languages (Deng et al., 2023a; Yong et al., 2023) (see also Section 3.2)."
-
24.11.00 · Risk Category
"The rapid integration of AI systems with advanced capabilities, such as greater autonomy, content generation, memorisation and planning skills (see Chapter 4) into personalised assistants also raises new and more specific challenges related to misinformation, disinformation and the broader integrity of our information environment. "
-
"Historical revisionism - Deliberate or unintentional reinterpretation of established/orthodox historical events or accounts held by societies, communities, academics."
-
"LLMs have been demonstrated to pursue consistent context [129]–[132], which may lead to erroneous generation when the prefixes contain false information. Typical examples include sycophancy [129], [130], false demonstrations-induced hallucinations [113], [133], and snowballing [131]. As LLMs are generally fine-tuned with instruction-following data and user feedback, they tend to reiterate user-provided opinions [129], [130], even though the opinions contain misinformation. Such a sycophantic behavior amplifies the likelihood of generating hallucinations, since the model may prioritize user op
-
information-based harms capture concerns of misinformation, disinformation, and malinformation. Algorithmic systems, especially generative models and recommender, systems can lead to these information harms
-
45.02.02 · Risk Sub-Category
Safety risks in AI Applications
Cyberspace risks (Risks of confusing facts, misleading users, and bypassing authentication)
"AI systems and their outputs, if not clearly labeled, can make it difficult for users to discern whether they are interacting with AI and to identify the source of generated content. This can impede users' ability to determine the authenticity of information, leading to misjudgment and misunderstanding. Additionally, AI-generated highly realistic images, audio, and videos may circumvent existing identity verification mechanisms, such as facial recognition and voice recognition, rendering these authentication processes ineffective."
-
"disseminating false or misleading information about people"
-
—
-
31.01.05 · Risk Sub-Category
Clickbait and feeding the surveillance advertising ecosystem
"Beyond misinformation and disinformation, generative AI can be used to create clickbait headlines and articles, which manipulate how users navigate the internet and applications. For example, generative AI is being used to create full articles, regardless of their veracity, grammar, or lack of common sense, to drive search engine optimization and create more webpages that users will click on. These mechanisms attempt to maximize clicks and engagement at the truth’s expense, degrading users’ experiences in the process. Generative AI continues to feed this harmful cycle by spreading misinformat
-
55.04.05 · Risk Sub-Category
Worsened epistemic processes for society
Reduced decision-making capacity as a result of decreased trust in information
"In addition, the increased awareness of these trends in information production and distribution could make it harder for anyone to evaluate the trustworthiness of any information source, reducing overall trust in information. In all of these scenarios, it would be much harder for humanity to make good decisions on important issues, particularly due to declining trust in credible multipartisan sources, which could hamper attempts at cooperation and collective action. The vaccine and mask hesitancy that exacerbated Covid-19, for example, were likely the result of insufficient trust in public he
-
"Radicalisation - Adoption of extreme political, social, or religious ideals and aspirations due to the nature or misuse of an algorithmic system, potentially resulting in abuse, violence, or terrorism."
-
"Institutional trust loss - Erosion of trust in public institutions and weakened checks and balances due to mis/disinformation, influence operations, over-dependence on technology, etc."
-
61.02.20 · Risk Sub-Category
Sources of systemic risks from general-purpose AI
Detection challenges in content
"The difficulty in distinguishing synthetic content from authentic material adds to information risks."
-
"Contaminating publicly available information with false or inaccurate information (i.e., the generative tool's output is disseminated beyond the end user)"
-
"Eroding trust in public information and knowledge"
-
"Pollution of a space/ecosystem that is expected to be free of AI involvement/influence (e.g., creative material submission portals, job applications)"
-
"Frontier AI can cheaply generate realistic content which can falsely portray people and events. There is potential risk of compromised decision-making by individuals and institutions who rely on inaccurate or misleading publicly available information, as well as lower overall trust in true information."
-
03.05.00 · Risk Category
"Some uses of AI have been deeply concerning, namely voice cloning [58] and the generation of deep fake videos [59]. For example, in March 2022, in the early days of the Russian invasion of Ukraine, hackers broadcast via the Ukrainian news website Ukraine 24 a deep fake video of President Volodymyr Zelensky capitulating and calling on his soldiers to lay down their weapons [60]. The necessary software to create these fakes is readily available on the Internet, and the hardware requirements are modest by today’s standards [61]. Other nefarious uses of AI include accelerating password cracking [
-
04.07.00 · Risk Category
LMs, due to their remarkable capabilities, carry the same potential for malice as other technological products. For instance, they may be used in information warfare to generate deceptive information or unlawful content, thereby having a significant impact on individuals and society. As current LMs are increasingly built as agents to accomplish user objectives, they may disregard the moral and safety guidelines if operating without adequate supervision. Instead, they may execute user commands mechanically without considering the potential damage. They might interact unpredictably with humans a
-
"AI’s potential for both beneficial and harmful applications complicates efforts to manage its societal impacts effectively."
-
"Benign intermediate for harmful end objective"
-
Political harms emerge when “people are disenfranchised and deprived of appropriate political power and influence” [186, p. 162]. These harms focus on the domain of government, and focus on how algorithmic systems govern through individualized nudges or micro-directives [187], that may destabilize governance systems, erode human rights, be used as weapons of war [188], and enact surveillant regimes that disproportionately target and harm people of color
-
19.02.00 · Risk Category
"Informational and communicational AI risks refer particularly to informational manipulation through AI systems that influence the provision of information (Rahwan, 2018; Wirtz & Müller, 2019), AIbased disinformation and computational propaganda, as well as targeted censorship through AI systems that use respectively modified algorithms, and thus restrict freedom of speech."
-
19.02.01 · Risk Sub-Category
Informational and Communicational AI Risks
Manipulation and control of information provision (e.g., personalised adds, filtered news)
—
-
24.03.14 · Risk Sub-Category
Authoritarian Surveillance, Censorship, and Use: Delegation of Decision-Making Authority to Malicious Actors
"Finally, the principal value proposition of AI assistants is that they can either enhance or automate decision-making capabilities of people in society, thus lowering the cost and increasing the accuracy of decision-making for its user. However, benefiting from this enhancement necessarily means delegating some degree of agency away from a human and towards an automated decision-making system—motivating research fields such as value alignment. This introduces a whole new form of malicious use which does not break the tripwire of what one might call an ‘attack’ (social engineering, cyber offen
-
-
-
46.03.00 · Risk Category
"The distortion of the information ecosystem, including the spread of misinformation, fake news, and other forms of deceptive content [28], is categorized as “Information Manipulation.”"
-
-
-
46.04.00 · Risk Category
"Lastly, broader harms that can impact communities, societal structures, and critical infrastructures, including threats to democratic processes, social cohesion, and technological systems, are captured under “Societal, Socio-technical, and Infrastructural Damage.”"
-
-
-
-
-
—
-
—
-
58.08.00 · Risk Category
"Political and Economic - Manipulation of political beliefs, damage to political institutions and the effective delivery of government services."
-
"Large-scale influence on communication and information systems, and epistemic processes more generally."
-
"This group comprises 11% of the articles and centers on risks stemming from AI systems designed with malicious intent or that can end up in a threat to human life. It can be divided into two key themes: threats to law and democracy, and transhumanism."
-
48.01.00 · Risk Category
"Eased access to or synthesis of materially nefarious information or design capabilities related to chemical, biological, radiological, or nuclear (CBRN) weapons or other dangerous materials or agents."
-
—
-
50.04.08 · Risk Sub-Category
Legal and Rights-Related Risks
Criminal Activities (Other Unlawful/Criminal Activities)
—
-
53.03.06 · Risk Sub-Category
Failures in or misuse of intermediary (non-AGI) AI systems, resulting in catastrophe
"Deployment of “prepotent” AI systems that are non-general but capable of outperforming human collective efforts on various key dimensions;170 → Militarization of AI enabling mass attacks using swarms of lethal autonomous weapons systems;171 → Military use of AI leading to (intentional or unintentional) nuclear escalation, either because machine learning systems are directly integrated in nuclear command and control systems in ways that result in escalation172 or because conventional AI-enabled systems (e.g., autonomous ships) are deployed in ways that result in provocation and escalation;173
-
"Business operations/infrastructure damage - Damage, disruption, or destruction of a business system and/or its components due to malfunction, cyberattacks, etc."
-
68.02.00 · Risk Category
"Cyber risks, especially in the context of cyber offense, are an existing threat that may be exacerbated by AI. [108] demonstrated that teams of LLM agents can exploit zero-day vulnerabilities when given a description of the vulnerability and toy capture-the-flag problems. While cyber risks are not typically regarded as catastrophic, [3] argues that cyberwarfare is an underappreciated risk that poses a credible threat of catastrophic harm."
-
"Biological risks encompass the dangerous modification of pathogens and unethical manipulation of genetic material, potentially leading to unforeseen biohazardous outcomes."
-
"AI-generated impersonation for identity theft might be found at the intersection of “Harm to the Person” and “Deception.”"
-
46.02.00 · Risk Category
"Then, we have the potential for financial loss, fraud, market manipulation, and other economic harms, which fall under “Financial and Economic Damage.”
-
-
Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.