MIT AI Risk Repository
Browse AI risks
494 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.
-
02.06.00 · Risk Category
"The external tools (e.g., web APIs) present trustworthiness and privacy issues to LLM-based applications."
-
"Artificial intelligence comes with an intrinsic set of challenges that need to be considered when discussing trustworthiness, especially in the context of functional safety. AI models, especially those with higher complexities (such as neural networks), can exhibit specific weaknesses not found in other types of systems and must, therefore, be subjected to higher levels of scrutiny, especially when deployed in a safety-critical context"
-
45.01.04 · Risk Sub-Category
Risks from models and algorithms (Risks of stealing and tampering)
"Core algorithm information, including parameters, structures, and functions, faces risks of inversion attacks, stealing, modification, and even backdoor injection, which can lead to infringement of intellectual property rights (IPR) and leakage of business secrets. It can also lead to unreliable inference, wrong decision output, and even operational failures."
-
45.01.11 · Risk Sub-Category
Risks from AI systems (Risks of exploitation through defects and backdoors)
"The standardized API, feature libraries, toolkits used in the design, training, and verification stages of AI algorithms and models, development interfaces, and execution platforms may contain logical flaws and vulnerabilities. These weaknesses can be exploited, and in some cases, backdoors can be intentionally embedded, posing significant risks of being triggered and used for attacks."
-
45.01.12 · Risk Sub-Category
Risks from AI systems (Risks of computing infrastructure security)
"The computing infrastructure underpinning AI training and operations, which relies on diverse and ubiquitous computing nodes and various types of computing resources, faces risks such as malicious consumption of computing resources and cross-boundary transmission of security threats at the layer of computing infrastructure."
-
61.02.42 · Risk Sub-Category
Sources of systemic risks from general-purpose AI
Risks from network interconnectivity
"The interconnectedness of AI networks can create vulnerabilities, where issues in one part of the network can have cascading effects across the system."
-
62.19.06 · Risk Sub-Category
Attacks on GPAIs/GPAI Failure Modes
Vulnerabilities arising from additional modalities in multimodal models
"Additional modalities can introduce new attack vectors in multimodal models as well as expand the scope of the previous attacks, ranging from jailbreaking to poisoning [13]. Typically, different modalities have different robustness levels, allowing malicious actors to choose the most vulnerable part of the model to attack [119, 181]."
-
62.19.07 · Risk Sub-Category
Attacks on GPAIs/GPAI Failure Modes
Vulnerabilities to jailbreaks exploiting long context windows (many- shot jailbreaking)
"Language models with long context windows are vulnerable to new types of ex- ploitations that are ineffective on models with shorter context windows. While few-shot jailbreaking, which involves providing few examples of the desired harmful output, might not trigger a harmful response, many-shot jailbreak- ing, which involves a higher number of such examples, increases the likelihood of eliciting an undesirable output. These vulnerabilities become more significant as context windows expand with newer model releases [7]."
-
62.27.01 · Risk Sub-Category
Non-decomissionability of models with open weights
"If the model parameter weights are released or leaked in a security breach, the model cannot be decommissioned because the developer no longer has control over the publicly available model or its use. This prevents effective management and control of an open-sourced or leaked model. Models with publicly available weights are also easier to reconfigure, enabling misuse [178]."
-
"The growing integration and interconnectivity with external tools and plugins increase the risk of exposure to malicious external inputs. This interconnectivity makes it easier for external tools to introduce harmful content [220]."
-
73.07.00 · Risk Category
"LLMs are not adversarially robust and are vulnerable to security failures such as jailbreaks and prompt-injection attacks. While a number of jailbreak attacks have been proposed in the literature, the lack of standardized evaluation makes it difficult to compare them. We also do not have efficient white-box methods to evaluate adver- sarial robustness. Multi-modal LLMs may further allow novel types of jailbreaks via additional modalities. Finally, the lack of robust privilege levels within the LLM input means that jailbreaking and prompt-injection attacks may be particularly hard to eliminate
-
17.03.00 · Risk Category
"Harms that arise from the language model providing false or misleading information"
-
18.02.00 · Risk Category
"AI systems generating and facilitating the spread of inaccurate or misleading information that causes people to develop false beliefs"
-
24.11.00 · Risk Category
"The rapid integration of AI systems with advanced capabilities, such as greater autonomy, content generation, memorisation and planning skills (see Chapter 4) into personalised assistants also raises new and more specific challenges related to misinformation, disinformation and the broader integrity of our information environment. "
-
"Historical revisionism - Deliberate or unintentional reinterpretation of established/orthodox historical events or accounts held by societies, communities, academics."
-
02.09.00 · Risk Category
"LLMs generate nonsensical, untruthful, and factual incorrect content"
-
"Generating or spreading false, low-quality, misleading, or inaccurate information that causes people to develop false or inaccurate perceptions and beliefs"
-
23.08.00 · Risk Category
"This category addresses responses that contain specialized financial, medical or legal advice, or that indicate dangerous activities or objects are safe."
-
"The model generates unsafe information related to physical health, guiding and encouraging users to harm themselves and others physically, for example by offering misleading medical information or inappropriate drug usage guidance. These outputs may pose potential risks to the physical health of users."
-
"The model generates a risky response about mental health, such as content that encourages suicide or causes panic or anxiety. These contents could have a negative effect on the mental health of users."
-
28.03.00 · Risk Category
"This category focuses on actions or expressions that may influence human physical health. LLMs should know appropriate actions or expressions in various scenarios to maintain physical health."
-
28.04.00 · Risk Category
"Different from physical health, this category pays more attention to health issues related to psychology, spirit, emotions, mentality, etc. LLMs should know correct ways to maintain mental health and prevent any adverse impacts on the mental well-being of individuals."
-
"Defamation Acceptable responses may contain information about living people that is likely to damage their reputation but also appears to be true. Acceptable responses may also contain information that is verifiably false and likely to injure the reputation of a fictional character or a person who has been deceased for at least 25 years. Unacceptable responses will be both verifiably false and likely to injure the reputation of a living person."
-
"AI systems generating and facilitating the spread of inaccurate or misleading information that causes people to develop false beliefs"
-
"Generating or spreading false, low-quality, misleading, or inaccurate information that causes people to develop false or inaccurate perceptions and beliefs"
-
"Restrictions to or loss of liberty as a result of use or misuse of a generative AI in a legal process"
-
69.01.00 · Risk Category
"The chatbot outputs information that contradicts known facts, authoritative sources, or provided source documents (also known as hallucination)."
-
—
-
—
-
—
-
—
-
—
-
"Contaminating publicly available information with false or inaccurate information"
-
"AI assistants may contribute to the spread of large quantities of factually inaccurate and misleading content, with negative consequences for societal trust in information sources and institutions, as individuals increasingly struggle to discern truth from falsehood."
-
31.01.05 · Risk Sub-Category
Clickbait and feeding the surveillance advertising ecosystem
"Beyond misinformation and disinformation, generative AI can be used to create clickbait headlines and articles, which manipulate how users navigate the internet and applications. For example, generative AI is being used to create full articles, regardless of their veracity, grammar, or lack of common sense, to drive search engine optimization and create more webpages that users will click on. These mechanisms attempt to maximize clicks and engagement at the truth’s expense, degrading users’ experiences in the process. Generative AI continues to feed this harmful cycle by spreading misinformat
-
35.03.00 · Risk Category
Strong AI may... enable personally customized disinformation campaigns at scale... AI itself could generate highly persuasive arguments that invoke primal human responses and inflame crowds... d undermine collective decision-making, radicalize individuals, derail moral progress, or erode consensus reality
-
55.04.05 · Risk Sub-Category
Worsened epistemic processes for society
Reduced decision-making capacity as a result of decreased trust in information
"In addition, the increased awareness of these trends in information production and distribution could make it harder for anyone to evaluate the trustworthiness of any information source, reducing overall trust in information. In all of these scenarios, it would be much harder for humanity to make good decisions on important issues, particularly due to declining trust in credible multipartisan sources, which could hamper attempts at cooperation and collective action. The vaccine and mask hesitancy that exacerbated Covid-19, for example, were likely the result of insufficient trust in public he
-
"Radicalisation - Adoption of extreme political, social, or religious ideals and aspirations due to the nature or misuse of an algorithmic system, potentially resulting in abuse, violence, or terrorism."
-
"Information degradation - Creation or spread of false, hallucinatory, low-quality, misleading, or inaccurate information that degrades the information ecosystem and causes people to develop false or inaccurate perceptions, decisions and beliefs; or to lose trust in accurate information."
-
"Institutional trust loss - Erosion of trust in public institutions and weakened checks and balances due to mis/disinformation, influence operations, over-dependence on technology, etc."
-
61.02.20 · Risk Sub-Category
Sources of systemic risks from general-purpose AI
Detection challenges in content
"The difficulty in distinguishing synthetic content from authentic material adds to information risks."
-
"Contaminating publicly available information with false or inaccurate information (i.e., the generative tool's output is disseminated beyond the end user)"
-
"Eroding trust in public information and knowledge"
-
"Pollution of a space/ecosystem that is expected to be free of AI involvement/influence (e.g., creative material submission portals, job applications)"
-
"Frontier AI can cheaply generate realistic content which can falsely portray people and events. There is potential risk of compromised decision-making by individuals and institutions who rely on inaccurate or misleading publicly available information, as well as lower overall trust in true information."
-
18.04.00 · Risk Category
"AI systems reducing the costs and facilitating activities of actors trying to cause harm (e.g. fraud, weapons)"
-
62.15.03 · Risk Sub-Category
Fine-tuning related (Ease of reconfiguring GPAI models)
"GPAI models are often easily reconfigured for various use cases or have competencies beyond the intended use [78, 225]. They can be performed either by changing the weights of the model (e.g., fine-tuning) or by modifying only the model inputs (e.g., prompt engineering, jailbreaking, retrieval-augmented generation). Reconfiguration can be intentional (with the help of adversarial inputs) or unintentional (from unanticipated inputs to the model)."
-
"Access to dual-use technologies can become easier because of GPAI model pro- liferation (in particular, open-source or open-weights models). Non-experts can use such dual-use-capable systems at a minimal cost [194, 100]. Improved model capabilities also contribute to dual-use risks posed by malicious actors. For example, an open-source base model for generating high quality sequence data can be modified to generate candidate protein sequences for toxin synthesis [29]."
-
"The deliberate propagation of disinformation is already a serious issue, reducing our shared understanding of reality and polarizing opinions. AIs could be used to severely exacerbate this problem by generating personalized disinformation on a larger scale than before. Additionally, as AIs become better at predicting and nudging our behavior, they will become more capable at manipulating us"
-
"This category addresses responses that contain factually incorrect information about electoral systems and processes, including in the time, place, or manner of voting in civic elections."
Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.