MIT AI Risk Repository
Browse AI risks
977 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.
-
"Denial of or loss of access to welfare benefits, pensions, housing, etc due to the malfunction, use or misuse of a technology system"
-
69.06.02 · Risk Sub-Category
Toxic and disrespectful content
Discriminatory and exclusionary language
-
-
"Like virtual applications of AI, EAI can display bias towards and dis- criminate against users. When EAI systems are placed in positions of power, their biases could have significant impacts on fairness in everyday interactions and on general social dynamics [105, 106]."
-
73.04.01 · Risk Sub-Category
LLM-Systems Can Be Untrustworthy
Harms of Representation and Other Biases
"A pretrained LLM generally has many of the stereotypical biases commonly present in the human society (Touvron et al., 2023). This makes it difficult for users to trust that LLMs will work well for them and not produce unfair or biased responses. Appropriate finetuning can effectively limit the bias displayed in LLM outputs in a variety of situations, e.g. when models are explicitly prompted with stereotypes (Wang et al., 2023k), but it does not ‘solve’ the problem. Even after finetuning, biases often resurface when deliberately elicited (Wang et al., 2023k), or under novel scenarios, e.g. in
-
02.01.00 · Risk Category
"The LLM-generated content sometimes contains biased, toxic, and private information"
-
"Toxicity means the generated content contains rude, disrespectful, and even illegal information"
-
02.11.00 · Risk Category
"Inputting a prompt contain an unsafe topic (e.g., notsuitable-for-work (NSFW) content) by a benign user. "
-
04.01.00 · Risk Category
This typically refers to rude, harmful, or inappropriate expressions.
-
04.04.00 · Risk Category
The controversial views expressed by large models are also a widely discussed concern. Bang et al. (2021) evaluated several large models and found that they occasionally express inappropriate or extremist views when discussing political top-ics. Furthermore, models like ChatGPT (OpenAI, 2022) that claim political neutrality and aim to provide objective information for users have been shown to exhibit notable left-leaning political biases in areas like economics, social policy, foreign affairs, and civil liberties.
-
05.03.00 · Risk Category
Generating unethical, fraudulent, toxic, violent, pornographic, or other harmful content is a further predominant concern, again focusing notably on LLMs and text-to-image models. Numerous studies highlight the risks associated with the intentional creation of disinformation, fake news, propaganda, or deepfakes, underscoring their significant threat to the integrity of public discourse and the trust in credible media. Additionally, papers explore the potential for generative models to aid in criminal activities, incidents of self-harm, identity theft, or impersonation. Furthermore, the literat
-
13.01.02 · Risk Sub-Category
Impacts: The Technical Base System
Cultural Values and Sensitive Content
"Cultural values are specific to groups and sensitive content is normative. Sensitive topics also vary by culture and can include hate speech, which itself is contingent on cultural norms of acceptability."
-
16.01.02 · Risk Sub-Category
Risk area 1: Discrimination, Hate speech and Exclusion
Hate speech and offensive language
"LMs may generate language that includes profanities, identity attacks, insults, threats, language that incites violence, or language that causes justified offence as such language is prominent online [57, 64, 143,191]. This language risks causing offence, psychological harm, and inciting hate or violence."
-
"LM’s may predict hate speech or other language that is “toxic”. While there is no single agreed definition of what constitutes hate speech or toxic speech (Fortuna and Nunes, 2018; Persily and Tucker, 2020; Schmidt and Wiegand, 2017), proposed definitions often include profanities, identity attacks, sleights, insults, threats, sexually explicit content, demeaning language, language that incites violence, or ‘hostile and malicious language targeted at a person or group because of their actual or perceived innate characteristics’ (Fortuna and Nunes, 2018; Gorwa et al., 2020; PerspectiveAPI)"
-
"Generating content that violates community standards, including harming or inciting hatred or violence against individuals and groups (e.g. gore, child sexual abuse material, profanities, identity attacks)"
-
23.01.00 · Risk Category
"This category addresses responses that enable, encourage, or endorse the commission of violent crimes."
-
23.02.00 · Risk Category
"This category addresses responses that enable, encourage, or endorse the commission of non-violent crimes."
-
23.03.00 · Risk Category
"This category addresses responses that enable, encourage, or endorse the commission of sex-related crimes"
-
23.04.00 · Risk Category
"This category addresses responses that contain, describe, enable, encourage, or endorse the sexual abuse of children (i.e., under 18 years old) or the acquisition of child sexual abuse material."
-
23.06.00 · Risk Category
"This category addresses responses that enable, encourage, or endorse acts of intentional self-harm."
-
"This category addresses responses that demean or dehumanize people on the basis of their sensitive, personal characteristics."
-
23.13.00 · Risk Category
"This category addresses responses that contain erotica. This includes depictions of sexual behavior, genitalia, and overtly sexualized body parts."
-
"Second, because LLMs are trained on internet text data, there is also a risk that model weights encode functions which, if deployed in particular contexts, would violate social norms of that context. Following the principles of contextual integrity, it may be that models deviate from information sharing norms as a result of their training. Overcoming this challenge requires two types of infrastructure: one for keeping track of social norms in context, and another for ensuring that models adhere to them. Keeping track of what social norms are presently at play is an active research area. Surfa
-
"Insulting content generated by LMs is a highly visible and frequently mentioned safety issue. Mostly, it is unfriendly, disrespectful, or ridiculous content that makes users uncomfortable and drives them away. It is extremely hazardous and could have negative social consequences."
-
"The model output contains illegal and criminal attitudes, behaviors, or motivations, such as incitement to commit crimes, fraud, and rumor propagation. These contents may hurt users and have negative societal repercussions."
-
"For some sensitive and controversial topics (especially on politics), LMs tend to generate biased, misleading, and inaccurate content. For example, there may be a tendency to support a specific political position, leading to discrimination or exclusion of other political viewpoints."
-
28.01.00 · Risk Category
"This category is about threat, insult, scorn, profanity, sarcasm, impoliteness, etc. LLMs are required to identify and oppose these offensive contents or actions."
-
Avoiding unsafe and illegal outputs, and leaking private information
-
LLMs are found to generate answers that contain violent content or generate content that responds to questions that solicit information about violent behaviors
-
LLMs have been shown to be a convenient tool for soliciting advice on accessing, purchasing (illegally), and creating illegal substances, as well as for dangerous use of them
-
LLMs can be leveraged to solicit answers that contain harmful content to children and youth
-
LLMs have the capability to generate sex-explicit conversations, and erotic texts, and to recommend websites with sexual content
-
30.06.00 · Risk Category
LLMs are expected to reflect social values by avoiding the use of offensive language toward specific groups of users, being sensitive to topics that can create instability, as well as being sympathetic when users are seeking emotional support
-
language being rude, disrespectful, threatening, or identity-attacking toward certain groups of the user population (culture, race, and gender etc)
-
"Harmful or inappropriate content produced by generative AI includes but is not limited to violent content, the use of offensive language, discriminative content, and pornography. Although OpenAI has set up a content policy for ChatGPT, harmful or inappropriate content can still appear due to reasons such as algorithmic limitations or jailbreaking (i.e., removal of restrictions imposed). The language models’ ability to understand or generate harmful or offensive content is referred to as toxicity (Zhuo et al., 2023). Toxicity can bring harm to society and damage the harmony of the community. H
-
45.02.01 · Risk Sub-Category
Safety risks in AI Applications
Cyberspace risks (Risks of information and content safety)
"AI-generated or synthesized content can lead to the spread of false information, discrimination and bias, privacy leakage, and infringement issues, threatening the safety of citizens' lives and property, national security, ideological security, and causing ethical risks. If users’ inputs contain harmful content, the model may output illegal or damaging information without robust security mechanisms."
-
48.03.00 · Risk Category
"Eased production of and access to violent, inciting, radicalizing, or threatening content as well as recommendations to carry out self-harm or conduct illegal activities. Includes difficulty controlling public exposure to hateful and disparaging or stereotyping content."
-
48.11.00 · Risk Category
"Eased production of and access to obscene, degrading, and/or abusive imagery which can cause harm, including synthetic child sexual abuse material (CSAM), and nonconsensual intimate images (NCII) of adults."
-
-
-
50.02.01 · Risk Sub-Category
Violence and extremism (Supporting malicious organized groups)
—
-
—
-
—
-
—
-
50.02.08 · Risk Sub-Category
Hate/Toxicity (Hate Speech: Inciting/Promoting/Expressing Hatred)
—
-
—
-
—
-
—
-
—
-
—
-
—
-
—
Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.