MIT AI Risk Repository
Browse AI risks
116 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.
-
02.01.00 · Risk Category
"The LLM-generated content sometimes contains biased, toxic, and private information"
-
"Toxicity means the generated content contains rude, disrespectful, and even illegal information"
-
"Following previous studies [96], [97], toxic data in LLMs is defined as rude, disrespectful, or unreasonable language that is opposite to a polite, positive, and healthy language environment, including hate speech, offensive utterance, profanities, and threats [91]."
-
02.11.00 · Risk Category
"Inputting a prompt contain an unsafe topic (e.g., notsuitable-for-work (NSFW) content) by a benign user. "
-
04.01.00 · Risk Category
This typically refers to rude, harmful, or inappropriate expressions.
-
04.04.00 · Risk Category
The controversial views expressed by large models are also a widely discussed concern. Bang et al. (2021) evaluated several large models and found that they occasionally express inappropriate or extremist views when discussing political top-ics. Furthermore, models like ChatGPT (OpenAI, 2022) that claim political neutrality and aim to provide objective information for users have been shown to exhibit notable left-leaning political biases in areas like economics, social policy, foreign affairs, and civil liberties.
-
05.03.00 · Risk Category
Generating unethical, fraudulent, toxic, violent, pornographic, or other harmful content is a further predominant concern, again focusing notably on LLMs and text-to-image models. Numerous studies highlight the risks associated with the intentional creation of disinformation, fake news, propaganda, or deepfakes, underscoring their significant threat to the integrity of public discourse and the trust in credible media. Additionally, papers explore the potential for generative models to aid in criminal activities, incidents of self-harm, identity theft, or impersonation. Furthermore, the literat
-
13.01.02 · Risk Sub-Category
Impacts: The Technical Base System
Cultural Values and Sensitive Content
"Cultural values are specific to groups and sensitive content is normative. Sensitive topics also vary by culture and can include hate speech, which itself is contingent on cultural norms of acceptability."
-
16.01.00 · Risk Category
"Speech can create a range of harms, such as promoting social stereotypes that perpetuate the derogatory representation or unfair treatment of marginalised groups [22], inciting hate or violence [57], causing profound offence [199], or reinforcing social norms that exclude or marginalise identities [15,58]. LMs that faithfully mirror harmful language present in the training data can reproduce these harms. Unfair treatment can also emerge from LMs that perform better for some social groups than others [18]. These risks have been widely known, observed and documented in LMs. Mitigation approache
-
16.01.02 · Risk Sub-Category
Risk area 1: Discrimination, Hate speech and Exclusion
Hate speech and offensive language
"LMs may generate language that includes profanities, identity attacks, insults, threats, language that incites violence, or language that causes justified offence as such language is prominent online [57, 64, 143,191]. This language risks causing offence, psychological harm, and inciting hate or violence."
-
"LM’s may predict hate speech or other language that is “toxic”. While there is no single agreed definition of what constitutes hate speech or toxic speech (Fortuna and Nunes, 2018; Persily and Tucker, 2020; Schmidt and Wiegand, 2017), proposed definitions often include profanities, identity attacks, sleights, insults, threats, sexually explicit content, demeaning language, language that incites violence, or ‘hostile and malicious language targeted at a person or group because of their actual or perceived innate characteristics’ (Fortuna and Nunes, 2018; Gorwa et al., 2020; PerspectiveAPI)"
-
"Generating content that violates community standards, including harming or inciting hatred or violence against individuals and groups (e.g. gore, child sexual abuse material, profanities, identity attacks)"
-
23.01.00 · Risk Category
"This category addresses responses that enable, encourage, or endorse the commission of violent crimes."
-
—
-
—
-
—
-
—
-
23.01.05 · Risk Sub-Category
Terror (Terror groups, Terror actors, Terrorist actions)
—
-
23.02.00 · Risk Category
"This category addresses responses that enable, encourage, or endorse the commission of non-violent crimes."
-
—
-
—
-
—
-
—
-
—
-
23.02.06 · Risk Sub-Category
Sale, purchase or use of illegal, highly regulated or unregistered items (Dangerous substances, Dangerous weapons, Prohibited wildlife, Fake goods)
—
-
23.03.00 · Risk Category
"This category addresses responses that enable, encourage, or endorse the commission of sex-related crimes"
-
—
-
—
-
—
-
—
-
—
-
23.03.06 · Risk Sub-Category
Indecent exposure (i.e. showing sexual body parts or "flashing")
—
-
23.04.00 · Risk Category
"This category addresses responses that contain, describe, enable, encourage, or endorse the sexual abuse of children (i.e., under 18 years old) or the acquisition of child sexual abuse material."
-
23.04.01 · Risk Sub-Category
Inappropriate/harmful relationships with children. This includes adults forming romantic relationships with children or grooming them
—
-
23.04.02 · Risk Sub-Category
Sexual abuse of children, including the sexualisation of children
—
-
23.04.03 · Risk Sub-Category
Child Sexual Abuse Material (CSAM). This includes erotic materials involving children
—
-
—
-
—
-
—
-
—
-
23.05.05 · Risk Sub-Category
Indiscriminate weapons (CBRNE)
Nuclear weapons (ex: fission or fusion bombs)
—
-
23.05.06 · Risk Sub-Category
Indiscriminate weapons (CBRNE)
High-yield explosive weapons (ex: cluster munitions)
—
-
23.06.00 · Risk Category
"This category addresses responses that enable, encourage, or endorse acts of intentional self-harm."
-
—
-
—
-
—
-
23.06.04 · Risk Sub-Category
Dangerous challenges and hoaxes that can lead individuals to harm themselves
—
-
"This category addresses responses that demean or dehumanize people on the basis of their sensitive, personal characteristics."
-
23.07.01 · Risk Sub-Category
Demeaning and derogatory remarks (Mental capacity, Sensory and physical attributes, Behavioral attributes)
—
-
23.07.02 · Risk Sub-Category
Inflicting physical harm (Mass violence and genocide, Murder and individual violence)
—
Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.