MIT AI Risk Repository
Browse AI risks
2,500 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.
-
16.01.02 · Risk Sub-Category
Risk area 1: Discrimination, Hate speech and Exclusion
Hate speech and offensive language
"LMs may generate language that includes profanities, identity attacks, insults, threats, language that incites violence, or language that causes justified offence as such language is prominent online [57, 64, 143,191]. This language risks causing offence, psychological harm, and inciting hate or violence."
-
"LM’s may predict hate speech or other language that is “toxic”. While there is no single agreed definition of what constitutes hate speech or toxic speech (Fortuna and Nunes, 2018; Persily and Tucker, 2020; Schmidt and Wiegand, 2017), proposed definitions often include profanities, identity attacks, sleights, insults, threats, sexually explicit content, demeaning language, language that incites violence, or ‘hostile and malicious language targeted at a person or group because of their actual or perceived innate characteristics’ (Fortuna and Nunes, 2018; Gorwa et al., 2020; PerspectiveAPI)"
-
"Generating content that violates community standards, including harming or inciting hatred or violence against individuals and groups (e.g. gore, child sexual abuse material, profanities, identity attacks)"
-
23.01.00 · Risk Category
"This category addresses responses that enable, encourage, or endorse the commission of violent crimes."
-
—
-
—
-
—
-
—
-
23.01.05 · Risk Sub-Category
Terror (Terror groups, Terror actors, Terrorist actions)
—
-
23.02.00 · Risk Category
"This category addresses responses that enable, encourage, or endorse the commission of non-violent crimes."
-
—
-
—
-
—
-
—
-
—
-
23.02.06 · Risk Sub-Category
Sale, purchase or use of illegal, highly regulated or unregistered items (Dangerous substances, Dangerous weapons, Prohibited wildlife, Fake goods)
—
-
23.03.00 · Risk Category
"This category addresses responses that enable, encourage, or endorse the commission of sex-related crimes"
-
—
-
—
-
—
-
—
-
—
-
23.03.06 · Risk Sub-Category
Indecent exposure (i.e. showing sexual body parts or "flashing")
—
-
23.04.00 · Risk Category
"This category addresses responses that contain, describe, enable, encourage, or endorse the sexual abuse of children (i.e., under 18 years old) or the acquisition of child sexual abuse material."
-
23.04.01 · Risk Sub-Category
Inappropriate/harmful relationships with children. This includes adults forming romantic relationships with children or grooming them
—
-
23.04.02 · Risk Sub-Category
Sexual abuse of children, including the sexualisation of children
—
-
23.04.03 · Risk Sub-Category
Child Sexual Abuse Material (CSAM). This includes erotic materials involving children
—
-
—
-
—
-
—
-
—
-
23.05.05 · Risk Sub-Category
Indiscriminate weapons (CBRNE)
Nuclear weapons (ex: fission or fusion bombs)
—
-
23.05.06 · Risk Sub-Category
Indiscriminate weapons (CBRNE)
High-yield explosive weapons (ex: cluster munitions)
—
-
23.06.00 · Risk Category
"This category addresses responses that enable, encourage, or endorse acts of intentional self-harm."
-
—
-
—
-
—
-
23.06.04 · Risk Sub-Category
Dangerous challenges and hoaxes that can lead individuals to harm themselves
—
-
"This category addresses responses that demean or dehumanize people on the basis of their sensitive, personal characteristics."
-
23.07.01 · Risk Sub-Category
Demeaning and derogatory remarks (Mental capacity, Sensory and physical attributes, Behavioral attributes)
—
-
23.07.02 · Risk Sub-Category
Inflicting physical harm (Mass violence and genocide, Murder and individual violence)
—
-
23.07.03 · Risk Sub-Category
Inflicting emotional harm (Emotional abuse, Undermine and make feel unwelcome)
—
-
23.07.04 · Risk Sub-Category
Exclusion and isolation (Social exclusion, Political exclusion, Economic exclusion)
—
-
23.13.00 · Risk Category
"This category addresses responses that contain erotica. This includes depictions of sexual behavior, genitalia, and overtly sexualized body parts."
-
"Second, because LLMs are trained on internet text data, there is also a risk that model weights encode functions which, if deployed in particular contexts, would violate social norms of that context. Following the principles of contextual integrity, it may be that models deviate from information sharing norms as a result of their training. Overcoming this challenge requires two types of infrastructure: one for keeping track of social norms in context, and another for ensuring that models adhere to them. Keeping track of what social norms are presently at play is an active research area. Surfa
-
"Insulting content generated by LMs is a highly visible and frequently mentioned safety issue. Mostly, it is unfriendly, disrespectful, or ridiculous content that makes users uncomfortable and drives them away. It is extremely hazardous and could have negative social consequences."
-
"The model output contains illegal and criminal attitudes, behaviors, or motivations, such as incitement to commit crimes, fraud, and rumor propagation. These contents may hurt users and have negative societal repercussions."
-
"For some sensitive and controversial topics (especially on politics), LMs tend to generate biased, misleading, and inaccurate content. For example, there may be a tendency to support a specific political position, leading to discrimination or exclusion of other political viewpoints."
-
28.01.00 · Risk Category
"This category is about threat, insult, scorn, profanity, sarcasm, impoliteness, etc. LLMs are required to identify and oppose these offensive contents or actions."
-
Avoiding unsafe and illegal outputs, and leaking private information
Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.