MIT AI Risk Repository
Browse AI risks
116 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.
-
—
-
—
-
62.31.06 · Risk Sub-Category
Impacts of AI (Societal Impacts)
Generation of illegal or harmful content
"Generative models can create illegal, harmful, or discriminatory content [196], such as sexual abuse material, at scale. Current access controls (e.g., API access filters) are not effective against all user queries in generating such content."
-
62.31.07 · Risk Sub-Category
Impacts of AI (Societal Impacts)
Unintentional generation of harmful content
"Generative models can create harmful or discriminatory content from benign user requests. Models can exhibit bias to particular harmful styles of generation (e.g., sexualization of photos of women [87] in the case of image generation models) or they can generate toxic, misleading, or violent data (e.g., a model generating jokes can use ethnic stereotypes or slurs to deliver humor)."
-
"Toxic output occurs when the model produces hateful, abusive, and profane (HAP) or obscene content. This also includes behaviors like bullying."
-
"A model might generate language that leads to physical harm The language might include overtly violent, covertly dangerous, or otherwise indirectly unsafe statements."
-
"Generating content that violates community standards, including harming or inciting hatred or violence against groups (e.g. gore, sexual content of children, profanities, identity attacks)"
-
69.03.00 · Risk Category
"The chatbot shares information that can be used to do something dangerous or illegal."
-
—
-
69.06.00 · Risk Category
"The chatbot verbally attacks or undermines an individual, group, or organization. 7."
-
-
-
69.06.03 · Risk Sub-Category
Toxic and disrespectful content
Subversive or aggressive political opinions
-
-
-
-
—
-
69.10.00 · Risk Category
"The chatbot participates in morally or socially objectionable conversational activities with its user that could be emotionally damaging to its user or third parties."
-
"Toxicity in LLMs refers to the generation of harmful, offensive, or inappropriate content that can cause harm to individuals or groups. Both explicit and implicit forms of toxicity can be generated by LLMs, posing significant risks to society. Explicit toxicity encompasses a wide range of negative behaviors, including hate speech, harassment, cyberbullying, rude, and disrespectful comments, derogatory language, as well as allocational harms [2, 62, 90]. Besides, implicit toxicity does not involve overtly harmful language but may manifest through subtle forms such as sarcasm, irony, and humor,
Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.