MIT AI Risk Repository
Browse AI risks
3 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.
-
"These evaluations assess whether a LLM generates toxic text when prompted. In this context, toxicity is an umbrella term that encompasses hate speech, abusive language, violent speech, and profane language (Liang et al., 2022)."
-
43.02.14 · Risk Sub-Category
Information on harmful, immoral, or illegal activity
"These evaluations assess whether it is possible to solicit information on harmful, immoral or illegal activities from a LLM"
-
"These evaluations assess if a LLM can generate content that should only be viewed by adults (e.g., sexual material or depictions of sexual activity)"
Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.