MIT AI Risk Repository

Browse AI risks

116 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

116 entries · page 1 of 3

  1. 02.01.00 · Risk Category

    Harmful Content

    "The LLM-generated content sometimes contains biased, toxic, and private information"

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  2. 02.01.02 · Risk Sub-Category

    Harmful Content

    Toxicity

    "Toxicity means the generated content contains rude, disrespectful, and even illegal information"

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  3. 02.08.01 · Risk Sub-Category

    Toxicity and Bias Tendencies

    Toxic Training Data

    "Following previous studies [96], [97], toxic data in LLMs is defined as rude, disrespectful, or unreasonable language that is opposite to a polite, positive, and healthy language environment, including hate speech, offensive utterance, profanities, and threats [91]."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  4. "Inputting a prompt contain an unsafe topic (e.g., notsuitable-for-work (NSFW) content) by a benign user. "

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  5. This typically refers to rude, harmful, or inappropriate expressions.

    From Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements (Deng2023)

  6. 04.04.00 · Risk Category

    Controversial Opinions

    The controversial views expressed by large models are also a widely discussed concern. Bang et al. (2021) evaluated several large models and found that they occasionally express inappropriate or extremist views when discussing political top-ics. Furthermore, models like ChatGPT (OpenAI, 2022) that claim political neutrality and aim to provide objective information for users have been shown to exhibit notable left-leaning political biases in areas like economics, social policy, foreign affairs, and civil liberties.

    From Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements (Deng2023)

  7. Generating unethical, fraudulent, toxic, violent, pornographic, or other harmful content is a further predominant concern, again focusing notably on LLMs and text-to-image models. Numerous studies highlight the risks associated with the intentional creation of disinformation, fake news, propaganda, or deepfakes, underscoring their significant threat to the integrity of public discourse and the trust in credible media. Additionally, papers explore the potential for generative models to aid in criminal activities, incidents of self-harm, identity theft, or impersonation. Furthermore, the literat

    From Mapping the Ethics of Generative AI: A Comprehensive Scoping Review (Hagendorff2024)

  8. 13.01.02 · Risk Sub-Category

    Impacts: The Technical Base System

    Cultural Values and Sensitive Content

    "Cultural values are specific to groups and sensitive content is normative. Sensitive topics also vary by culture and can include hate speech, which itself is contingent on cultural norms of acceptability."

    From Evaluating the Social Impact of Generative AI Systems in Systems and Society (Solaiman2023)

  9. "Speech can create a range of harms, such as promoting social stereotypes that perpetuate the derogatory representation or unfair treatment of marginalised groups [22], inciting hate or violence [57], causing profound offence [199], or reinforcing social norms that exclude or marginalise identities [15,58]. LMs that faithfully mirror harmful language present in the training data can reproduce these harms. Unfair treatment can also emerge from LMs that perform better for some social groups than others [18]. These risks have been widely known, observed and documented in LMs. Mitigation approache

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  10. 16.01.02 · Risk Sub-Category

    Risk area 1: Discrimination, Hate speech and Exclusion

    Hate speech and offensive language

    "LMs may generate language that includes profanities, identity attacks, insults, threats, language that incites violence, or language that causes justified offence as such language is prominent online [57, 64, 143,191]. This language risks causing offence, psychological harm, and inciting hate or violence."

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  11. 17.01.03 · Risk Sub-Category

    Discrimination, Exclusion and Toxicity

    Toxic language

    "LM’s may predict hate speech or other language that is “toxic”. While there is no single agreed definition of what constitutes hate speech or toxic speech (Fortuna and Nunes, 2018; Persily and Tucker, 2020; Schmidt and Wiegand, 2017), proposed definitions often include profanities, identity attacks, sleights, insults, threats, sexually explicit content, demeaning language, language that incites violence, or ‘hostile and malicious language targeted at a person or group because of their actual or perceived innate characteristics’ (Fortuna and Nunes, 2018; Gorwa et al., 2020; PerspectiveAPI)"

    From Ethical and social risks of harm from language models (Weidinger2021)

  12. 18.01.03 · Risk Sub-Category

    Representation & Toxicity Harms

    Toxic content

    "Generating content that violates community standards, including harming or inciting hatred or violence against individuals and groups (e.g. gore, child sexual abuse material, profanities, identity attacks)"

    From Sociotechnical Safety Evaluation of Generative AI Systems (Weidinger2023)

  13. 23.01.00 · Risk Category

    Violent crimes

    "This category addresses responses that enable, encourage, or endorse the commission of violent crimes."

    From Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024)

  14. 23.01.01 · Risk Sub-Category

    Violent crimes

    Mass violence

  15. 23.01.02 · Risk Sub-Category

    Violent crimes

    Murder

  16. 23.01.03 · Risk Sub-Category

    Violent crimes

    Physical assault against a person

  17. 23.01.04 · Risk Sub-Category

    Violent crimes

    Violent domestic abuse

  18. 23.01.05 · Risk Sub-Category

    Violent crimes

    Terror (Terror groups, Terror actors, Terrorist actions)

  19. 23.02.00 · Risk Category

    Non-violent crimes

    "This category addresses responses that enable, encourage, or endorse the commission of non-violent crimes."

    From Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024)

  20. 23.02.01 · Risk Sub-Category

    Non-violent crimes

    Theft

  21. 23.02.02 · Risk Sub-Category

    Non-violent crimes

    Modern slavery and human trafficking

  22. 23.02.03 · Risk Sub-Category

    Non-violent crimes

    Child abuse (non-sexual)

  23. 23.02.04 · Risk Sub-Category

    Non-violent crimes

    Property damage

  24. 23.02.05 · Risk Sub-Category

    Non-violent crimes

    Financial crime (Scams, Fraud)

  25. 23.02.06 · Risk Sub-Category

    Non-violent crimes

    Sale, purchase or use of illegal, highly regulated or unregistered items (Dangerous substances, Dangerous weapons, Prohibited wildlife, Fake goods)

  26. 23.03.00 · Risk Category

    Sex-related crimes

    "This category addresses responses that enable, encourage, or endorse the commission of sex-related crimes"

    From Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024)

  27. 23.03.01 · Risk Sub-Category

    Sex-related crimes

    Sexual assault (e.g., rape)

  28. 23.03.02 · Risk Sub-Category

    Sex-related crimes

    Sexual harassment

  29. 23.03.03 · Risk Sub-Category

    Sex-related crimes

    Sex trafficking

  30. 23.03.04 · Risk Sub-Category

    Sex-related crimes

    Prostitution

  31. 23.03.05 · Risk Sub-Category

    Sex-related crimes

    Intimate-image based abuse

  32. 23.03.06 · Risk Sub-Category

    Sex-related crimes

    Indecent exposure (i.e. showing sexual body parts or "flashing")

  33. "This category addresses responses that contain, describe, enable, encourage, or endorse the sexual abuse of children (i.e., under 18 years old) or the acquisition of child sexual abuse material."

    From Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024)

  34. 23.04.01 · Risk Sub-Category

    Child sexual exploitation

    Inappropriate/harmful relationships with children. This includes adults forming romantic relationships with children or grooming them

  35. 23.04.02 · Risk Sub-Category

    Child sexual exploitation

    Sexual abuse of children, including the sexualisation of children

  36. 23.04.03 · Risk Sub-Category

    Child sexual exploitation

    Child Sexual Abuse Material (CSAM). This includes erotic materials involving children

  37. 23.05.01 · Risk Sub-Category

    Indiscriminate weapons (CBRNE)

    CBRNE (all)

  38. 23.05.02 · Risk Sub-Category

    Indiscriminate weapons (CBRNE)

    Chemical weapons (ex: nerve gas)

  39. 23.05.03 · Risk Sub-Category

    Indiscriminate weapons (CBRNE)

    Biological weapons (ex: anthrax)

  40. 23.05.04 · Risk Sub-Category

    Indiscriminate weapons (CBRNE)

    Radiological weapons (ex: dirty bombs)

  41. 23.05.05 · Risk Sub-Category

    Indiscriminate weapons (CBRNE)

    Nuclear weapons (ex: fission or fusion bombs)

  42. 23.05.06 · Risk Sub-Category

    Indiscriminate weapons (CBRNE)

    High-yield explosive weapons (ex: cluster munitions)

  43. 23.06.00 · Risk Category

    Suicide & self-harm

    "This category addresses responses that enable, encourage, or endorse acts of intentional self-harm."

    From Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024)

  44. 23.06.01 · Risk Sub-Category

    Suicide & self-harm

    Suicide

  45. 23.06.02 · Risk Sub-Category

    Suicide & self-harm

    Self-harm

  46. 23.06.03 · Risk Sub-Category

    Suicide & self-harm

    Eating disorders

  47. 23.06.04 · Risk Sub-Category

    Suicide & self-harm

    Dangerous challenges and hoaxes that can lead individuals to harm themselves

  48. 23.07.00 · Risk Category

    Hate

    "This category addresses responses that demean or dehumanize people on the basis of their sensitive, personal characteristics."

    From Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024)

  49. 23.07.01 · Risk Sub-Category

    Hate

    Demeaning and derogatory remarks (Mental capacity, Sensory and physical attributes, Behavioral attributes)

  50. 23.07.02 · Risk Sub-Category

    Hate

    Inflicting physical harm (Mass violence and genocide, Murder and individual violence)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.