MIT AI Risk Repository

Browse AI risks

224 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset Also filtered by framework Gipiškis2024 ×

224 entries · page 3 of 5

  1. 16.01.02 · Risk Sub-Category

    Risk area 1: Discrimination, Hate speech and Exclusion

    Hate speech and offensive language

    "LMs may generate language that includes profanities, identity attacks, insults, threats, language that incites violence, or language that causes justified offence as such language is prominent online [57, 64, 143,191]. This language risks causing offence, psychological harm, and inciting hate or violence."

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  2. 17.01.03 · Risk Sub-Category

    Discrimination, Exclusion and Toxicity

    Toxic language

    "LM’s may predict hate speech or other language that is “toxic”. While there is no single agreed definition of what constitutes hate speech or toxic speech (Fortuna and Nunes, 2018; Persily and Tucker, 2020; Schmidt and Wiegand, 2017), proposed definitions often include profanities, identity attacks, sleights, insults, threats, sexually explicit content, demeaning language, language that incites violence, or ‘hostile and malicious language targeted at a person or group because of their actual or perceived innate characteristics’ (Fortuna and Nunes, 2018; Gorwa et al., 2020; PerspectiveAPI)"

    From Ethical and social risks of harm from language models (Weidinger2021)

  3. 18.01.03 · Risk Sub-Category

    Representation & Toxicity Harms

    Toxic content

    "Generating content that violates community standards, including harming or inciting hatred or violence against individuals and groups (e.g. gore, child sexual abuse material, profanities, identity attacks)"

    From Sociotechnical Safety Evaluation of Generative AI Systems (Weidinger2023)

  4. 23.01.00 · Risk Category

    Violent crimes

    "This category addresses responses that enable, encourage, or endorse the commission of violent crimes."

    From Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024)

  5. 23.01.01 · Risk Sub-Category

    Violent crimes

    Mass violence

  6. 23.01.02 · Risk Sub-Category

    Violent crimes

    Murder

  7. 23.01.03 · Risk Sub-Category

    Violent crimes

    Physical assault against a person

  8. 23.01.04 · Risk Sub-Category

    Violent crimes

    Violent domestic abuse

  9. 23.01.05 · Risk Sub-Category

    Violent crimes

    Terror (Terror groups, Terror actors, Terrorist actions)

  10. 23.02.00 · Risk Category

    Non-violent crimes

    "This category addresses responses that enable, encourage, or endorse the commission of non-violent crimes."

    From Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024)

  11. 23.02.01 · Risk Sub-Category

    Non-violent crimes

    Theft

  12. 23.02.02 · Risk Sub-Category

    Non-violent crimes

    Modern slavery and human trafficking

  13. 23.02.03 · Risk Sub-Category

    Non-violent crimes

    Child abuse (non-sexual)

  14. 23.02.04 · Risk Sub-Category

    Non-violent crimes

    Property damage

  15. 23.02.05 · Risk Sub-Category

    Non-violent crimes

    Financial crime (Scams, Fraud)

  16. 23.02.06 · Risk Sub-Category

    Non-violent crimes

    Sale, purchase or use of illegal, highly regulated or unregistered items (Dangerous substances, Dangerous weapons, Prohibited wildlife, Fake goods)

  17. 23.03.00 · Risk Category

    Sex-related crimes

    "This category addresses responses that enable, encourage, or endorse the commission of sex-related crimes"

    From Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024)

  18. 23.03.01 · Risk Sub-Category

    Sex-related crimes

    Sexual assault (e.g., rape)

  19. 23.03.02 · Risk Sub-Category

    Sex-related crimes

    Sexual harassment

  20. 23.03.03 · Risk Sub-Category

    Sex-related crimes

    Sex trafficking

  21. 23.03.04 · Risk Sub-Category

    Sex-related crimes

    Prostitution

  22. 23.03.05 · Risk Sub-Category

    Sex-related crimes

    Intimate-image based abuse

  23. 23.03.06 · Risk Sub-Category

    Sex-related crimes

    Indecent exposure (i.e. showing sexual body parts or "flashing")

  24. "This category addresses responses that contain, describe, enable, encourage, or endorse the sexual abuse of children (i.e., under 18 years old) or the acquisition of child sexual abuse material."

    From Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024)

  25. 23.04.01 · Risk Sub-Category

    Child sexual exploitation

    Inappropriate/harmful relationships with children. This includes adults forming romantic relationships with children or grooming them

  26. 23.04.02 · Risk Sub-Category

    Child sexual exploitation

    Sexual abuse of children, including the sexualisation of children

  27. 23.04.03 · Risk Sub-Category

    Child sexual exploitation

    Child Sexual Abuse Material (CSAM). This includes erotic materials involving children

  28. 23.05.01 · Risk Sub-Category

    Indiscriminate weapons (CBRNE)

    CBRNE (all)

  29. 23.05.02 · Risk Sub-Category

    Indiscriminate weapons (CBRNE)

    Chemical weapons (ex: nerve gas)

  30. 23.05.03 · Risk Sub-Category

    Indiscriminate weapons (CBRNE)

    Biological weapons (ex: anthrax)

  31. 23.05.04 · Risk Sub-Category

    Indiscriminate weapons (CBRNE)

    Radiological weapons (ex: dirty bombs)

  32. 23.05.05 · Risk Sub-Category

    Indiscriminate weapons (CBRNE)

    Nuclear weapons (ex: fission or fusion bombs)

  33. 23.05.06 · Risk Sub-Category

    Indiscriminate weapons (CBRNE)

    High-yield explosive weapons (ex: cluster munitions)

  34. 23.06.00 · Risk Category

    Suicide & self-harm

    "This category addresses responses that enable, encourage, or endorse acts of intentional self-harm."

    From Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024)

  35. 23.06.01 · Risk Sub-Category

    Suicide & self-harm

    Suicide

  36. 23.06.02 · Risk Sub-Category

    Suicide & self-harm

    Self-harm

  37. 23.06.03 · Risk Sub-Category

    Suicide & self-harm

    Eating disorders

  38. 23.06.04 · Risk Sub-Category

    Suicide & self-harm

    Dangerous challenges and hoaxes that can lead individuals to harm themselves

  39. 23.07.00 · Risk Category

    Hate

    "This category addresses responses that demean or dehumanize people on the basis of their sensitive, personal characteristics."

    From Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024)

  40. 23.07.01 · Risk Sub-Category

    Hate

    Demeaning and derogatory remarks (Mental capacity, Sensory and physical attributes, Behavioral attributes)

  41. 23.07.02 · Risk Sub-Category

    Hate

    Inflicting physical harm (Mass violence and genocide, Murder and individual violence)

  42. 23.07.03 · Risk Sub-Category

    Hate

    Inflicting emotional harm (Emotional abuse, Undermine and make feel unwelcome)

  43. 23.07.04 · Risk Sub-Category

    Hate

    Exclusion and isolation (Social exclusion, Political exclusion, Economic exclusion)

  44. 23.13.00 · Risk Category

    Sexual Content

    "This category addresses responses that contain erotica. This includes depictions of sexual behavior, genitalia, and overtly sexualized body parts."

    From Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024)

  45. 24.08.02 · Risk Sub-Category

    Privacy

    Violation of social norms

    "Second, because LLMs are trained on internet text data, there is also a risk that model weights encode functions which, if deployed in particular contexts, would violate social norms of that context. Following the principles of contextual integrity, it may be that models deviate from information sharing norms as a result of their training. Overcoming this challenge requires two types of infrastructure: one for keeping track of social norms in context, and another for ensuring that models adhere to them. Keeping track of what social norms are presently at play is an active research area. Surfa

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  46. 27.01.01 · Risk Sub-Category

    Typical safety scenarios

    Insult

    "Insulting content generated by LMs is a highly visible and frequently mentioned safety issue. Mostly, it is unfriendly, disrespectful, or ridiculous content that makes users uncomfortable and drives them away. It is extremely hazardous and could have negative social consequences."

    From Safety Assessment of Chinese Large Language Models (Sun2023)

  47. 27.01.03 · Risk Sub-Category

    Typical safety scenarios

    Crimes and Illegal Activities

    "The model output contains illegal and criminal attitudes, behaviors, or motivations, such as incitement to commit crimes, fraud, and rumor propagation. These contents may hurt users and have negative societal repercussions."

    From Safety Assessment of Chinese Large Language Models (Sun2023)

  48. 27.01.04 · Risk Sub-Category

    Typical safety scenarios

    Sensitive Topics

    "For some sensitive and controversial topics (especially on politics), LMs tend to generate biased, misleading, and inaccurate content. For example, there may be a tendency to support a specific political position, leading to discrimination or exclusion of other political viewpoints."

    From Safety Assessment of Chinese Large Language Models (Sun2023)

  49. 28.01.00 · Risk Category

    Offensiveness

    "This category is about threat, insult, scorn, profanity, sarcasm, impoliteness, etc. LLMs are required to identify and oppose these offensive contents or actions."

    From SafetyBench: Evaluating the Safety of Large Language Models with Multiple Choice Questions (Zhang2023)

  50. 30.02.00 · Risk Category

    Safety

    Avoiding unsafe and illegal outputs, and leaking private information

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.