MIT AI Risk Repository

Browse AI risks

494 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset

494 entries · page 1 of 10

  1. 28.02.00 · Risk Category

    Unfairness and Bias

    "This type of safety problem is mainly about social bias across various topics such as race, gender, religion, etc. LLMs are expected to identify and avoid unfair and biased expressions and actions."

    From SafetyBench: Evaluating the Safety of Large Language Models with Multiple Choice Questions (Zhang2023)

  2. 50.04.02 · Risk Sub-Category

    Legal and Rights-Related Risks

    Discrimination/Bias (Discriminatory Activities)

  3. 58.07.11 · Risk Sub-Category

    Societal and Cultural

    Stereotyping

    "Stereotyping - Derogatory or otherwise harmful stereotyping or homogenisation of individuals, groups, societies or cultures due to the mis-representation, over-representation, under-representation, or non- representation of specific identities, groups, or perspectives."

    From A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms (Abercrombie2024)

  4. Social bias is an unfairly negative attitude towards a social group or individuals based on one-sided or inaccurate information, typically pertaining to widely disseminated negative stereotypes regarding gender, race, religion, etc.

    From Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements (Deng2023)

  5. 13.02.02 · Risk Sub-Category

    Impacts: People and Society

    Inequality, Marginalization, and Violence

    "Generative AI systems are capable of exacerbating inequality, as seen in sections on 4.1.1 Bias, Stereotypes, and Representational Harms and 4.1.2 Cultural Values and Sensitive Content, and Disparate Performance. When deployed or updated, systems' impacts on people and groups can directly and indirectly be used to harm and exploit vulnerable and marginalized groups."

    From Evaluating the Social Impact of Generative AI Systems in Systems and Society (Solaiman2023)

  6. 18.02.02 · Risk Sub-Category

    Misinformation Harms

    Erosion of trust in public information

    "Eroding trust in public information and knowledge"

    From Sociotechnical Safety Evaluation of Generative AI Systems (Weidinger2023)

  7. 27.01.02 · Risk Sub-Category

    Typical safety scenarios

    Unfairness and discrinimation

    "The model produces unfair and discriminatory data, such as social bias based on race, gender, religion, appearance, etc. These contents may discomfort certain groups and undermine social stability and peace."

    From Safety Assessment of Chinese Large Language Models (Sun2023)

  8. 30.03.03 · Risk Sub-Category

    Fairness

    Preference Bias

    LLMs are exposed to vast groups of people, and their political biases may pose a risk of manipulation of socio-political processes

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  9. 43.01.02 · Risk Sub-Category

    Safety & Trustworthiness

    Bias

    7 types of bias evaluated: Demographical representation: These evaluations assess whether there is disparity in the rates at which different demographic groups are mentioned in LLM generated text. This ascertains over- representation, under-representation, or erasure of specific demographic groups; (2) Stereotype bias: These evaluations assess whether there is disparity in the rates at which different demographic groups are associated with stereotyped terms (e.g., occupations) in a LLM's generated output; (3) Fairness: These evaluations assess whether sensitive attributes (e.g., sex and race)

    From Cataloguing LLM Evaluations (InfoComm2023)

  10. 45.01.02 · Risk Sub-Category

    AI's inherent safety risks

    Risks from models and algorithms (Risks of bias and discrimination)

    "During the algorithm design and training process, personal biases may be introduced, either intentionally or unintentionally. Additionally, poor-quality datasets can lead to biased or discriminatory outcomes in the algorithm's design and outputs, including discriminatory content regarding ethnicity, religion, nationality, and region."

    From AI Safety Governance Framework (TC2602024)

  11. 50.02.09 · Risk Sub-Category

    Content Safety Risks

    Hate/Toxicity (Perpetuating Harmful Beliefs)

  12. 50.04.03 · Risk Sub-Category

    Legal and Rights-Related Risks

    Discrimination/Bias (Protected Characteristics)

  13. 58.06.03 · Risk Sub-Category

    Human rights and civil liberties

    Discrimination

    "Discrimination - Unfair or inadequate treatment or arbitrary distinction based on a person’s race, ethnicity, age, gender, sexual preference, religion, national origin, marital status, disability, language, or other protected groups."

    From A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms (Abercrombie2024)

  14. 61.01.03 · Risk Sub-Category

    Types of systemic risks from general-purpose AI

    Discrimination

    "The creation, perpetuation or exacerbation of inequalities and biases at a large-scale."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  15. 62.18.04 · Risk Sub-Category

    Model Evaluations (Interpretability/Explainability)

    Biases are not accurately reflected in explanations

    "Existing explainability techniques can be insufficient for detecting discriminatory biases. Manipulation methods can hide underlying biases from these tech- niques, generating misleading explanations [192, 112]. Such explanations ex- clude sensitive or prohibitive attributes, such as race or gender, and instead include desired attributes, even though they do not accurately represent the underlying model."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  16. 62.36.04 · Risk Sub-Category

    Impacts of AI (Bias)

    Systemic bias across specific communities

    "AI systems may exhibit unfair or unfavorable outputs across a range of tasks against specific communities of people, either implicitly or explicitly. Bias can lead to forms of exclusion or erasure (e.g., mislabelling for categorization-based tasks) and violence (e.g., sexual violence against women from deepfake pornog- raphy)."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  17. 66.06.04 · Risk Sub-Category

    Representation and Toxicity

    Cultural disposession

    "Intentional and/or unintentional erasure of cultural goods and values, such as ways of speaking, expressing humour, or sounds and voices that contribute to a cultural identity, or their inappropriate re-use in other cultures"

    From A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025)

  18. 66.10.02 · Risk Sub-Category

    Human Rights and Civil Liberties

    Benefits / entitlements loss

    "Denial of or loss of access to welfare benefits, pensions, housing, etc due to the malfunction, use or misuse of a technology system"

    From A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025)

  19. 69.06.02 · Risk Sub-Category

    Toxic and disrespectful content

    Discriminatory and exclusionary language

  20. This typically refers to rude, harmful, or inappropriate expressions.

    From Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements (Deng2023)

  21. 04.04.00 · Risk Category

    Controversial Opinions

    The controversial views expressed by large models are also a widely discussed concern. Bang et al. (2021) evaluated several large models and found that they occasionally express inappropriate or extremist views when discussing political top-ics. Furthermore, models like ChatGPT (OpenAI, 2022) that claim political neutrality and aim to provide objective information for users have been shown to exhibit notable left-leaning political biases in areas like economics, social policy, foreign affairs, and civil liberties.

    From Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements (Deng2023)

  22. 23.01.00 · Risk Category

    Violent crimes

    "This category addresses responses that enable, encourage, or endorse the commission of violent crimes."

    From Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024)

  23. 23.02.00 · Risk Category

    Non-violent crimes

    "This category addresses responses that enable, encourage, or endorse the commission of non-violent crimes."

    From Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024)

  24. 23.03.00 · Risk Category

    Sex-related crimes

    "This category addresses responses that enable, encourage, or endorse the commission of sex-related crimes"

    From Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024)

  25. "This category addresses responses that contain, describe, enable, encourage, or endorse the sexual abuse of children (i.e., under 18 years old) or the acquisition of child sexual abuse material."

    From Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024)

  26. 23.06.00 · Risk Category

    Suicide & self-harm

    "This category addresses responses that enable, encourage, or endorse acts of intentional self-harm."

    From Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024)

  27. 23.07.00 · Risk Category

    Hate

    "This category addresses responses that demean or dehumanize people on the basis of their sensitive, personal characteristics."

    From Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024)

  28. 23.13.00 · Risk Category

    Sexual Content

    "This category addresses responses that contain erotica. This includes depictions of sexual behavior, genitalia, and overtly sexualized body parts."

    From Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024)

  29. 27.01.01 · Risk Sub-Category

    Typical safety scenarios

    Insult

    "Insulting content generated by LMs is a highly visible and frequently mentioned safety issue. Mostly, it is unfriendly, disrespectful, or ridiculous content that makes users uncomfortable and drives them away. It is extremely hazardous and could have negative social consequences."

    From Safety Assessment of Chinese Large Language Models (Sun2023)

  30. 27.01.03 · Risk Sub-Category

    Typical safety scenarios

    Crimes and Illegal Activities

    "The model output contains illegal and criminal attitudes, behaviors, or motivations, such as incitement to commit crimes, fraud, and rumor propagation. These contents may hurt users and have negative societal repercussions."

    From Safety Assessment of Chinese Large Language Models (Sun2023)

  31. 27.01.04 · Risk Sub-Category

    Typical safety scenarios

    Sensitive Topics

    "For some sensitive and controversial topics (especially on politics), LMs tend to generate biased, misleading, and inaccurate content. For example, there may be a tendency to support a specific political position, leading to discrimination or exclusion of other political viewpoints."

    From Safety Assessment of Chinese Large Language Models (Sun2023)

  32. 28.01.00 · Risk Category

    Offensiveness

    "This category is about threat, insult, scorn, profanity, sarcasm, impoliteness, etc. LLMs are required to identify and oppose these offensive contents or actions."

    From SafetyBench: Evaluating the Safety of Large Language Models with Multiple Choice Questions (Zhang2023)

  33. 30.02.00 · Risk Category

    Safety

    Avoiding unsafe and illegal outputs, and leaking private information

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  34. 30.06.00 · Risk Category

    Social Norm

    LLMs are expected to reflect social values by avoiding the use of offensive language toward specific groups of users, being sensitive to topics that can create instability, as well as being sympathetic when users are seeking emotional support

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  35. 30.06.01 · Risk Sub-Category

    Social Norm

    Toxicity

    language being rude, disrespectful, threatening, or identity-attacking toward certain groups of the user population (culture, race, and gender etc)

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  36. 33.01.01 · Risk Sub-Category

    Ethical Concerns

    Harmful or inappropriate content

    "Harmful or inappropriate content produced by generative AI includes but is not limited to violent content, the use of offensive language, discriminative content, and pornography. Although OpenAI has set up a content policy for ChatGPT, harmful or inappropriate content can still appear due to reasons such as algorithmic limitations or jailbreaking (i.e., removal of restrictions imposed). The language models’ ability to understand or generate harmful or offensive content is referred to as toxicity (Zhuo et al., 2023). Toxicity can bring harm to society and damage the harmony of the community. H

    From Generative AI and ChatGPT: Applications, Challenges, and AI-Human Collaboration (Nah2023)

  37. 43.01.01 · Risk Sub-Category

    Safety & Trustworthiness

    Toxicity generation

    "These evaluations assess whether a LLM generates toxic text when prompted. In this context, toxicity is an umbrella term that encompasses hate speech, abusive language, violent speech, and profane language (Liang et al., 2022)."

    From Cataloguing LLM Evaluations (InfoComm2023)

  38. 43.02.14 · Risk Sub-Category

    Undesirable Use Cases

    Information on harmful, immoral, or illegal activity

    "These evaluations assess whether it is possible to solicit information on harmful, immoral or illegal activities from a LLM"

    From Cataloguing LLM Evaluations (InfoComm2023)

  39. 45.01.08 · Risk Sub-Category

    AI's inherent safety risks

    Risks from data (Risks of improper content and poisoning in training data)

    "If the training data includes illegal or harmful information, such as false, biased, or IPR-infringing content, or lacks diversity in its sources, the output may include harmful content like illegal, malicious, or extreme information. Training data is also at risk of being poisoned through tampering, error injection, or misleading actions by attackers. This can interfere with the model's probability distribution, reducing its accuracy and reliability."

    From AI Safety Governance Framework (TC2602024)

  40. 45.02.01 · Risk Sub-Category

    Safety risks in AI Applications

    Cyberspace risks (Risks of information and content safety)

    "AI-generated or synthesized content can lead to the spread of false information, discrimination and bias, privacy leakage, and infringement issues, threatening the safety of citizens' lives and property, national security, ideological security, and causing ethical risks. If users’ inputs contain harmful content, the model may output illegal or damaging information without robust security mechanisms."

    From AI Safety Governance Framework (TC2602024)

  41. "Eased production of and access to violent, inciting, radicalizing, or threatening content as well as recommendations to carry out self-harm or conduct illegal activities. Includes difficulty controlling public exposure to hateful and disparaging or stereotyping content."

    From Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST2024)

  42. 50.02.00 · Risk Category

    Content Safety Risks

    -

  43. 50.02.01 · Risk Sub-Category

    Content Safety Risks

    Violence and extremism (Supporting malicious organized groups)

  44. 50.02.02 · Risk Sub-Category

    Content Safety Risks

    Violence and extremism (Celebrating suffering)

  45. 50.02.03 · Risk Sub-Category

    Content Safety Risks

    Violence and extremism (Violent Acts)

  46. 50.02.08 · Risk Sub-Category

    Content Safety Risks

    Hate/Toxicity (Hate Speech: Inciting/Promoting/Expressing Hatred)

  47. 50.02.10 · Risk Sub-Category

    Content Safety Risks

    Hate/Toxicity (Offensive Language)

  48. 50.02.11 · Risk Sub-Category

    Content Safety Risks

    Sexual Content (Adult Content)

  49. 50.02.12 · Risk Sub-Category

    Content Safety Risks

    Sexual Content (Erotic)

  50. 50.02.13 · Risk Sub-Category

    Content Safety Risks

    Sexual Content (Non-Consensual Nudity)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.