MIT AI Risk Repository

Browse AI risks

662 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset

662 entries · page 2 of 14

  1. 56.01.00 · Risk Category

    Discrimination

    "More broadly, bad decisions or errors by AI tools could lead to discrimination or deeper inequality"

    From Future Risks of Frontier AI (GOS2023)

  2. "Discriminative data bias describes the systematic discrimination of groups of persons in the form of data shortcomings, such as distributional representation or incorrectness. Data bias can manifest in the model and lead to unfair decisions if not appropriately treated. Note, that the term bias is often used in other contexts, such as data representation. However, these issues are treated by other AI hazards in this list."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  3. 62.35.03 · Risk Sub-Category

    Impacts of AI (Bias)

    Biases in AI-based content moderation algorithms

    "AI-based content moderation algorithms, while intended to filter harmful con- tent, can perpetuate biases. For example, gender biases within these systems may lead to the disproportionate suppression or “shadowbanning” of content featuring women [132]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  4. 62.36.04 · Risk Sub-Category

    Impacts of AI (Bias)

    Systemic bias across specific communities

    "AI systems may exhibit unfair or unfavorable outputs across a range of tasks against specific communities of people, either implicitly or explicitly. Bias can lead to forms of exclusion or erasure (e.g., mislabelling for categorization-based tasks) and violence (e.g., sexual violence against women from deepfake pornog- raphy)."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  5. 62.36.05 · Risk Sub-Category

    Impacts of AI (Bias)

    Unintentional bias amplification

    "Dataset bias may be unintentionally amplified [60] where the outputs of the AI model trained on a dataset are more biased than the dataset itself."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  6. 65.19.01 · Risk Sub-Category

    Output risks (Fairness)

    Output bias

    "Generated content might unfairly represent certain groups or individuals."

    From AI Risk Atlas (IBM2025)

  7. 65.19.02 · Risk Sub-Category

    Output risks (Fairness)

    Decision bias

    "Decision bias occurs when one group is unfairly advantaged over another due to decisions of the model. This might be caused by biases in the data and also amplified as a result of the model’s training."

    From AI Risk Atlas (IBM2025)

  8. 66.06.02 · Risk Sub-Category

    Representation and Toxicity

    Stereotyping

    "Derogatory or otherwise harmful stereotyping or homogenisation of individuals, groups, societies or cultures due to the mis-representation, over-representation, under-representation, or non-representation of specific identities, groups or perspectives"

    From A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025)

  9. "Frontier AI models can contain and magnify biases ingrained in the data they are trained on, reflecting societal and historical inequalities and stereotypes.177 These biases, often subtle and deeply embedded, compromise the equitable and ethical use of AI systems, making it difficult for AI to improve fairness in decisions.178 Removing attributes like race and gender from training data has generally proven ineffective as a remedy for algorithmic bias, as models can infer these attributes from other information such as names, locations, and other seemingly unrelated factors."

    From Capabilities and Risks from Frontier AI (DSIT2023)

  10. 69.06.02 · Risk Sub-Category

    Toxic and disrespectful content

    Discriminatory and exclusionary language

  11. "The chatbot gives information that, while not obviously false or harmful, could lead to biased decision-making."

    From Emerging Risks and Mitigations for Public Chatbots: LILAC v1 (Stanley2024)

  12. 70.04.01 · Risk Sub-Category

    Social Risks

    Bias and discrimination

    "Like virtual applications of AI, EAI can display bias towards and dis- criminate against users. When EAI systems are placed in positions of power, their biases could have significant impacts on fairness in everyday interactions and on general social dynamics [105, 106]."

    From Embodied AI: Emerging Risks and Opportunities for Policy Action (Perlo2025)

  13. 73.04.01 · Risk Sub-Category

    LLM-Systems Can Be Untrustworthy

    Harms of Representation and Other Biases

    "A pretrained LLM generally has many of the stereotypical biases commonly present in the human society (Touvron et al., 2023). This makes it difficult for users to trust that LLMs will work well for them and not produce unfair or biased responses. Appropriate finetuning can effectively limit the bias displayed in LLM outputs in a variety of situations, e.g. when models are explicitly prompted with stereotypes (Wang et al., 2023k), but it does not ‘solve’ the problem. Even after finetuning, biases often resurface when deliberately elicited (Wang et al., 2023k), or under novel scenarios, e.g. in

    From Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  14. 02.01.00 · Risk Category

    Harmful Content

    "The LLM-generated content sometimes contains biased, toxic, and private information"

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  15. 02.01.02 · Risk Sub-Category

    Harmful Content

    Toxicity

    "Toxicity means the generated content contains rude, disrespectful, and even illegal information"

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  16. 02.08.01 · Risk Sub-Category

    Toxicity and Bias Tendencies

    Toxic Training Data

    "Following previous studies [96], [97], toxic data in LLMs is defined as rude, disrespectful, or unreasonable language that is opposite to a polite, positive, and healthy language environment, including hate speech, offensive utterance, profanities, and threats [91]."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  17. 04.04.00 · Risk Category

    Controversial Opinions

    The controversial views expressed by large models are also a widely discussed concern. Bang et al. (2021) evaluated several large models and found that they occasionally express inappropriate or extremist views when discussing political top-ics. Furthermore, models like ChatGPT (OpenAI, 2022) that claim political neutrality and aim to provide objective information for users have been shown to exhibit notable left-leaning political biases in areas like economics, social policy, foreign affairs, and civil liberties.

    From Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements (Deng2023)

  18. 13.01.02 · Risk Sub-Category

    Impacts: The Technical Base System

    Cultural Values and Sensitive Content

    "Cultural values are specific to groups and sensitive content is normative. Sensitive topics also vary by culture and can include hate speech, which itself is contingent on cultural norms of acceptability."

    From Evaluating the Social Impact of Generative AI Systems in Systems and Society (Solaiman2023)

  19. "Speech can create a range of harms, such as promoting social stereotypes that perpetuate the derogatory representation or unfair treatment of marginalised groups [22], inciting hate or violence [57], causing profound offence [199], or reinforcing social norms that exclude or marginalise identities [15,58]. LMs that faithfully mirror harmful language present in the training data can reproduce these harms. Unfair treatment can also emerge from LMs that perform better for some social groups than others [18]. These risks have been widely known, observed and documented in LMs. Mitigation approache

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  20. 16.01.02 · Risk Sub-Category

    Risk area 1: Discrimination, Hate speech and Exclusion

    Hate speech and offensive language

    "LMs may generate language that includes profanities, identity attacks, insults, threats, language that incites violence, or language that causes justified offence as such language is prominent online [57, 64, 143,191]. This language risks causing offence, psychological harm, and inciting hate or violence."

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  21. 17.01.03 · Risk Sub-Category

    Discrimination, Exclusion and Toxicity

    Toxic language

    "LM’s may predict hate speech or other language that is “toxic”. While there is no single agreed definition of what constitutes hate speech or toxic speech (Fortuna and Nunes, 2018; Persily and Tucker, 2020; Schmidt and Wiegand, 2017), proposed definitions often include profanities, identity attacks, sleights, insults, threats, sexually explicit content, demeaning language, language that incites violence, or ‘hostile and malicious language targeted at a person or group because of their actual or perceived innate characteristics’ (Fortuna and Nunes, 2018; Gorwa et al., 2020; PerspectiveAPI)"

    From Ethical and social risks of harm from language models (Weidinger2021)

  22. 18.01.03 · Risk Sub-Category

    Representation & Toxicity Harms

    Toxic content

    "Generating content that violates community standards, including harming or inciting hatred or violence against individuals and groups (e.g. gore, child sexual abuse material, profanities, identity attacks)"

    From Sociotechnical Safety Evaluation of Generative AI Systems (Weidinger2023)

  23. 23.01.00 · Risk Category

    Violent crimes

    "This category addresses responses that enable, encourage, or endorse the commission of violent crimes."

    From Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024)

  24. 23.02.00 · Risk Category

    Non-violent crimes

    "This category addresses responses that enable, encourage, or endorse the commission of non-violent crimes."

    From Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024)

  25. 23.03.00 · Risk Category

    Sex-related crimes

    "This category addresses responses that enable, encourage, or endorse the commission of sex-related crimes"

    From Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024)

  26. "This category addresses responses that contain, describe, enable, encourage, or endorse the sexual abuse of children (i.e., under 18 years old) or the acquisition of child sexual abuse material."

    From Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024)

  27. 23.06.00 · Risk Category

    Suicide & self-harm

    "This category addresses responses that enable, encourage, or endorse acts of intentional self-harm."

    From Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024)

  28. 23.07.00 · Risk Category

    Hate

    "This category addresses responses that demean or dehumanize people on the basis of their sensitive, personal characteristics."

    From Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024)

  29. 23.13.00 · Risk Category

    Sexual Content

    "This category addresses responses that contain erotica. This includes depictions of sexual behavior, genitalia, and overtly sexualized body parts."

    From Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024)

  30. 24.08.02 · Risk Sub-Category

    Privacy

    Violation of social norms

    "Second, because LLMs are trained on internet text data, there is also a risk that model weights encode functions which, if deployed in particular contexts, would violate social norms of that context. Following the principles of contextual integrity, it may be that models deviate from information sharing norms as a result of their training. Overcoming this challenge requires two types of infrastructure: one for keeping track of social norms in context, and another for ensuring that models adhere to them. Keeping track of what social norms are presently at play is an active research area. Surfa

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  31. 27.01.01 · Risk Sub-Category

    Typical safety scenarios

    Insult

    "Insulting content generated by LMs is a highly visible and frequently mentioned safety issue. Mostly, it is unfriendly, disrespectful, or ridiculous content that makes users uncomfortable and drives them away. It is extremely hazardous and could have negative social consequences."

    From Safety Assessment of Chinese Large Language Models (Sun2023)

  32. 27.01.03 · Risk Sub-Category

    Typical safety scenarios

    Crimes and Illegal Activities

    "The model output contains illegal and criminal attitudes, behaviors, or motivations, such as incitement to commit crimes, fraud, and rumor propagation. These contents may hurt users and have negative societal repercussions."

    From Safety Assessment of Chinese Large Language Models (Sun2023)

  33. 27.01.04 · Risk Sub-Category

    Typical safety scenarios

    Sensitive Topics

    "For some sensitive and controversial topics (especially on politics), LMs tend to generate biased, misleading, and inaccurate content. For example, there may be a tendency to support a specific political position, leading to discrimination or exclusion of other political viewpoints."

    From Safety Assessment of Chinese Large Language Models (Sun2023)

  34. 28.01.00 · Risk Category

    Offensiveness

    "This category is about threat, insult, scorn, profanity, sarcasm, impoliteness, etc. LLMs are required to identify and oppose these offensive contents or actions."

    From SafetyBench: Evaluating the Safety of Large Language Models with Multiple Choice Questions (Zhang2023)

  35. 30.02.00 · Risk Category

    Safety

    Avoiding unsafe and illegal outputs, and leaking private information

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  36. 30.02.01 · Risk Sub-Category

    Safety

    Violence

    LLMs are found to generate answers that contain violent content or generate content that responds to questions that solicit information about violent behaviors

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  37. 30.02.02 · Risk Sub-Category

    Safety

    Unlawful Conduct

    LLMs have been shown to be a convenient tool for soliciting advice on accessing, purchasing (illegally), and creating illegal substances, as well as for dangerous use of them

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  38. 30.02.03 · Risk Sub-Category

    Safety

    Harms to Minor

    LLMs can be leveraged to solicit answers that contain harmful content to children and youth

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  39. 30.02.04 · Risk Sub-Category

    Safety

    Adult Content

    LLMs have the capability to generate sex-explicit conversations, and erotic texts, and to recommend websites with sexual content

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  40. 30.06.00 · Risk Category

    Social Norm

    LLMs are expected to reflect social values by avoiding the use of offensive language toward specific groups of users, being sensitive to topics that can create instability, as well as being sympathetic when users are seeking emotional support

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  41. 30.06.01 · Risk Sub-Category

    Social Norm

    Toxicity

    language being rude, disrespectful, threatening, or identity-attacking toward certain groups of the user population (culture, race, and gender etc)

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  42. 33.01.01 · Risk Sub-Category

    Ethical Concerns

    Harmful or inappropriate content

    "Harmful or inappropriate content produced by generative AI includes but is not limited to violent content, the use of offensive language, discriminative content, and pornography. Although OpenAI has set up a content policy for ChatGPT, harmful or inappropriate content can still appear due to reasons such as algorithmic limitations or jailbreaking (i.e., removal of restrictions imposed). The language models’ ability to understand or generate harmful or offensive content is referred to as toxicity (Zhuo et al., 2023). Toxicity can bring harm to society and damage the harmony of the community. H

    From Generative AI and ChatGPT: Applications, Challenges, and AI-Human Collaboration (Nah2023)

  43. 43.01.01 · Risk Sub-Category

    Safety & Trustworthiness

    Toxicity generation

    "These evaluations assess whether a LLM generates toxic text when prompted. In this context, toxicity is an umbrella term that encompasses hate speech, abusive language, violent speech, and profane language (Liang et al., 2022)."

    From Cataloguing LLM Evaluations (InfoComm2023)

  44. 43.02.14 · Risk Sub-Category

    Undesirable Use Cases

    Information on harmful, immoral, or illegal activity

    "These evaluations assess whether it is possible to solicit information on harmful, immoral or illegal activities from a LLM"

    From Cataloguing LLM Evaluations (InfoComm2023)

  45. "Eased production of and access to violent, inciting, radicalizing, or threatening content as well as recommendations to carry out self-harm or conduct illegal activities. Includes difficulty controlling public exposure to hateful and disparaging or stereotyping content."

    From Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST2024)

  46. 50.02.01 · Risk Sub-Category

    Content Safety Risks

    Violence and extremism (Supporting malicious organized groups)

  47. 50.02.02 · Risk Sub-Category

    Content Safety Risks

    Violence and extremism (Celebrating suffering)

  48. 50.02.03 · Risk Sub-Category

    Content Safety Risks

    Violence and extremism (Violent Acts)

  49. 50.02.04 · Risk Sub-Category

    Content Safety Risks

    Violence and extremism (Depicting violence)

  50. 50.02.08 · Risk Sub-Category

    Content Safety Risks

    Hate/Toxicity (Hate Speech: Inciting/Promoting/Expressing Hatred)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.