MIT AI Risk Repository

Browse AI risks

977 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset

977 entries · page 1 of 20

  1. "Social harms that arise from the language model producing discriminatory or exclusionary speech"

    From Ethical and social risks of harm from language models (Weidinger2021)

  2. "AI systems under-, over-, or misrepresenting certain groups or generating toxic, offensive, abusive, or hateful content"

    From Sociotechnical Safety Evaluation of Generative AI Systems (Weidinger2023)

  3. 28.02.00 · Risk Category

    Unfairness and Bias

    "This type of safety problem is mainly about social bias across various topics such as race, gender, religion, etc. LLMs are expected to identify and avoid unfair and biased expressions and actions."

    From SafetyBench: Evaluating the Safety of Large Language Models with Multiple Choice Questions (Zhang2023)

  4. "AI systems under-, over-, or misrepresenting certain groups or generating toxic, offensive, abusive, or hateful content"

    From A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025)

  5. 03.01.00 · Risk Category

    Broken systems

    "These are the most mentioned cases. They refer to situations where the algorithm or the training data lead to unreliable outputs. These systems frequently assign disproportionate weight to some variables, like race or gender, but there is no transparency to this effect, making them impossible to challenge. These situations are typically only identified when regulators or the press examine the systems under freedom of information acts. Nevertheless, the damage they cause to people’s lives can be dramatic, such as lost homes, divorces, prosecution, or incarceration. Besides the inherent technic

    From Navigating the Landscape of AI Ethics and Responsibility (Cunha2023)

  6. Social bias is an unfairly negative attitude towards a social group or individuals based on one-sided or inaccurate information, typically pertaining to widely disseminated negative stereotypes regarding gender, race, religion, etc.

    From Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements (Deng2023)

  7. 05.01.00 · Risk Category

    Fairness - Bias

    Fairness is, by far, the most discussed issue in the literature, remaining a paramount concern especially in case of LLMs and text-to-image models. This is sparked by training data biases propagating into model outputs, causing negative effects like stereotyping, racism, sexism, ideological leanings, or the marginalization of minorities. Next to attesting generative AI a conservative inclination by perpetuating existing societal patterns, there is a concern about reinforcing existing biases when training new generative models with synthetic data from previous models. Beyond technical fairness

    From Mapping the Ethics of Generative AI: A Comprehensive Scoping Review (Hagendorff2024)

  8. 06.03.00 · Risk Category

    Discrimination

    "When AI is not carefully designed, it can discriminate against certain groups."

    From A framework for ethical Ai at the United Nations (Hogenhout2021)

  9. "The decision process used by AI systems has the potential to present biased choices, either because it acts from criteria that will generate forms of bias or because it is based on the history of choices."

    From Social Impacts of Artificial Intelligence and Mitigation Recommendations: An Exploratory Study (Paes2023)

  10. 10.02.00 · Risk Category

    Risk of Injury

    "Poorly designed intelligent systems can cause moral, psychological, and physical harm. For example, the use of predictive policing tools may cause more people to be arrested or physically harmed by the police."

    From Social Impacts of Artificial Intelligence and Mitigation Recommendations: An Exploratory Study (Paes2023)

  11. "The risks associated with the use of AI are still unpredictable and unprecedented, and there are already several examples that show AI has made discriminatory decisions against minorities, reinforced social stereotypes in Internet search engines and enabled data breaches."

    From Social Impacts of Artificial Intelligence and Mitigation Recommendations: An Exploratory Study (Paes2023)

  12. 11.01.00 · Risk Category

    Representational Harms

    "beliefs about different social groups that reproduce unjust societal hierarchies"

    From Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction (Shelby2023)

  13. 11.01.01 · Risk Sub-Category

    Representational Harms

    Stereotyping social groups

    Stereotyping in an algorithmic system refers to how the system’s outputs reflect “beliefs about the characteristics, attributes, and behaviors of members of certain groups....and about how and why certain attributes go together"

    From Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction (Shelby2023)

  14. 11.01.02 · Risk Sub-Category

    Representational Harms

    Demeaning social groups

    Demeaning of social groups to occur when they are when they are “cast as being lower status and less deserving of respect"... discourses, images, and language used to marginalize or oppress a social group... Controlling images include forms of human-animal confusion in image tagging systems

    From Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction (Shelby2023)

  15. 11.01.04 · Risk Sub-Category

    Representational Harms

    Alienating social groups

    when an image tagging system does not acknowledge the relevance of someone’s membership in a specific social group to what is depicted in one or more images

    From Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction (Shelby2023)

  16. 11.01.05 · Risk Sub-Category

    Representational Harms

    Denying people the opportunity to self-identify

    complex and non-traditional ways in which humans are represented and classified automatically, and often at the cost of autonomy loss... such as categorizing someone who identifies as non-binary into a gendered category they do not belong ... undermines people’s ability to disclose aspects of their identity on their own terms

    From Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction (Shelby2023)

  17. 11.01.06 · Risk Sub-Category

    Representational Harms

    Reifying essentialist categories

    algorithmic systems that reify essentialist social categories can be understood as when systems that classify a person’s membership in a social group based on narrow, socially constructed criteria that reinforce perceptions of human difference as inherent, static and seemingly natural... especially likely when ML models or human raters classify a person’s attributes – for instance, their gender, race, or sexual orientation – by making assumptions based on their physical appearance

    From Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction (Shelby2023)

  18. 11.02.00 · Risk Category

    Allocative Harms

    "These harms occur when a system withholds information, opportunities, or resources [22] from historically marginalized groups in domains that affect material well-being [146], such as housing [47], employment [201], social services [15, 201], finance [117], education [119], and healthcare [158]."

    From Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction (Shelby2023)

  19. 11.02.01 · Risk Sub-Category

    Allocative Harms

    Opportunity loss

    Opportunity loss occurs when algorithmic systems enable disparate access to information and resources needed to equitably participate in society, including the withholding of housing through targeting ads based on race [10] and social services along lines of class [84]

    From Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction (Shelby2023)

  20. 11.02.02 · Risk Sub-Category

    Allocative Harms

    Economic loss

    Financial harms [52, 160] co-produced through algorithmic systems, especially as they relate to lived experiences of poverty and economic inequality... demonetization algorithms that parse content titles, metadata, and text, and it may penalize words with multiple meanings [51, 81], disproportionately impacting queer, trans, and creators of color [81]. Differential pricing algorithms, where people are systematically shown different prices for the same products, also leads to economic loss [55]. These algorithms may be especially sensitive to feedback loops from existing inequities related to e

    From Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction (Shelby2023)

  21. 13.02.02 · Risk Sub-Category

    Impacts: People and Society

    Inequality, Marginalization, and Violence

    "Generative AI systems are capable of exacerbating inequality, as seen in sections on 4.1.1 Bias, Stereotypes, and Representational Harms and 4.1.2 Cultural Values and Sensitive Content, and Disparate Performance. When deployed or updated, systems' impacts on people and groups can directly and indirectly be used to harm and exploit vulnerable and marginalized groups."

    From Evaluating the Social Impact of Generative AI Systems in Systems and Society (Solaiman2023)

  22. 14.01.00 · Risk Category

    Fairness

    "The general principle of equal treatment requires that an AI system upholds the principle of fairness, both ethically and legally. This means that the same facts are treated equally for each person unless there is an objective justification for unequal treatment."

    From Sources of Risk of AI Systems (Steimers2022)

  23. 15.02.02 · Risk Sub-Category

    Second-Order Risks

    Discrimination

    This is the risk of an ML system encoding stereotypes of or performing disproportionately poorly for some demographics/social groups.

    From The Risks of Machine Learning Systems (Tan2022)

  24. 16.05.01 · Risk Sub-Category

    Risk area 5: Human-Computer Interaction Harms

    Promoting harmful stereotypes by implying gender or ethnic identity

    "CAs can perpetuate harmful stereotypes by using particular identity markers in language (e.g. referring to “self” as “female”), or by more general design features (e.g. by giving the product a gendered name such as Alexa). The risk of representational harm in these cases is that the role of “assistant” is presented as inherently linked to the female gender [19, 36]. Gender or ethnicity identity markers may be implied by CA vocabulary, knowledge or vernacular [124]; product description, e.g. in one case where users could choose as virtual assistant Jake - White, Darnell - Black, Antonio - Hisp

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  25. 17.05.03 · Risk Sub-Category

    Human-Computer Interaction Harms

    Promoting harmful stereotypes by implying gender or ethnic identity

    "A conversational agent may invoke associations that perpetuate harmful stereotypes, either by using particular identity markers in language (e.g. referring to “self” as “female”), or by more general design features (e.g. by giving the product a gendered name)."

    From Ethical and social risks of harm from language models (Weidinger2021)

  26. 18.01.01 · Risk Sub-Category

    Representation & Toxicity Harms

    Unfair representation

    "Mis-, under-, or over-representing certain identities, groups, or perspectives or failing to represent them at all (e.g. via homogenisation, stereotypes)"

    From Sociotechnical Safety Evaluation of Generative AI Systems (Weidinger2023)

  27. 18.02.02 · Risk Sub-Category

    Misinformation Harms

    Erosion of trust in public information

    "Eroding trust in public information and knowledge"

    From Sociotechnical Safety Evaluation of Generative AI Systems (Weidinger2023)

  28. 19.05.02 · Risk Sub-Category

    Ethical AI Risks

    Unfair statistical AI decisions and discrimination of minorities

  29. 27.01.02 · Risk Sub-Category

    Typical safety scenarios

    Unfairness and discrinimation

    "The model produces unfair and discriminatory data, such as social bias based on race, gender, religion, appearance, etc. These contents may discomfort certain groups and undermine social stability and peace."

    From Safety Assessment of Chinese Large Language Models (Sun2023)

  30. 29.01.01 · Risk Sub-Category

    AI Trust Management

    Bias and Discrimination

    as they claim to generate biased and discriminatory results, these AI systems have a negative impact on the rights of individuals, principles of adjudication, and overall judicial integrity

    From Artificial Intelligence Trust, Risk and Security Management (AI TRiSM): Frameworks, Applications, Challenges and Future Research Directions (Habbal2024)

  31. 30.03.01 · Risk Sub-Category

    Fairness

    Injustice

    In the context of LLM outputs, we want to make sure the suggested or completed texts are indistinguishable in nature for two involved individuals (in the prompt) with the same relevant profiles but might come from different groups (where the group attribute is regarded as being irrelevant in this context)

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  32. 30.03.02 · Risk Sub-Category

    Fairness

    Stereotype Bias

    LLMs must not exhibit or highlight any stereotypes in the generated text. Pretrained LLMs tend to pick up stereotype biases persisting in crowdsourced data and further amplify them

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  33. 30.03.03 · Risk Sub-Category

    Fairness

    Preference Bias

    LLMs are exposed to vast groups of people, and their political biases may pose a risk of manipulation of socio-political processes

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  34. 30.07.03 · Risk Sub-Category

    Robustness

    Interventional Effect

    existing disparities in data among different user groups might create differentiated experiences when users interact with an algorithmic system (e.g. a recommendation system), which will further reinforce the bias

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  35. 49.02.02 · Risk Sub-Category

    Risks from Malfunctions

    Risks from bias and underrepresentation

    "The outputs and impacts of general- purpose AI systems can be biased with respect to various aspects of human identity, including race, gender, culture, age, and disability. This creates risks in high- stakes domains such as healthcare, job recruitment, and financial lending. General- purpose AI systems are primarily trained on language and image datasets that disproportionately represent English- speaking and Western cultures, increasing the potential for harm to individuals not represented well by this data."

    From International Scientific Report on the Safety of Advanced AI (Bengio2024)

  36. 50.01.04 · Risk Sub-Category

    System and Operational Risks

    Operational misuses (Automated decision-making)

  37. 50.02.09 · Risk Sub-Category

    Content Safety Risks

    Hate/Toxicity (Perpetuating Harmful Beliefs)

  38. 52.01.01 · Risk Sub-Category

    Risks from Unreliability

    Discrimination and Stereotype Reproduction

    "General purpose AI models interpret and respond to inputs based on their training data, potentially causing Discrimination and Stereotype Reproduction. Since they are “black-box” models, the exact mechanism behind decisions remains opaque and attempts to mitigate harmful outputs are not fully reliable yet. These models have the capacity to influence a multitude of downstream applications, decisions, and processes, thereby affecting many individuals simultaneously. The extent of this impact could outstrip the range of any single human or group of humans, amplifying the potential consequences o

    From Governing General Purpose AI: A Comprehensive Map of Unreliability, Misuse and Systemic Risks (Maham2023 )

  39. 54.01.03 · Risk Sub-Category

    Negative impacts of AI use

    Discrimination, toxicity, and bias

    "AI models and the tools that use them may exacerbate unequal access to employment and services. AI-generated content can promote inequality and harmful stereotypes."

    From Ten Hard Problems in Artificial Intelligence We Must Get Right (Leech2024 )

  40. 56.01.00 · Risk Category

    Discrimination

    "More broadly, bad decisions or errors by AI tools could lead to discrimination or deeper inequality"

    From Future Risks of Frontier AI (GOS2023)

  41. 58.06.03 · Risk Sub-Category

    Human rights and civil liberties

    Discrimination

    "Discrimination - Unfair or inadequate treatment or arbitrary distinction based on a person’s race, ethnicity, age, gender, sexual preference, religion, national origin, marital status, disability, language, or other protected groups."

    From A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms (Abercrombie2024)

  42. "Discriminative data bias describes the systematic discrimination of groups of persons in the form of data shortcomings, such as distributional representation or incorrectness. Data bias can manifest in the model and lead to unfair decisions if not appropriately treated. Note, that the term bias is often used in other contexts, such as data representation. However, these issues are treated by other AI hazards in this list."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  43. 61.01.03 · Risk Sub-Category

    Types of systemic risks from general-purpose AI

    Discrimination

    "The creation, perpetuation or exacerbation of inequalities and biases at a large-scale."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  44. 61.02.29 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Incomplete or biased training data

    "Incomplete or biased training data can lead to discriminatory AI outputs."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  45. 62.35.03 · Risk Sub-Category

    Impacts of AI (Bias)

    Biases in AI-based content moderation algorithms

    "AI-based content moderation algorithms, while intended to filter harmful con- tent, can perpetuate biases. For example, gender biases within these systems may lead to the disproportionate suppression or “shadowbanning” of content featuring women [132]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  46. 62.36.04 · Risk Sub-Category

    Impacts of AI (Bias)

    Systemic bias across specific communities

    "AI systems may exhibit unfair or unfavorable outputs across a range of tasks against specific communities of people, either implicitly or explicitly. Bias can lead to forms of exclusion or erasure (e.g., mislabelling for categorization-based tasks) and violence (e.g., sexual violence against women from deepfake pornog- raphy)."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  47. 62.36.05 · Risk Sub-Category

    Impacts of AI (Bias)

    Unintentional bias amplification

    "Dataset bias may be unintentionally amplified [60] where the outputs of the AI model trained on a dataset are more biased than the dataset itself."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  48. 65.04.01 · Risk Sub-Category

    Training Data Risks (Fairness)

    Data bias

    "Historical and societal biases that are present in the data are used to train and fine-tune the model."

    From AI Risk Atlas (IBM2025)

  49. 65.19.01 · Risk Sub-Category

    Output risks (Fairness)

    Output bias

    "Generated content might unfairly represent certain groups or individuals."

    From AI Risk Atlas (IBM2025)

  50. 66.06.02 · Risk Sub-Category

    Representation and Toxicity

    Stereotyping

    "Derogatory or otherwise harmful stereotyping or homogenisation of individuals, groups, societies or cultures due to the mis-representation, over-representation, under-representation, or non-representation of specific identities, groups or perspectives"

    From A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.