MIT AI Risk Repository · domain 1: Discrimination & Toxicity

1.1 Unfair discrimination and misrepresentation

Unequal treatment of individuals or groups by AI, often based on race, gender, or other sensitive characteristics, resulting in unfair outcomes and representation of those groups.

Risk entries
83
Frameworks citing it
12
Recorded incidents
118
Incidents since 2020
58
Causal entity (risk entries)
Causal entity (risk entries) 58 0 AI: 58 AI 58 Other: 13 Other 13 Human: 11 Human 11 Not coded: 1 Not coded 1
Causal entity (risk entries)
LabelValue
AI58
Other13
Human11
Not coded1
Intent (risk entries)
Intent (risk entries) 64 0 Unintentional: 64 Unintentional 64 Other: 16 Other 16 Intentional: 2 Intentional 2 Not coded: 1 Not coded 1
Intent (risk entries)
LabelValue
Unintentional64
Other16
Intentional2
Not coded1
Timing (risk entries)
Timing (risk entries) 50 0 Post-deployment: 50 Post-deployment 50 Other: 19 Other 19 Pre-deployment: 13 Pre-deployment 13 Not coded: 1 Not coded 1
Timing (risk entries)
LabelValue
Post-deployment50
Other19
Pre-deployment13
Not coded1
Recorded incidents per yearIncident date; current year partial
Recorded incidents per year 19 0 2012: 5 2012 5 2013: 2 2013 2 2014: 2 2014 2 2015: 6 2015 6 2016: 13 2016 13 2017: 9 2017 9 2018: 9 2018 9 2019: 8 2019 8 2020: 19 2020 19 2021: 8 2021 8 2022: 10 2022 10 2023: 11 2023 11 2024: 7 2024 7 2025: 3 2025 3
Recorded incidents per year
LabelValue
20125
20132
20142
20156
201613
20179
20189
20198
202019
20218
202210
202311
20247
20253
Entries by levelRisk categories, subcategories and additional evidence coded to this subdomain
Entries by level 61 0 Risk Category: 22 Risk Category 22 Risk Sub-Category: 61 Risk Sub-Category 61
Entries by level
LabelValue
Risk Category22
Risk Sub-Category61
  • Discrimination

    "Discrimination - Unfair or inadequate treatment or arbitrary distinction based on a person’s race, ethnicity, age, gender, sexual preference, religion, national origin, marital status, disability, la...

    A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms (Abercrombie2024) · Other · Other · Post-deployment

  • Harms of Representation and Other Biases

    "A pretrained LLM generally has many of the stereotypical biases commonly present in the human society (Touvron et al., 2023). This makes it difficult for users to trust that LLMs will work well for t...

    Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024) · AI · Unintentional · Post-deployment

  • Risks from bias and underrepresentation

    "The outputs and impacts of general- purpose AI systems can be biased with respect to various aspects of human identity, including race, gender, culture, age, and disability. This creates risks in hig...

    International Scientific Report on the Safety of Advanced AI (Bengio2024) · AI · Unintentional · Post-deployment

  • Bias

    "General-purpose AI systems can amplify social and political biases, causing concrete harm. They frequently display biases with respect to race, gender, culture, age, disability, political opinion, or...

    International AI Safety Report 2025 (Bengio2025) · Human · Unintentional · Pre-deployment

  • Bias

    "The training datasets of LLMs may contain biased information that leads LLMs to generate outputs with social biases"

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024) · AI · Unintentional · Other

  • Toxicity and Bias Tendencies

    "Extensive data collection in LLMs brings toxic content and stereotypical bias into the training data."

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024) · Human · Unintentional · Pre-deployment

  • Biased Training Data

    "Compared with the definition of toxicity, the definition of bias is more subjective and contextdependent. Based on previous work [97], [101], we describe the bias as disparities that could raise demo...

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024) · AI · Unintentional · Pre-deployment

  • Broken systems

    "These are the most mentioned cases. They refer to situations where the algorithm or the training data lead to unreliable outputs. These systems frequently assign disproportionate weight to some varia...

    Navigating the Landscape of AI Ethics and Responsibility (Cunha2023) · AI · Unintentional · Post-deployment

  • Unfairness and Discrimination

    Social bias is an unfairly negative attitude towards a social group or individuals based on one-sided or inaccurate information, typically pertaining to widely disseminated negative stereotypes regard...

    Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements (Deng2023) · Other · Other · Post-deployment

  • Bias, Fairness and Representational Harms

    "Frontier AI models can contain and magnify biases ingrained in the data they are trained on, reflecting societal and historical inequalities and stereotypes.177 These biases, often subtle and deeply...

    Capabilities and Risks from Frontier AI (DSIT2023) · AI · Unintentional · Other

  • Algorithm and data

    "More than 20% of the contributions are centered on the ethical dimensions of algorithms and data. This theme can be further categorized into two main subthemes: data bias and algorithm fairness, and...

    What Ethics Can Say on Artificial Intelligence: Insights from a Systematic Literature Review (Giarmoleo2024) · Human · Intentional · Pre-deployment

  • Discrimination and bias

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Not coded · Not coded · Not coded

  • Biases are not accurately reflected in explanations

    "Existing explainability techniques can be insufficient for detecting discriminatory biases. Manipulation methods can hide underlying biases from these tech- niques, generating misleading explanations...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Other · Other · Other

  • Biases in AI-based content moderation algorithms

    "AI-based content moderation algorithms, while intended to filter harmful con- tent, can perpetuate biases. For example, gender biases within these systems may lead to the disproportionate suppression...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · AI · Unintentional · Post-deployment

  • Systemic bias across specific communities

    "AI systems may exhibit unfair or unfavorable outputs across a range of tasks against specific communities of people, either implicitly or explicitly. Bias can lead to forms of exclusion or erasure (e...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · AI · Other · Post-deployment

  • Unintentional bias amplification

    "Dataset bias may be unintentionally amplified [60] where the outputs of the AI model trained on a dataset are more biased than the dataset itself."

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · AI · Unintentional · Post-deployment

  • Discrimination

    "More broadly, bad decisions or errors by AI tools could lead to discrimination or deeper inequality"

    Future Risks of Frontier AI (GOS2023) · AI · Unintentional · Post-deployment

  • Amplification of biases

    "Current Frontier AI mdoels amplify existing biases within their training data and can be manipulated into providing potentially harmful responses, for example abusive language or discriminatory respo...

    Future Risks of Frontier AI (GOS2023) · Human · Unintentional · Pre-deployment

  • Bias and discrimination (bias in training datasets)

    "AI experts consider training data to be the most salient source of bias in generative AI models. For example, GPT- 2’s training data comes from outbound links from Reddit, a social network often crit...

    Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024) · Other · Unintentional · Pre-deployment

  • Bias and discrimination (value lock and outcome homogenization)

    "Because models are not necessarily retrained to reflect evolving societal views, language models risk “value lock- ins,” which “reifies older, less inclusive understandings.”370 Therefore, the contin...

    Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024) · Human · Unintentional · Other

  • Bias and Discrimination

    as they claim to generate biased and discriminatory results, these AI systems have a negative impact on the rights of individuals, principles of adjudication, and overall judicial integrity

    Artificial Intelligence Trust, Risk and Security Management (AI TRiSM): Frameworks, Applications, Challenges and Future Research Directions (Habbal2024) · AI · Unintentional · Post-deployment

  • Fairness - Bias

    Fairness is, by far, the most discussed issue in the literature, remaining a paramount concern especially in case of LLMs and text-to-image models. This is sparked by training data biases propagating...

    Mapping the Ethics of Generative AI: A Comprehensive Scoping Review (Hagendorff2024) · AI · Unintentional · Post-deployment

  • Discrimination

    "When AI is not carefully designed, it can discriminate against certain groups."

    A framework for ethical Ai at the United Nations (Hogenhout2021) · AI · Unintentional · Post-deployment

  • Bias

    "The AI will only be as good as the data it is trained with. If the data contains bias (and much data does), then the AI will manifest that bias, too."

    A framework for ethical Ai at the United Nations (Hogenhout2021) · AI · Unintentional · Pre-deployment

  • Data bias

    "Historical and societal biases that are present in the data are used to train and fine-tune the model."

    AI Risk Atlas (IBM2025) · Human · Unintentional · Post-deployment