MIT AI Risk Repository · Risk Sub-Category · 47.02.08

Bias and discrimination (bias in training datasets)

Category: Ethical and social risks

Description

"AI experts consider training data to be the most salient source of bias in generative AI models. For example, GPT- 2’s training data comes from outbound links from Reddit, a social network often criticized for hosting anti-feminist content.351 As a result, AI models trained on such data are more likely to produce outputs that reflect these biases."

From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).

Classification

Causal entity
Other

Subdomain definition: Unequal treatment of individuals or groups by AI, often based on race, gender, or other sensitive characteristics, resulting in unfair outcomes and representation of those groups.

Real-world incidents in this subdomain

Browse all incidents in this subdomain

How other frameworks describe this risk

  • Discrimination

    A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms (Abercrombie2024)

  • Harms of Representation and Other Biases

    Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  • Risks from bias and underrepresentation

    International Scientific Report on the Safety of Advanced AI (Bengio2024)

  • Bias

    International AI Safety Report 2025 (Bengio2025)

  • Bias

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Toxicity and Bias Tendencies

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Biased Training Data

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Broken systems

    Navigating the Landscape of AI Ethics and Responsibility (Cunha2023)

Other entries from G'sell2024