MIT AI Risk Repository · Risk Sub-Category · 19.01.03

Lack of data, poor data quality, and biases in training data

Category: Technological, Data and Analytical AI Risks

Description

From Governance of artificial intelligence: A risk and guideline-based integrative framework (Wirtz2022), as extracted by the MIT AI Risk Repository (CC BY 4.0).

Classification

Causal entity
Human

Subdomain definition: Unequal treatment of individuals or groups by AI, often based on race, gender, or other sensitive characteristics, resulting in unfair outcomes and representation of those groups.

Real-world incidents in this subdomain

Browse all incidents in this subdomain

How other frameworks describe this risk

  • Discrimination

    A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms (Abercrombie2024)

  • Harms of Representation and Other Biases

    Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  • Risks from bias and underrepresentation

    International Scientific Report on the Safety of Advanced AI (Bengio2024)

  • Bias

    International AI Safety Report 2025 (Bengio2025)

  • Bias

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Toxicity and Bias Tendencies

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Biased Training Data

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Broken systems

    Navigating the Landscape of AI Ethics and Responsibility (Cunha2023)

Other entries from Wirtz2022