AIPolicyTracker

AI incident ·

Common Biases of Vector Embeddings

1 news report Synced from source · record last edited 5 Sep 2026

In brief

An AI system built by Microsoft Research, Google and 1 other and deployed by Microsoft Research and Boston University allegedly harmed Women and girls, Women and 1 other.

Risk domain
Discrimination and Toxicity Unfair discrimination and misrepresentation
Occurred
Coverage
1 reportJul 2016

What happened

Researchers from Boston University and Microsoft Research, New England demonstrated gender bias in the most common techniques used to embed words for natural language processing (NLP).

Laws that address this harm

Policy angle: Classified under Discrimination and Toxicity (Unfair discrimination and misrepresentation) in the MIT AI Risk Repository taxonomy; 5 recorded instruments address this use case in the United States.

Matched from the record's risk domain and country to the instruments recorded here. A reviewer can correct the match in the repository (data/external/incident_overrides.yaml).

News reports (1)

Titles link to the original publisher; report text is not reproduced here.

  1. Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings
    arxiv.org · Tolga Bolukbasi, Kai-Wei Chang, James Zou

Who was involved

Alleged harmed party
Women and girls Women Minority groups

AI systems implicated

Word embedding models

Classification (MIT AI Risk Repository taxonomy)

Causal entity
AI
Intent
Unintentional
Timing
Post-deployment
Harm level
none
Sectors
professional, scientific and technical activities
Countries
US

Risk entries describing this failure mode

Entries from the MIT AI Risk Repository coded to subdomain 1.1.

  • Discrimination

    "Discrimination - Unfair or inadequate treatment or arbitrary distinction based on a person’s race, ethnicity, age, gender, sexual preference, religion, national origin, marital status, disability, language, or other pro...

    A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms (Abercrombie2024)

  • Harms of Representation and Other Biases

    "A pretrained LLM generally has many of the stereotypical biases commonly present in the human society (Touvron et al., 2023). This makes it difficult for users to trust that LLMs will work well for them and not produce...

    Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  • Risks from bias and underrepresentation

    "The outputs and impacts of general- purpose AI systems can be biased with respect to various aspects of human identity, including race, gender, culture, age, and disability. This creates risks in high- stakes domains su...

    International Scientific Report on the Safety of Advanced AI (Bengio2024)

  • Bias

    "General-purpose AI systems can amplify social and political biases, causing concrete harm. They frequently display biases with respect to race, gender, culture, age, disability, political opinion, or other aspects of hu...

    International AI Safety Report 2025 (Bengio2025)

  • Biased Training Data

    "Compared with the definition of toxicity, the definition of bias is more subjective and contextdependent. Based on previous work [97], [101], we describe the bias as disparities that could raise demographic differences...

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Toxicity and Bias Tendencies

    "Extensive data collection in LLMs brings toxic content and stereotypical bias into the training data."

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Bias

    "The training datasets of LLMs may contain biased information that leads LLMs to generate outputs with social biases"

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Broken systems

    "These are the most mentioned cases. They refer to situations where the algorithm or the training data lead to unreliable outputs. These systems frequently assign disproportionate weight to some variables, like race or g...

    Navigating the Landscape of AI Ethics and Responsibility (Cunha2023)

Linked by editors or by text similarity in the source dataset.

Incidents in the same risk subdomain

All incidents in this subdomain

Source record: incident #12 on the AI Incident Database · all 1 report