AI incident #102 ·

Personal voice assistants struggle with black voices, new study shows

What happened

A study found that voice recognition tools from Apple, Amazon, Google, IBM, and Microsoft disproportionately made errors when transcribing black speakers.

Only the incident metadata is stored here. The underlying news reports are on the AI Incident Database (CC BY-SA 4.0); use the links above to read them.

News reports (2)

Coverage catalogued by the AI Incident Database. Titles link to the original publisher; the text is not reproduced here.

  1. Racial disparities in automated speech recognition
    pnas.org · Allison Koenecke, Andrew Nam, Emily Lake · AIID #1523

Who was involved

Alleged harmed party
Black People

Classification (MIT AI Risk Repository taxonomy)

Causal entity
AI
Intent
Unintentional
Timing
Post-deployment
Harm level
none
Sectors
administrative and support service activities, information and communication
Countries
US

Risk entries describing this failure mode

Entries from the MIT AI Risk Repository coded to subdomain 1.3.

  • Bias and discrimination (value embedding)

    "Generative AI models may also be subject to the “value embedding” phenomenon.361 “Value embedding” refers to the fact that developers of generative AI models strive to minimize biased outputs by retraining their models...

    Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024)

  • Impact on affected communities

    "It is important to include the perspectives or concerns of communities that are affected by model outcomes when designing and building models. Failing to include these perspectives makes it difficult to understand the r...

    AI Risk Atlas (IBM2025)

  • Unfair capability distribution

    "Performing worse for some groups than others in a way that harms the worse-off group"

    A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025)

  • Fairness

    Avoiding bias and ensuring no disparate performance

    Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  • Disparate Performance

    The LLM’s performances can differ significantly across different groups of users. For example, the question-answering capability showed significant performance differences across different racial and social status groups...

    Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  • Ideological Homogenization from Value Embedding

    "The increasing integration of general purpose AI models into every-day life raises concerns around their embedded normative values. The reach of a small number of AI models to a large number of people around the world c...

    Governing General Purpose AI: A Comprehensive Map of Unreliability, Misuse and Systemic Risks (Maham2023 )

  • Fairness

    This challenge appears when the learning model leads to a decision that is biased to some sensitive attributes... data itself could be biased, which results in unfair decisions. Therefore, this problem should be solved o...

    A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions (Saghiri2022)

  • Erasing social groups

    people, attributes, or artifacts associated with specific social groups are systematically absent or under-represented... Design choices [143] and training data [212] influence which people and experiences are legible to...

    Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction (Shelby2023)

Incidents in the same risk subdomain

All incidents in this subdomain