AI incident #732 ·

Whisper Speech-to-Text AI Reportedly Found to Create Violent Hallucinations

What happened

Researchers at Cornell reportedly found that OpenAI's Whisper, a speech-to-text system, can hallucinate violent language and fabricated details, especially with long pauses in speech, such as from those with speech impairments. Analyzing 13,000 clips, they determined 1% contained harmful hallucinations. These errors pose risks in hiring, legal trials, and medical documentation. The study suggests improving model training to reduce these hallucinations for diverse speaking patterns.

Only the incident metadata is stored here. The underlying news reports are on the AI Incident Database (CC BY-SA 4.0); use the links above to read them.

News reports (1)

Coverage catalogued by the AI Incident Database. Titles link to the original publisher; the text is not reproduced here.

Who was involved

Alleged developer
Openai
Alleged harmed party
Users Whose Speech Is Misinterpreted By Whisper, Professionals Relying On Accurate Transcriptions, Individuals With Speech Impairments, General Public, People With Disabilities

Classification (MIT AI Risk Repository taxonomy)

Causal entity
AI
Intent
Unintentional
Timing
Post-deployment
Harm level
Sectors
Countries

Risk entries describing this failure mode

Entries from the MIT AI Risk Repository coded to subdomain 1.3.

  • Bias and discrimination (value embedding)

    "Generative AI models may also be subject to the “value embedding” phenomenon.361 “Value embedding” refers to the fact that developers of generative AI models strive to minimize biased outputs by retraining their models...

    Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024)

  • Impact on affected communities

    "It is important to include the perspectives or concerns of communities that are affected by model outcomes when designing and building models. Failing to include these perspectives makes it difficult to understand the r...

    AI Risk Atlas (IBM2025)

  • Unfair capability distribution

    "Performing worse for some groups than others in a way that harms the worse-off group"

    A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025)

  • Disparate Performance

    The LLM’s performances can differ significantly across different groups of users. For example, the question-answering capability showed significant performance differences across different racial and social status groups...

    Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  • Fairness

    Avoiding bias and ensuring no disparate performance

    Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  • Ideological Homogenization from Value Embedding

    "The increasing integration of general purpose AI models into every-day life raises concerns around their embedded normative values. The reach of a small number of AI models to a large number of people around the world c...

    Governing General Purpose AI: A Comprehensive Map of Unreliability, Misuse and Systemic Risks (Maham2023 )

  • Fairness

    This challenge appears when the learning model leads to a decision that is biased to some sensitive attributes... data itself could be biased, which results in unfair decisions. Therefore, this problem should be solved o...

    A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions (Saghiri2022)

  • Quality-of-Service Harms

    "These harms occur when algorithmic systems disproportionately underperform for certain groups of people along social categories of difference such as disability, ethnicity, gender identity, and race."

    Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction (Shelby2023)

Incidents in the same risk subdomain

All incidents in this subdomain