AI incident #827 ·
AI Transcription Tool Whisper Reportedly Inserting Fabricated Content in Medical Transcripts
What happened
OpenAI's AI-powered transcription tool Whisper, used to translate and transcribe audio content such as patient consultations with doctors, is advertised as having near “human level robustness and accuracy.” However, software engineers, developers and academic researchers have alleged that it is prone to making up chunks of text or even entire sentences and that some of the hallucinations can include racial commentary, violent rhetoric, and even imagined medical treatments.
Only the incident metadata is stored here. The underlying news reports are on the AI Incident Database (CC BY-SA 4.0); use the links above to read them.
News reports (1)
Coverage catalogued by the AI Incident Database. Titles link to the original publisher; the text is not reproduced here.
Who was involved
Classification (MIT AI Risk Repository taxonomy)
- Risk domain
- Misinformation
- Risk subdomain
- 3.1 False or misleading information
- Causal entity
- AI
- Intent
- Unintentional
- Timing
- Post-deployment
- Harm level
- —
- Sectors
- —
- Countries
- —
Risk entries describing this failure mode
Entries from the MIT AI Risk Repository coded to subdomain 3.1.
- Pursuing Consistent Context
"LLMs have been demonstrated to pursue consistent context [129]–[132], which may lead to erroneous generation when the prefixes contain false information. Typical examples include sycophancy [129], [130], false demonstra...
- Knowledge Gaps
"Since the training corpora of LLMs can not contain all possible world knowledge [114]–[119], and it is challenging for LLMs to grasp the long-tail knowledge within their training data [120], [121], LLMs inherently posse...
- Hallucinations
"LLMs generate nonsensical, untruthful, and factual incorrect content"
- Faithfulness Errors
"The LLM-generated content could contain inaccurate information" which is is not true to the source material or input used
- Defective Decoding Process
In general, LLMs employ the Transformer architecture [32] and generate content in an autoregressive manner, where the prediction of the next token is conditioned on the previously generated token sequence. Such a scheme...
- Untruthful Content
"The LLM-generated content could contain inaccurate information"
- Factuality Errors
"The LLM-generated content could contain inaccurate information" which is factually incorrect
- Noisy Training Data
"Another important source of hallucinations is the noise in training data, which introduces errors in the knowledge stored in model parameters [111]–[113]. Generally, the training data inherently harbors misinformation....
Incidents in the same risk subdomain
- Nonfiction Book 'The Future of Truth' Reportedly Included AI-Generated and Misattributed Quotations
- Claude Console Reportedly Generated Phantom Legal Quotations in Trump Layoffs Court Filing
- Purportedly AI-Enhanced Images of Iranian Women Protesters Were Reportedly Spread With Unverified Execution Claims
- South Africa Draft National AI Policy Reportedly Included Fictitious References Believed to Be AI Hallucinations
- Purportedly AI-Generated Image Reportedly Misled Daejeon Authorities Searching for Escaped Wolf Neukgu
- Gemini and Grok Reportedly Misidentified Authentic Minab School-Strike Graveyard Photo as Unrelated Disaster Imagery
Other incidents involving Openai
- Nippon Life Alleged ChatGPT Practiced Law Without a License in Illinois Disability Case
- ChatGPT Reportedly Found to Reproduce Protected German Lyrics in Copyright Case
- Alleged Harmful Health Outcomes Following Reported Use of Purported ChatGPT-Generated Medical Advice in Hyderabad
- Lawsuit Alleged ChatGPT (GPT-4o) Encouraged Colorado Man's Suicide After Prolonged 'AI Companion' Chats
- Large-Scale Mental Health Crises Allegedly Associated with ChatGPT Interactions
- OpenAI ChatGPT Models Reportedly Jailbroken to Provide Chemical, Biological, and Nuclear Weapons Instructions