AI incident #1401 ·
Washington State DOL's AI Phone System Reportedly Failed to Provide Spanish-Language Service to Callers Requesting Spanish
What happened
For months, callers to the Washington State Department of Licensing who selected Spanish reportedly received AI-generated English responses spoken with a Spanish accent rather than actual Spanish-language service. The agency reportedly apologized and said staff configuration caused the error, which purportedly created accessibility problems for callers seeking language support.
Only the incident metadata is stored here. The underlying news reports are on the AI Incident Database (CC BY-SA 4.0); use the links above to read them.
News reports (5)
Coverage catalogued by the AI Incident Database. Titles link to the original publisher; the text is not reproduced here.
Who was involved
- Alleged deployer
- Washington State Department Of Licensing, Amazon
- Alleged developer
- Amazon
- Alleged harmed party
- Spanish Language Speakers, General Public Of Washington State, General Public
Classification (MIT AI Risk Repository taxonomy)
- Risk domain
- Discrimination and Toxicity
- Risk subdomain
- 1.3 Unequal performance across groups
- Causal entity
- Human
- Intent
- Unintentional
- Timing
- Post-deployment
- Harm level
- —
- Sectors
- —
- Countries
- —
Risk entries describing this failure mode
Entries from the MIT AI Risk Repository coded to subdomain 1.3.
- Bias and discrimination (value embedding)
"Generative AI models may also be subject to the “value embedding” phenomenon.361 “Value embedding” refers to the fact that developers of generative AI models strive to minimize biased outputs by retraining their models...
- Impact on affected communities
"It is important to include the perspectives or concerns of communities that are affected by model outcomes when designing and building models. Failing to include these perspectives makes it difficult to understand the r...
- Unfair capability distribution
"Performing worse for some groups than others in a way that harms the worse-off group"
- Disparate Performance
The LLM’s performances can differ significantly across different groups of users. For example, the question-answering capability showed significant performance differences across different racial and social status groups...
- Fairness
Avoiding bias and ensuring no disparate performance
- Ideological Homogenization from Value Embedding
"The increasing integration of general purpose AI models into every-day life raises concerns around their embedded normative values. The reach of a small number of AI models to a large number of people around the world c...
- Fairness
This challenge appears when the learning model leads to a decision that is biased to some sensitive attributes... data itself could be biased, which results in unfair decisions. Therefore, this problem should be solved o...
- Quality-of-Service Harms
"These harms occur when algorithmic systems disproportionately underperform for certain groups of people along social categories of difference such as disability, ethnicity, gender identity, and race."
Incidents in the same risk subdomain
- UK Facial Recognition System Reportedly Exhibits Higher False Positive Rates for Black and Asian Subjects
- Infinite Campus AI-Driven Student Risk Model Leads to Cuts in Support for Nevada's Low-Income Schools
- Police Use of Facial Recognition Software Causes Wrongful Arrests Without Defendant Knowledge
- Department for Work and Pensions (DWP) Algorithm Wrongly Flags 200,000 for Housing Benefit Fraud
- Facewatch Reported to Have Wrongfully Flagged Home Bargains Customer as Shoplifter
- Whisper Speech-to-Text AI Reportedly Found to Create Violent Hallucinations