MIT AI Risk Repository · Risk Sub-Category · 16.01.04
Lower performance for some languages and social groups
Category: Risk area 1: Discrimination, Hate speech and Exclusion
Description
"LMs are typically trained in few languages, and perform less well in other languages [95, 162]. In part, this is due to unavailability of training data: there are many widely spoken languages for which no systematic efforts have been made to create labelled training datasets, such as Javanese which is spoken by more than 80 million people [95]. Training data is particularly missing for languages that are spoken by groups who are multilingual and can use a technology in English, or for languages spoken by groups who are not the primary target demographic for new technologies."
From Taxonomy of Risks posed by Language Models (Weidinger2022), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Causal entity
- AI
- Intent
- Unintentional
- Timing
- Post-deployment
Subdomain definition: Accuracy and effectiveness of AI decisions and actions is dependent on group membership, where decisions in AI system design and biased training data lead to unequal outcomes, reduced benefits, increased effort, and alienation of users.
Real-world incidents in this subdomain
- Washington State DOL's AI Phone System Reportedly Failed to Provide Spanish-Language Service to Callers Requesting Spanish
- UK Facial Recognition System Reportedly Exhibits Higher False Positive Rates for Black and Asian Subjects
- Infinite Campus AI-Driven Student Risk Model Leads to Cuts in Support for Nevada's Low-Income Schools
- Police Use of Facial Recognition Software Causes Wrongful Arrests Without Defendant Knowledge
- Department for Work and Pensions (DWP) Algorithm Wrongly Flags 200,000 for Housing Benefit Fraud
- Facewatch Reported to Have Wrongfully Flagged Home Bargains Customer as Shoplifter
How other frameworks describe this risk
Other entries from Weidinger2022
- Risk area 1: Discrimination, Hate speech and Exclusion
- Social stereotypes and unfair discrimination
- Social stereotypes and unfair discrimination.
- Hate speech and offensive language
- Exclusionary norms
- Exclusionary norms
- Exclusionary norms
- Exclusionary norms
- Exclusionary norms
- Lower performance for some languages and social groups
- Lower performance for some languages and social groups
- Risk area 2: Information Hazards