MIT AI Risk Repository · Risk Sub-Category · 16.05.01
Promoting harmful stereotypes by implying gender or ethnic identity
Category: Risk area 5: Human-Computer Interaction Harms
Description
"CAs can perpetuate harmful stereotypes by using particular identity markers in language (e.g. referring to “self” as “female”), or by more general design features (e.g. by giving the product a gendered name such as Alexa). The risk of representational harm in these cases is that the role of “assistant” is presented as inherently linked to the female gender [19, 36]. Gender or ethnicity identity markers may be implied by CA vocabulary, knowledge or vernacular [124]; product description, e.g. in one case where users could choose as virtual assistant Jake - White, Darnell - Black, Antonio - Hisp
From Taxonomy of Risks posed by Language Models (Weidinger2022), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Causal entity
- AI
- Intent
- Unintentional
- Timing
- Post-deployment
Subdomain definition: Unequal treatment of individuals or groups by AI, often based on race, gender, or other sensitive characteristics, resulting in unfair outcomes and representation of those groups.
Real-world incidents in this subdomain
- DOGE Reportedly Relied on Unvetted ChatGPT Outputs in Canceling National Endowment for the Humanities Grants
- Sora Video Generator Has Reportedly Been Creating Biased Human Representations Across Race, Gender, and Disability
- Meta AI Characters Allegedly Exhibited Racism, Fabricated Identities, and Exploited User Trust
- Alleged AI-Generated Photo Alteration Leads to Inappropriate Modifications in Speaker's Conference Picture
- Algorithmic Bias in French Welfare System Allegedly Discriminates Against Marginalized Groups
- Department for Work and Pensions (DWP) AI Systems Allegedly Discriminate Against Single Mothers
How other frameworks describe this risk
Other entries from Weidinger2022
- Risk area 1: Discrimination, Hate speech and Exclusion
- Social stereotypes and unfair discrimination
- Social stereotypes and unfair discrimination.
- Hate speech and offensive language
- Exclusionary norms
- Exclusionary norms
- Exclusionary norms
- Exclusionary norms
- Exclusionary norms
- Lower performance for some languages and social groups
- Lower performance for some languages and social groups
- Lower performance for some languages and social groups