MIT AI Risk Repository · Risk Sub-Category · 69.06.02
Discriminatory and exclusionary language
Category: Toxic and disrespectful content
Description
-
From Emerging Risks and Mitigations for Public Chatbots: LILAC v1 (Stanley2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Causal entity
- AI
- Intent
- Other
- Timing
- Post-deployment
Subdomain definition: Unequal treatment of individuals or groups by AI, often based on race, gender, or other sensitive characteristics, resulting in unfair outcomes and representation of those groups.
Real-world incidents in this subdomain
- DOGE Reportedly Relied on Unvetted ChatGPT Outputs in Canceling National Endowment for the Humanities Grants
- Sora Video Generator Has Reportedly Been Creating Biased Human Representations Across Race, Gender, and Disability
- Meta AI Characters Allegedly Exhibited Racism, Fabricated Identities, and Exploited User Trust
- Alleged AI-Generated Photo Alteration Leads to Inappropriate Modifications in Speaker's Conference Picture
- Algorithmic Bias in French Welfare System Allegedly Discriminates Against Marginalized Groups
- Department for Work and Pensions (DWP) AI Systems Allegedly Discriminate Against Single Mothers
How other frameworks describe this risk
Other entries from Stanley2024
- False information
- Hallucinated responses (in general)
- Hallucinated responses (in general)
- About a topic or source (which the user repeats)
- About a topic or source (which the user repeats)
- About a policy (which the user acts on)
- About a policy (which the user acts on)
- About a person or their activities
- About a person or their activities
- Spreads and self-perpetuates mis/disinformation
- Spreads and self-perpetuates mis/disinformation
- Performative utterances