MIT AI Risk Repository · Risk Category · 39.03.00
Data Issues
Description
Data heterogeneity, data insufficiency, imbalanced data, untrusted data, biased data, and data uncertainty are other data issues that may cause various difficulties in datadriven machine learning algorithms.. Bias is a human feature that may affect data gathering and labeling. Sometimes, bias is present in historical, cultural, or geographical data. Consequently, bias may lead to biased models which can provide inappropriate analysis. Despite being aware of the existence of bias, avoiding biased models is a challenging task
From A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions (Saghiri2022), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Causal entity
- AI
- Intent
- Unintentional
- Timing
- Other
Subdomain definition: Unequal treatment of individuals or groups by AI, often based on race, gender, or other sensitive characteristics, resulting in unfair outcomes and representation of those groups.
Real-world incidents in this subdomain
- DOGE Reportedly Relied on Unvetted ChatGPT Outputs in Canceling National Endowment for the Humanities Grants
- Sora Video Generator Has Reportedly Been Creating Biased Human Representations Across Race, Gender, and Disability
- Meta AI Characters Allegedly Exhibited Racism, Fabricated Identities, and Exploited User Trust
- Alleged AI-Generated Photo Alteration Leads to Inappropriate Modifications in Speaker's Conference Picture
- Algorithmic Bias in French Welfare System Allegedly Discriminates Against Marginalized Groups
- Department for Work and Pensions (DWP) AI Systems Allegedly Discriminate Against Single Mothers