MIT AI Risk Repository · Risk Sub-Category · 62.18.04
Biases are not accurately reflected in explanations
Category: Model Evaluations (Interpretability/Explainability)
Description
"Existing explainability techniques can be insufficient for detecting discriminatory biases. Manipulation methods can hide underlying biases from these tech- niques, generating misleading explanations [192, 112]. Such explanations ex- clude sensitive or prohibitive attributes, such as race or gender, and instead include desired attributes, even though they do not accurately represent the underlying model."
From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
Subdomain definition: Unequal treatment of individuals or groups by AI, often based on race, gender, or other sensitive characteristics, resulting in unfair outcomes and representation of those groups.
Real-world incidents in this subdomain
- DOGE Reportedly Relied on Unvetted ChatGPT Outputs in Canceling National Endowment for the Humanities Grants
- Sora Video Generator Has Reportedly Been Creating Biased Human Representations Across Race, Gender, and Disability
- Meta AI Characters Allegedly Exhibited Racism, Fabricated Identities, and Exploited User Trust
- Alleged AI-Generated Photo Alteration Leads to Inappropriate Modifications in Speaker's Conference Picture
- Algorithmic Bias in French Welfare System Allegedly Discriminates Against Marginalized Groups
- Department for Work and Pensions (DWP) AI Systems Allegedly Discriminate Against Single Mothers