MIT AI Risk Repository · Risk Sub-Category · 30.03.04
Disparate Performance
Category: Fairness
Description
The LLM’s performances can differ significantly across different groups of users. For example, the question-answering capability showed significant performance differences across different racial and social status groups. The fact-checking abilities can differ for different tasks and languages
From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Causal entity
- AI
- Intent
- Unintentional
- Timing
- Other
Subdomain definition: Accuracy and effectiveness of AI decisions and actions is dependent on group membership, where decisions in AI system design and biased training data lead to unequal outcomes, reduced benefits, increased effort, and alienation of users.
Real-world incidents in this subdomain
- Washington State DOL's AI Phone System Reportedly Failed to Provide Spanish-Language Service to Callers Requesting Spanish
- UK Facial Recognition System Reportedly Exhibits Higher False Positive Rates for Black and Asian Subjects
- Infinite Campus AI-Driven Student Risk Model Leads to Cuts in Support for Nevada's Low-Income Schools
- Police Use of Facial Recognition Software Causes Wrongful Arrests Without Defendant Knowledge
- Department for Work and Pensions (DWP) Algorithm Wrongly Flags 200,000 for Housing Benefit Fraud
- Facewatch Reported to Have Wrongfully Flagged Home Bargains Customer as Shoplifter