MIT AI Risk Repository · Risk Sub-Category · 02.01.01

Bias

Category: Harmful Content

Description

"The training datasets of LLMs may contain biased information that leads LLMs to generate outputs with social biases"

From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).

Classification

Causal entity
AI
Timing
Other

Subdomain definition: Unequal treatment of individuals or groups by AI, often based on race, gender, or other sensitive characteristics, resulting in unfair outcomes and representation of those groups.

Real-world incidents in this subdomain

Browse all incidents in this subdomain

How other frameworks describe this risk

Other entries from Cui2024