MIT AI Risk Repository · domain 1: Discrimination & Toxicity
1.1 Unfair discrimination and misrepresentation
Unequal treatment of individuals or groups by AI, often based on race, gender, or other sensitive characteristics, resulting in unfair outcomes and representation of those groups.
- 83
- 12
- 118
- 58
| Label | Value |
|---|---|
| AI | 58 |
| Other | 13 |
| Human | 11 |
| Not coded | 1 |
| Label | Value |
|---|---|
| Unintentional | 64 |
| Other | 16 |
| Intentional | 2 |
| Not coded | 1 |
| Label | Value |
|---|---|
| Post-deployment | 50 |
| Other | 19 |
| Pre-deployment | 13 |
| Not coded | 1 |
| Label | Value |
|---|---|
| 2012 | 5 |
| 2013 | 2 |
| 2014 | 2 |
| 2015 | 6 |
| 2016 | 13 |
| 2017 | 9 |
| 2018 | 9 |
| 2019 | 8 |
| 2020 | 19 |
| 2021 | 8 |
| 2022 | 10 |
| 2023 | 11 |
| 2024 | 7 |
| 2025 | 3 |
| Label | Value |
|---|---|
| Risk Category | 22 |
| Risk Sub-Category | 61 |
Risk entries
Browse and export all- Discrimination
"Discrimination - Unfair or inadequate treatment or arbitrary distinction based on a person’s race, ethnicity, age, gender, sexual preference, religion, national origin, marital status, disability, la...
- Harms of Representation and Other Biases
"A pretrained LLM generally has many of the stereotypical biases commonly present in the human society (Touvron et al., 2023). This makes it difficult for users to trust that LLMs will work well for t...
- Risks from bias and underrepresentation
"The outputs and impacts of general- purpose AI systems can be biased with respect to various aspects of human identity, including race, gender, culture, age, and disability. This creates risks in hig...
- Bias
"General-purpose AI systems can amplify social and political biases, causing concrete harm. They frequently display biases with respect to race, gender, culture, age, disability, political opinion, or...
- Bias
"The training datasets of LLMs may contain biased information that leads LLMs to generate outputs with social biases"
- Toxicity and Bias Tendencies
"Extensive data collection in LLMs brings toxic content and stereotypical bias into the training data."
- Biased Training Data
"Compared with the definition of toxicity, the definition of bias is more subjective and contextdependent. Based on previous work [97], [101], we describe the bias as disparities that could raise demo...
- Broken systems
"These are the most mentioned cases. They refer to situations where the algorithm or the training data lead to unreliable outputs. These systems frequently assign disproportionate weight to some varia...
- Unfairness and Discrimination
Social bias is an unfairly negative attitude towards a social group or individuals based on one-sided or inaccurate information, typically pertaining to widely disseminated negative stereotypes regard...
- Bias, Fairness and Representational Harms
"Frontier AI models can contain and magnify biases ingrained in the data they are trained on, reflecting societal and historical inequalities and stereotypes.177 These biases, often subtle and deeply...
- Algorithm and data
"More than 20% of the contributions are centered on the ethical dimensions of algorithms and data. This theme can be further categorized into two main subthemes: data bias and algorithm fairness, and...
- Discrimination and bias
- Biases are not accurately reflected in explanations
"Existing explainability techniques can be insufficient for detecting discriminatory biases. Manipulation methods can hide underlying biases from these tech- niques, generating misleading explanations...
- Biases in AI-based content moderation algorithms
"AI-based content moderation algorithms, while intended to filter harmful con- tent, can perpetuate biases. For example, gender biases within these systems may lead to the disproportionate suppression...
- Systemic bias across specific communities
"AI systems may exhibit unfair or unfavorable outputs across a range of tasks against specific communities of people, either implicitly or explicitly. Bias can lead to forms of exclusion or erasure (e...
- Unintentional bias amplification
"Dataset bias may be unintentionally amplified [60] where the outputs of the AI model trained on a dataset are more biased than the dataset itself."
- Discrimination
"More broadly, bad decisions or errors by AI tools could lead to discrimination or deeper inequality"
- Amplification of biases
"Current Frontier AI mdoels amplify existing biases within their training data and can be manipulated into providing potentially harmful responses, for example abusive language or discriminatory respo...
- Bias and discrimination (bias in training datasets)
"AI experts consider training data to be the most salient source of bias in generative AI models. For example, GPT- 2’s training data comes from outbound links from Reddit, a social network often crit...
- Bias and discrimination (value lock and outcome homogenization)
"Because models are not necessarily retrained to reflect evolving societal views, language models risk “value lock- ins,” which “reifies older, less inclusive understandings.”370 Therefore, the contin...
- Bias and Discrimination
as they claim to generate biased and discriminatory results, these AI systems have a negative impact on the rights of individuals, principles of adjudication, and overall judicial integrity
- Fairness - Bias
Fairness is, by far, the most discussed issue in the literature, remaining a paramount concern especially in case of LLMs and text-to-image models. This is sparked by training data biases propagating...
- Discrimination
"When AI is not carefully designed, it can discriminate against certain groups."
- Bias
"The AI will only be as good as the data it is trained with. If the data contains bias (and much data does), then the AI will manifest that bias, too."
- Data bias
"Historical and societal biases that are present in the data are used to train and fine-tune the model."