MIT AI Risk Repository · Risk Sub-Category · 63.06.02

Undesirable Dispositions from Human Data

Category: Selection Pressures

Description

"Undesirable Dispositions from Human Data. It is well-understood that models trained on human data – such as being pre-trained on human-written text or fine-tuned on human feedback – can exhibit human biases. For these reasons, there has already been considerable attention to measuring biases related to protected characteristics such as sex and ethnicity (e.g., Ferrara, 2023; Liang et al., 2021; Nadeem et al., 2020; Nangia et al., 2020), which can be amplified in multi-agent settings (Acerbi & Stubbersfield, 2023, see also Case Study 7). More recently, there has been increasing attention paid

From Multi-Agent Risks from Advanced AI (Hammond2025), as extracted by the MIT AI Risk Repository (CC BY 4.0).

Classification

Causal entity
Other

Subdomain definition: Risks from multi-agent interactions, due to incentives (which can lead to conflict or collusion) and/or the structure of multi-agent systems, which can create cascading failures, selection pressures, new security vulnerabilities, and a lack of shared information and trust.

How other frameworks describe this risk

Other entries from Hammond2025