MIT AI Risk Repository · Risk Sub-Category · 34.01.04

Limitations of Human Feedback

Category: Causes of Misalignment

Description

"Limitations of Human Feedback. During the training of LLMs, inconsistencies can arise from human dataannotators (e.g., the varied cultural backgrounds of these annotators can introduce implicit biases (Peng et al.,2022)) (OpenAI, 2023a). Moreover, they might even introduce biases deliberately, leading to untruthful preferencedata (Casper et al., 2023b). For complex tasks that are hard for humans to evaluate (e.g., the value ofgame state), these challenges become even more salient (Irving et al., 2018)."

From AI Alignment: A Comprehensive Survey (Ji2023), as extracted by the MIT AI Risk Repository (CC BY 4.0).

Classification

Subdomain
7.0
Causal entity
Human

How other frameworks describe this risk

Other entries from Ji2023