MIT AI Risk Repository · Risk Sub-Category · 34.01.04
Limitations of Human Feedback
Category: Causes of Misalignment
Description
"Limitations of Human Feedback. During the training of LLMs, inconsistencies can arise from human dataannotators (e.g., the varied cultural backgrounds of these annotators can introduce implicit biases (Peng et al.,2022)) (OpenAI, 2023a). Moreover, they might even introduce biases deliberately, leading to untruthful preferencedata (Casper et al., 2023b). For complex tasks that are hard for humans to evaluate (e.g., the value ofgame state), these challenges become even more salient (Irving et al., 2018)."
From AI Alignment: A Comprehensive Survey (Ji2023), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Subdomain
- 7.0
- Causal entity
- Human
- Intent
- Unintentional
- Timing
- Pre-deployment