MIT AI Risk Repository · Risk Sub-Category · 02.09.02

Noisy Training Data

Category: Hallucinations

Description

"Another important source of hallucinations is the noise in training data, which introduces errors in the knowledge stored in model parameters [111]–[113]. Generally, the training data inherently harbors misinformation. When training on large-scale corpora, this issue becomes more serious because it is difficult to eliminate all the noise from the massive pre-training data."

From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).

Classification

Causal entity
AI

Subdomain definition: AI systems that inadvertently generate or spread incorrect or deceptive information, which can lead to inaccurate beliefs in users and undermine their autonomy. Humans that make decisions based on false beliefs can experience physical, emotional or material harms

Real-world incidents in this subdomain

Browse all incidents in this subdomain

How other frameworks describe this risk

Other entries from Cui2024