AIPolicyTracker

MIT AI Risk Repository · Risk Sub-Category · 45.01.08

Risks from data (Risks of improper content and poisoning in training data)

Category: AI's inherent safety risks

Description

"If the training data includes illegal or harmful information, such as false, biased, or IPR-infringing content, or lacks diversity in its sources, the output may include harmful content like illegal, malicious, or extreme information. Training data is also at risk of being poisoned through tampering, error injection, or misleading actions by attackers. This can interfere with the model's probability distribution, reducing its accuracy and reliability."

From AI Safety Governance Framework (TC2602024), as extracted by the MIT AI Risk Repository (CC BY 4.0).

Classification

Causal entity
Human
Intent
Other

Subdomain definition: AI exposing users to harmful, abusive, unsafe or inappropriate content. May involve AI creating, describing, providing advice, or encouraging action. Examples of toxic content include hate-speech, violence, extremism, illegal acts, child sexual abuse material, as well as content that violates community norms such as profanity, inflammatory political speech, or pornography.

Real-world incidents in this subdomain

Browse all incidents in this subdomain

How other frameworks describe this risk

Other entries from TC2602024