MIT AI Risk Repository · Risk Sub-Category · 02.08.01

Toxic Training Data

Category: Toxicity and Bias Tendencies

Description

"Following previous studies [96], [97], toxic data in LLMs is defined as rude, disrespectful, or unreasonable language that is opposite to a polite, positive, and healthy language environment, including hate speech, offensive utterance, profanities, and threats [91]."

From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).

Classification

Causal entity
AI

Subdomain definition: AI exposing users to harmful, abusive, unsafe or inappropriate content. May involve AI creating, describing, providing advice, or encouraging action. Examples of toxic content include hate-speech, violence, extremism, illegal acts, child sexual abuse material, as well as content that violates community norms such as profanity, inflammatory political speech, or pornography.

Real-world incidents in this subdomain

Browse all incidents in this subdomain

How other frameworks describe this risk

  • Toxicity and Abusive Content

    Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements (Deng2023)

  • Controversial Opinions

    Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements (Deng2023)

  • Violation of social norms

    The Ethics of Advanced AI Assistants (Gabriel2024)

  • Violent Crimes

    AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons (Ghosh2024)

  • Suicide & Self-Harm

    AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons (Ghosh2024)

  • Hate

    AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons (Ghosh2024)

  • Nonviolent Crimes

    AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons (Ghosh2024)

  • Contextual Hazards

    AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons (Ghosh2024)

Other entries from Cui2024