MIT AI Risk Repository · Risk Category · 02.01.00

Harmful Content

Description

"The LLM-generated content sometimes contains biased, toxic, and private information"

From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).

Classification

Causal entity
AI

Subdomain definition: AI exposing users to harmful, abusive, unsafe or inappropriate content. May involve AI creating, describing, providing advice, or encouraging action. Examples of toxic content include hate-speech, violence, extremism, illegal acts, child sexual abuse material, as well as content that violates community norms such as profanity, inflammatory political speech, or pornography.

Real-world incidents in this subdomain

Browse all incidents in this subdomain

How other frameworks describe this risk

  • Toxicity and Abusive Content

    Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements (Deng2023)

  • Controversial Opinions

    Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements (Deng2023)

  • Violation of social norms

    The Ethics of Advanced AI Assistants (Gabriel2024)

  • Violent Crimes

    AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons (Ghosh2024)

  • Sex-Related Crimes

    AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons (Ghosh2024)

  • Suicide & Self-Harm

    AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons (Ghosh2024)

  • Hate

    AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons (Ghosh2024)

  • Nonviolent Crimes

    AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons (Ghosh2024)

Other entries from Cui2024