MIT AI Risk Repository

Browse AI risks

7 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset Also filtered by framework Cui2024 ×

7 entries

  1. 02.01.01 · Risk Sub-Category

    Harmful Content

    Bias

    "The training datasets of LLMs may contain biased information that leads LLMs to generate outputs with social biases"

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  2. "Extensive data collection in LLMs brings toxic content and stereotypical bias into the training data."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  3. 02.08.02 · Risk Sub-Category

    Toxicity and Bias Tendencies

    Biased Training Data

    "Compared with the definition of toxicity, the definition of bias is more subjective and contextdependent. Based on previous work [97], [101], we describe the bias as disparities that could raise demographic differences among various groups, which may involve demographic word prevalence and stereotypical contents. Concretely, in massive corpora, the prevalence of different pronouns and identities could influence an LLM’s tendency about gender, nationality, race, religion, and culture [4]. For instance, the pronoun He is over-represented compared with the pronoun She in the training corpora, le

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  4. 02.01.00 · Risk Category

    Harmful Content

    "The LLM-generated content sometimes contains biased, toxic, and private information"

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  5. 02.01.02 · Risk Sub-Category

    Harmful Content

    Toxicity

    "Toxicity means the generated content contains rude, disrespectful, and even illegal information"

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  6. 02.08.01 · Risk Sub-Category

    Toxicity and Bias Tendencies

    Toxic Training Data

    "Following previous studies [96], [97], toxic data in LLMs is defined as rude, disrespectful, or unreasonable language that is opposite to a polite, positive, and healthy language environment, including hate speech, offensive utterance, profanities, and threats [91]."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  7. "Inputting a prompt contain an unsafe topic (e.g., notsuitable-for-work (NSFW) content) by a benign user. "

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.