MIT AI Risk Repository

Browse AI risks

14 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset Also filtered by framework Liu2024 ×

14 entries

  1. 30.03.01 · Risk Sub-Category

    Fairness

    Injustice

    In the context of LLM outputs, we want to make sure the suggested or completed texts are indistinguishable in nature for two involved individuals (in the prompt) with the same relevant profiles but might come from different groups (where the group attribute is regarded as being irrelevant in this context)

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  2. 30.03.02 · Risk Sub-Category

    Fairness

    Stereotype Bias

    LLMs must not exhibit or highlight any stereotypes in the generated text. Pretrained LLMs tend to pick up stereotype biases persisting in crowdsourced data and further amplify them

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  3. 30.03.03 · Risk Sub-Category

    Fairness

    Preference Bias

    LLMs are exposed to vast groups of people, and their political biases may pose a risk of manipulation of socio-political processes

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  4. 30.07.03 · Risk Sub-Category

    Robustness

    Interventional Effect

    existing disparities in data among different user groups might create differentiated experiences when users interact with an algorithmic system (e.g. a recommendation system), which will further reinforce the bias

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  5. 30.02.00 · Risk Category

    Safety

    Avoiding unsafe and illegal outputs, and leaking private information

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  6. 30.02.01 · Risk Sub-Category

    Safety

    Violence

    LLMs are found to generate answers that contain violent content or generate content that responds to questions that solicit information about violent behaviors

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  7. 30.02.02 · Risk Sub-Category

    Safety

    Unlawful Conduct

    LLMs have been shown to be a convenient tool for soliciting advice on accessing, purchasing (illegally), and creating illegal substances, as well as for dangerous use of them

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  8. 30.02.03 · Risk Sub-Category

    Safety

    Harms to Minor

    LLMs can be leveraged to solicit answers that contain harmful content to children and youth

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  9. 30.02.04 · Risk Sub-Category

    Safety

    Adult Content

    LLMs have the capability to generate sex-explicit conversations, and erotic texts, and to recommend websites with sexual content

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  10. 30.06.00 · Risk Category

    Social Norm

    LLMs are expected to reflect social values by avoiding the use of offensive language toward specific groups of users, being sensitive to topics that can create instability, as well as being sympathetic when users are seeking emotional support

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  11. 30.06.01 · Risk Sub-Category

    Social Norm

    Toxicity

    language being rude, disrespectful, threatening, or identity-attacking toward certain groups of the user population (culture, race, and gender etc)

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  12. 30.06.03 · Risk Sub-Category

    Social Norm

    Cultural Insensitivity

    it is important to build high-quality locally collected datasets that reflect views from local users to align a model’s value system

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  13. 30.03.00 · Risk Category

    Fairness

    Avoiding bias and ensuring no disparate performance

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  14. 30.03.04 · Risk Sub-Category

    Fairness

    Disparate Performance

    The LLM’s performances can differ significantly across different groups of users. For example, the question-answering capability showed significant performance differences across different racial and social status groups. The fact-checking abilities can differ for different tasks and languages

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.